Method and device for drivable area detection based on enhanced disparity map and multi-scale uncertainty perception
The method uses enhanced disparity maps and multi-scale uncertainty perception to improve drivable area segmentation in autonomous vehicles, addressing lighting and road marking challenges for accurate decision-making.
Patent Information
- Application Number
- CN202310879474.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-18
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-07-18
AI Technical Summary
The existing method of dividing the driving area is prone to error judgments when the lighting conditions are poor or the road markings are unclear, which affects the safety of autonomous driving.
Using a feasible area detection method based on enhanced parallax map and multi-scale uncertainty perception, images are acquired through binocular cameras, feature enhancement is performed using stereo matching algorithms and surface normal vector methods, and a feasible area segmentation model is constructed, including an encoder, multi-scale uncertainty perception module and decoder, feature extraction and weight reassignment are performed, and the segmentation result of the feasible area is finally obtained.
It improves the segmentation accuracy and robustness in complex environments, effectively improves the segmentation accuracy of traditional methods under light changes, enhances the distinction between the feasible and the inoperable areas, and improves the safety of autonomous driving.
Smart Images

Figure CN116824534B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of autonomous driving, and particularly to a drivable area detection method and device based on an enhanced disparity map and multi-scale uncertainty perception. Background Art
[0002] Autonomous driving technology is a rapidly developing field in recent years, and its application scope involves many aspects, such as intelligent transportation systems, automobile manufacturing, and urban planning. Among them, in autonomous driving, drivable area segmentation is a key issue, which can effectively distinguish the driving area of a vehicle from the non-driving area, thereby helping the vehicle make more accurate decisions. At present, the existing drivable area segmentation methods are mainly based on computer vision technology and deep learning models. These methods mainly process the images captured by on-vehicle cameras, extract image features, and determine whether each pixel belongs to the driving area or the non-driving area. However, these methods still have certain limitations when processing images in different environments. For example, incorrect judgments may occur when the lighting conditions are poor, the light intensity changes, or the road markings are unclear, which seriously affects the safety of autonomous driving. Therefore, it is of great significance to design a method that can cope with complex environmental changes, especially the influence of lighting, and has both high robustness and high segmentation accuracy. Summary of the Invention
[0003] Aiming at the above-mentioned technical problems, the purpose of the embodiments of the present application is to propose a drivable area detection method and device based on an enhanced disparity map and multi-scale uncertainty perception to solve the technical problems mentioned in the above background art section.
[0004] In a first aspect, the present invention provides a drivable area detection method based on an enhanced disparity map and multi-scale uncertainty perception, including the following steps:
[0005] Obtain binocular images of a road collected by a binocular camera, correct the binocular images to obtain corrected images, and process the corrected images using a stereo matching algorithm to obtain a disparity map;
[0006] Enhance the features of the disparity map using the surface normal vector method to obtain a surface normal vector map;
[0007] Construct and train a drivable area segmentation model to obtain a trained drivable area segmentation model. The drivable area segmentation model includes an encoder module, a multi-scale uncertainty perception module, and a decoder module connected in sequence. The encoder module is used to extract features from the surface normal vector map to obtain a first feature vector. The input multi-scale uncertainty perception module is used to reset the weights corresponding to each pixel point in the first feature vector to obtain a second feature vector. The decoder module is used to decode the second feature vector to obtain the segmentation result of the drivable area;
[0008] Input the surface normal vector map into the trained drivable area segmentation model to obtain the segmentation result of the drivable area.
[0009] Preferably, the calibration includes distortion calibration and stereo calibration.
[0010] Preferably, the surface normal vector method is used to enhance the features of the disparity map to obtain the surface normal vector map, specifically including:
[0011] Perform convolution operations on each pixel point in the disparity map using a horizontal image gradient filter and a vertical image gradient filter to obtain the corresponding horizontal normal vector and vertical normal vector:
[0012]
[0013]
[0014] where D represents the disparity map, and fx and fy represent the camera focal lengths in the x direction and y direction respectively;
[0015] For any point A on the disparity map, take several adjacent pixel points and represent them as a set: N A =[G1,G2,......,G m T , and calculate the distance sets between several pixel points and point A in the X, Y, and Z directions: G m -A = [ΔX m ,ΔY m ,ΔZ m T , where m ∈ [1,12];
[0016] Select one of the pixel points adjacent to point A as a reference, and combine the horizontal normal vector S X and the vertical normal vector S Y to obtain an expression form of the normal vector S Z of point A in the Z direction:
[0017]
[0018] Convert S Z into the form of spherical coordinates :
[0019]
[0020] where α is the inclination angle and β is the azimuth angle;
[0021] The inclination angle α is expressed as:
[0022]
[0023] Among them, i represents the i-th pixel point in the disparity map, i = 1, 2, ……, n, where n is the total number of pixel points in the disparity map, and S Zi represents the spherical coordinate in the Z direction; P i = S xi cosβ + S yi sinβ; S xi represents the spherical coordinate in the X direction; S yi represents the spherical coordinate in the Y direction;
[0024] The azimuth angle β is expressed as:
[0025]
[0026] Repeat the above steps for each pixel point in the disparity map to obtain the corresponding surface normal vector map.
[0027] Preferably, the encoder module is a pyramid structure composed of a fully connected layer, a batch normalization layer, a ReLu activation function layer, a max pooling layer, and four encoder layers connected in sequence. The encoder layer uses ResNet-50 as the Backbone, and each layer corresponds to the corresponding layer of ResNet.
[0028] Preferably, in the multi-scale uncertainty perception module, the first feature vector is convolved at three different scales respectively to obtain three sets of convolution results. The three sets of convolution structures are respectively upsampled, and then through the Softplus activation function, feature maps at three scales are obtained: {[F0] l , [F1] l}, l ∈ 1, 2, 3, where F1 is the drivable area and F0 is the non-drivable area;
[0029] Perform an averaging operation on the feature maps at three scales to obtain the average value of the feature maps:
[0030]
[0031] According to F1 and F0, obtain the Dirichlet intensity map S:
[0032]
[0033] According to the Dirichlet intensity map S, F1, and F0, calculate the credibility B0 of the drivable area, the credibility B1 of the drivable area, and the uncertainty parameter U:
[0034]
[0035]
[0036]
[0037] Optimize the credibility B0 of the driving area, the credibility B1 of the drivable area, and the uncertainty parameter U according to the Dempster combination rule to obtain the optimized credibility b0, b1, and the optimized uncertainty parameter u:
[0038]
[0039]
[0040]
[0041] Calculate the new Dirichlet intensity map S' based on the optimized uncertainty parameter u:
[0042]
[0043] Calculate the non-drivable area attention parameter δ0 and the drivable area attention parameter δ1 based on the new Dirichlet intensity map S' and the optimized credibility b0, b1:
[0044]
[0045]
[0046] Among them, represents element-wise multiplication, that is, the multiplication operation is performed on each corresponding pixel point one by one;
[0047] Calculate the drivable area probability map P based on the drivable area attention parameter δ1 and the new Dirichlet intensity map:
[0048] Reassign the weight of each pixel point of the feature vector based on the drivable area probability map P as a reference to obtain the second feature vector.
[0049] Preferably, perform convolution on the first feature vector at three different scales, specifically including:
[0050] Perform convolution operations on the first feature vector using convolutional layers with kernel sizes of 1×1, 3×3, 3×3 and dilation rates of 0, 2, 5 respectively.
[0051] Preferably, the decoder module includes five decoder layers connected in sequence and a Softmax activation layer, and the decoder layer includes a transposed convolution layer, a batch normalization layer, a skip connection layer, and an activation function layer.
[0052] Second aspect, the present invention provides a drivable area detection device based on an enhanced disparity map and multi-scale uncertainty perception, including:
[0053] A disparity map acquisition module configured to acquire binocular images of a road collected by a binocular camera, correct the binocular images to obtain corrected images, and process the corrected images using a stereo matching algorithm to obtain a disparity map;
[0054] A feature enhancement module configured to enhance the features of the disparity map using the surface normal vector method to obtain a surface normal vector map;
[0055] A model construction module configured to construct and train a drivable area segmentation model to obtain a trained drivable area segmentation model. The drivable area segmentation model includes an encoder module, a multi-scale uncertainty perception module, and a decoder module connected in sequence. The encoder module is used to extract features from the surface normal vector map to obtain a first feature vector. The input multi-scale uncertainty perception module is used to reset the weights corresponding to each pixel point in the first feature vector to obtain a second feature vector. The decoder module is used to decode the second feature vector to obtain the segmentation result of the drivable area;
[0056] An execution module configured to input the surface normal vector map into the trained drivable area segmentation model to obtain the segmentation result of the drivable area.
[0057] Third aspect, the present invention provides an electronic device, including one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner of the first aspect.
[0058] Fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in any implementation manner of the first aspect is implemented.
[0059] Compared with the prior art, the present invention has the following beneficial effects:
[0060] (1) The present invention proposes a drivable area detection method based on an enhanced disparity map and multi-scale uncertainty perception. The method uses a binocular camera to obtain binocular images of a road in real time, and processes the captured binocular images using a binocular stereo matching algorithm to obtain a disparity map, effectively improving the problem that traditional RGB images are easily affected by light changes and complex environments, resulting in low segmentation accuracy.
[0061] (2) The present invention proposes a drivable area detection method based on enhanced disparity map and multi-scale uncertainty perception. The surface normal vector method is used to enhance the features of the input disparity map to obtain a surface normal vector map, thereby enhancing the image features of the disparity map, facilitating feature extraction, making the distinguishing features between the drivable area and the non-drivable area more obvious, and effectively improving the accuracy of the model's feature recognition of the input data.
[0062] (3) The present invention proposes a drivable area detection method based on enhanced disparity map and multi-scale uncertainty perception. The surface normal vector map is input into the decoder to obtain a first feature vector. The uncertainty perception module is used to process the drivable area and the non-drivable area, changing the energy intensity corresponding to its channels. Finally, the decoder restores the second feature vector to a pixel-level output with the same size as the original image to obtain the final detection result. The multi-scale uncertainty perception module extracts features from the first feature vector at three different scales and converts them into energy intensity representations. By calculating the credibility, the attention parameters of the drivable area are obtained. Finally, by calculating the drivable area probability map, the weights of the input feature vectors are redistributed, strengthening the weights corresponding to the drivable area, and improving the edge segmentation effect and segmentation accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0064] Figure 1 It is an exemplary device architecture diagram to which an embodiment of the present application can be applied;
[0065] Figure 2 It is a flowchart of the drivable area detection method based on enhanced disparity map and multi-scale uncertainty perception according to the embodiment of the present application;
[0066] Figure 3 It is a structural diagram of the drivable area segmentation model of the drivable area detection method based on enhanced disparity map and multi-scale uncertainty perception according to the embodiment of the present application;
[0067] Figure 4 It is a structural diagram of the multi-scale uncertainty perception module of the drivable area detection method based on enhanced disparity map and multi-scale uncertainty perception according to the embodiment of the present application;
[0068] Figure 5Schematic diagram of a drivable area detection device based on an enhanced disparity map and multi-scale uncertainty perception according to an embodiment of the present application;
[0069] Figure 6 It is a schematic structural diagram of a computer device of an electronic device suitable for implementing the embodiments of the present application. Detailed implementation manners
[0070] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0071] Figure 1 An exemplary device architecture 100 is shown that can apply the drivable area detection method based on an enhanced disparity map and multi-scale uncertainty perception or the drivable area detection device based on an enhanced disparity map and multi-scale uncertainty perception according to the embodiments of the present application.
[0072] As Figure 1 shown, the device architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0073] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various applications may be installed on the terminal devices 101, 102, 103, such as data processing applications, file processing applications, etc.
[0074] The terminal devices 101, 102, 103 may be hardware or software. When the terminal devices 101, 102, 103 are hardware, they may be various electronic devices, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc. When the terminal devices 101, 102, 103 are software, they may be installed in the above-listed electronic devices. It may be implemented as multiple software or software modules (such as software or software modules for providing distributed services), or may be implemented as a single software or software module. No specific limitation is made here.
[0075] The server 105 can be a server that provides various services, such as a background data processing server for processing files or data uploaded by the terminal devices 101, 102, and 103. The background data processing server can process the acquired files or data to generate a processing result.
[0076] It should be noted that the method for detecting a drivable area based on an enhanced disparity map and multi-scale uncertainty perception provided by the embodiments of the present application can be executed by the server 105, or can be executed by the terminal devices 101, 102, and 103. Correspondingly, the device for detecting a drivable area based on an enhanced disparity map and multi-scale uncertainty perception can be set in the server 105, or can be set in the terminal devices 101, 102, and 103.
[0077] It should be understood that Figure 1 the numbers of the terminal devices, network, and server in
[0078] Figure 2 FIG. shows a method for detecting a drivable area based on an enhanced disparity map and multi-scale uncertainty perception provided by an embodiment of the present application, which is characterized by including the following steps:
[0079] S1. Obtain binocular images of a road collected by a binocular camera, correct the binocular images to obtain corrected images, and process the corrected images using a stereo matching algorithm to obtain a disparity map.
[0080] In a specific embodiment, the correction includes distortion correction and stereo correction.
[0081] Specifically, the binocular camera captures binocular images of the road passed by the car in real time, and a set of corresponding binocular images can be obtained each time, and the binocular images are transmitted to the computer of the car for subsequent processing. The binocular camera should be fixed on the top of the car, and the field of view needs to cover all the roads in front of the car to ensure that it will not be blocked by the components of the car itself. The baseline of the binocular camera needs to be determined according to requirements during installation, and then its internal and external parameters are calibrated. If the baseline needs to be adjusted subsequently, the internal and external parameters need to be recalibrated after the adjustment. The binocular camera captures the binocular images in front at a speed of 15 FPS / second, and this parameter value can be adjusted according to the actual situation.
[0082] Combining the binocular images with the internal and external parameters of the camera, the binocular images are subjected to distortion correction and stereo correction to eliminate the parallax in the horizontal direction, so that the matching points correspond to the pixel positions on the same object surface seen by the two cameras. After correction, the coordinate systems of the left and right cameras become consistent, thus simplifying the subsequent calculation process. At the same time, the corrected images have the same resolution and image size, facilitating the processing and optimization of subsequent algorithms. Then, a stereo matching algorithm is used to process the corrected images to obtain a disparity map.
[0083] S2. Use the surface normal vector method to enhance the features of the disparity map to obtain a surface normal vector map.
[0084] In a specific embodiment, step S2 specifically includes:
[0085] Perform convolution operations on each pixel point in the disparity map using a horizontal image gradient filter and a vertical image gradient filter to obtain the corresponding horizontal normal vector and vertical normal vector:
[0086]
[0087]
[0088] where D represents the disparity map, and fx and fy respectively represent the camera focal lengths in the x - direction and y - direction;
[0089] For any point A on the disparity map, take a number of adjacent pixel points, which are represented as a set: N A =[G1,G2,......,G m T , and calculate the distance sets of a number of pixel points to point A in the X, Y, and Z directions: G m -A = [ΔX m ,ΔY m ,ΔZ m T , where m ∈ [1, 12];
[0090] Select one of the pixel points adjacent to point A as a reference, and combine the horizontal normal vector S X and the vertical normal vector S Y to obtain an expression form of the normal vector S Z of point A in the Z - direction:
[0091]
[0092] Convert S Z into the form of spherical coordinates :
[0093]
[0094] Among them, α is the dip angle and β is the azimuth angle;
[0095] The dip angle α is expressed as:
[0096]
[0097] Among them, i represents the i-th pixel point in the disparity map, i = 1, 2, ……, n, and n is the total number of pixel points in the disparity map. S Zi represents the spherical coordinate in the Z direction; P i = S xi cosθ + S yi sinβ; S xi represents the spherical coordinate in the X direction; S yi represents the spherical coordinate in the Y direction;
[0098] The azimuth angle β is expressed as:
[0099]
[0100] Repeat the above steps for each pixel point in the disparity map to obtain the corresponding surface normal vector map.
[0101] Specifically, use a horizontal image gradient filter and a vertical image gradient filter to perform a convolution operation on each pixel point in the disparity map, so as to obtain the corresponding horizontal normal vector S X and vertical normal vector S Y . The horizontal image gradient filter and the vertical image gradient filter can be respectively expressed as:
[0102]
[0103] The parameter settings of the filter and the gradient of the corresponding matrix can be modified according to actual requirements. Only one example is shown here. Further, for any point A on the disparity map, take the twelve adjacent pixel points, and they can be represented as a set: N A = [G1, G2,......, G 12 T , so the distance sets in the X, Y, and Z directions between these twelve points and point A can be obtained: G m -A = [ΔX m , ΔY m , ΔZ m T , where m ∈ [1, 12]. Select one of the twelve pixel points adjacent to point A as a reference, and combine the horizontal normal vector S X and the vertical normal vector S Y , a normal vector S of point A in the Z direction can be obtained Z . The S obtained here Z is not necessarily the best normal vector representation of point A in the Z direction, and it may affect the accuracy of the final surface normal vector map. Therefore, it is treated as an energy minimization problem here, and it is quantified by the inclination angle α and the azimuth angle β. S Z is transformed into the form of spherical coordinates to obtain the best adjacent point selection.
[0104] The above operations are performed on each pixel point in the disparity map to obtain its corresponding surface normal vector map, which is used as the input of the encoder module of the drivable area segmentation model. The structure of the drivable area segmentation model is as Figure 3 shown.
[0105] S3. Build a drivable area segmentation model and train it to obtain a trained drivable area segmentation model. The drivable area segmentation model includes an encoder module, a multi-scale uncertainty perception module, and a decoder module connected in sequence. The encoder module is used to extract features from the surface normal vector map to obtain a first feature vector. The input multi-scale uncertainty perception module is used to reset the weights corresponding to each pixel point in the first feature vector to obtain a second feature vector. The decoder module is used to decode the second feature vector to obtain the segmentation result of the drivable area.
[0106] In a specific embodiment, the encoder module is a pyramid structure composed of a fully connected layer, a batch normalization layer, a ReLu activation function layer, a max pooling layer, and four encoder layers connected in sequence. The encoder layer uses ResNet-50 as the Backbone, and each layer corresponds to each layer of ResNet.
[0107] Specifically, as the decoder layer deepens, the number of channels of the feature vector gradually decreases, and the spatial size gradually increases. After passing through the Softmax activation layer, a segmentation result with the same size as the original input image is obtained.
[0108] In a specific embodiment, in the multi-scale uncertainty perception module, the first feature vector is convolved at three different scales to obtain three sets of convolution results. The three sets of convolution structures are respectively upsampled, and then through the Softplus activation function, feature maps at three scales are obtained: {[F0] l ,[F1] l},l∈1,2,3, where F1 is the drivable area and F0 is the non-drivable area;
[0109] An averaging operation is performed on the feature maps at the three scales to obtain the average value of the feature maps:
[0110]
[0111] Obtain the Dirichlet intensity map S based on F1 and F0:
[0112]
[0113] Calculate the credibility B0 of the drivable area, the credibility B1 of the drivable area, and the uncertainty parameter U based on the Dirichlet intensity map S, F1, and F0:
[0114]
[0115]
[0116]
[0117] Optimize the credibility B0 of the driving area, the credibility B1 of the drivable area, and the uncertainty parameter U according to the Dempster combination rule to obtain the optimized credibility b0, b1, and the optimized uncertainty parameter u:
[0118]
[0119]
[0120]
[0121] Calculate the new Dirichlet intensity map S' according to the optimized uncertainty parameter u:
[0122]
[0123] Calculate the non-drivable area attention parameter δ0 and the drivable area attention parameter δ1 according to the new Dirichlet intensity map S' and the optimized credibility b0, b1:
[0124]
[0125]
[0126] Among them, represents element-wise multiplication, that is, the multiplication operation is performed on each corresponding pixel point one by one;
[0127] Calculate the drivable area probability map P according to the drivable area attention parameter δ1 and the new Dirichlet intensity map:
[0128] The weight of each pixel point of the feature vector is redistributed with the drivable area probability map P as a reference basis to obtain a second feature vector.
[0129] In a specific embodiment, the first feature vector is convolved at three different scales, specifically including:
[0130] Convolution operations are respectively performed on the first feature vector using convolutional layers with kernel sizes of 1×1, 3×3, 3×3 and dilation rates of 0, 2, 5.
[0131] Specifically, referring to Figure 4 , the encoder module outputs a first feature vector and transmits it to the multi-scale uncertainty perception module to reset the weights corresponding to each pixel point of the first feature vector, enhancing the weights of the pixel points corresponding to the drivable area, thereby improving the final segmentation accuracy. The multi-scale uncertainty perception module convolves the feature vector at three different scales respectively, performs upsampling on the three sets of convolution results respectively to restore them to the size of the original image, and then uses the non-linear activation function Softplus to process the feature map, fixing the range of the feature values therein between (0, +∞) to ensure its non-negativity. This Softplus activation function can be expressed as: f(x) = ln(1 + e x ). An averaging operation is performed on the processed feature map to obtain the drivable area F1 and the non-drivable area F0. Using F1 and F0 to obtain the Dirichlet intensity map S, combining S, F1 and F0, calculating the credibility B0 of the drivable area, the credibility B1 of the drivable area and the uncertainty parameter U, and further optimizing the above relevant parameters to make the credibility of the drivable area and the non-drivable area be more accurately distinguished and strengthen the contrast of the segmentation edge area. According to the optimized uncertainty parameter u, a new Dirichlet intensity map can be calculated, and the non-drivable area attention parameter δ0 and the drivable area attention parameter δ1 are calculated. The attention degree indicates how much attention the model should pay to a certain area, and the larger the value, the higher the possibility that this area is the target area. The drivable area probability map P is calculated through the drivable area attention parameter δ1 and the Dirichlet intensity map, and the weight of each pixel point of the first feature vector is redistributed with this as a reference basis to obtain a second feature vector, so that the weight of the drivable area is increased and the attention degree of the model to this area is strengthened, thereby improving the segmentation accuracy.
[0132] In a specific embodiment, the decoder module includes five decoder layers and a Softmax activation layer connected in sequence. The decoder layer includes a deconvolution layer, a batch normalization layer, a skip connection layer and an activation function layer.
[0133] Specifically, the second feature vector is input into the decoder module, and the feature vector is remapped into a high-dimensional vector identical to the original data and used to reconstruct the input data, finally obtaining the segmentation result of the drivable area.
[0134] S4. Input the surface normal vector map into the trained drivable area segmentation model to obtain the segmentation result of the drivable area.
[0135] Specifically, deploy the trained drivable area segmentation model in the computer device of the vehicle, and then the surface normal vector map obtained by processing the collected binocular images in real time can be input into the trained drivable area segmentation model, and the segmentation result of the drivable area can be obtained through real-time analysis.
[0136] The above steps S1 - S4 do not represent the order between steps, but are only used as step symbols.
[0137] For further reference Figure 5 , as an implementation of the methods shown in the above figures, an embodiment of a drivable area detection device based on an enhanced disparity map and multi-scale uncertainty perception is provided in the present application. This device embodiment corresponds to Figure 2 the method embodiment shown, and this device can be specifically applied to various electronic devices.
[0138] An embodiment of the present application provides a drivable area detection device based on an enhanced disparity map and multi-scale uncertainty perception, which is characterized by including:
[0139] A disparity map acquisition module 1, configured to acquire binocular images of a road collected by a binocular camera, correct the binocular images to obtain a corrected image, and process the corrected image using a stereo matching algorithm to obtain a disparity map;
[0140] A feature enhancement module 2, configured to enhance the features of the disparity map using the surface normal vector method to obtain a surface normal vector map;
[0141] A model construction module 3, configured to construct and train a drivable area segmentation model to obtain a trained drivable area segmentation model. The drivable area segmentation model includes an encoder module, a multi-scale uncertainty perception module, and a decoder module connected in sequence. The encoder module is used to extract features from the surface normal vector map to obtain a first feature vector, and the multi-scale uncertainty perception module is used to reset the weights corresponding to each pixel point in the first feature vector to obtain a second feature vector. The decoder module is used to decode the second feature vector to obtain the segmentation result of the drivable area;
[0142] An execution module 4, configured to input the surface normal vector map into the trained drivable area segmentation model to obtain the segmentation result of the drivable area.
[0143] Reference is made below to Figure 6 , which shows a schematic structural diagram of a computer device 600 suitable for use in implementing the electronic device (such as Figure 1 the server or terminal device shown) of the embodiments of the present application. Figure 6 The electronic device shown is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.
[0144] As Figure 6 shown, the computer device 600 includes a central processing unit (CPU) 601 and a graphics processing unit (GPU) 602, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 603 or programs loaded from a storage section 609 into a random access memory (RAM) 604. In the RAM 604, various programs and data required for the operation of the device 600 are also stored. The CPU 601, GPU 602, ROM 603, and RAM 604 are connected to each other via a bus 605. An input / output (I / O) interface 606 is also connected to the bus 605.
[0145] The following components are connected to the I / O interface 606: an input section 607 including a keyboard, a mouse, etc.; an output section 608 including, for example, a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 609 including a hard disk, etc.; and a communication section 610 including a network interface card such as a LAN card, a modem, etc. The communication section 610 performs communication processing via a network such as the Internet. A drive 611 can also be connected to the I / O interface 606 as needed. A removable medium 612, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 611 as needed so that a computer program read from it can be installed into the storage section 609 as needed.
[0146] Specifically, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 610, and / or installed from the removable medium 612. When the computer program is executed by the central processing unit (CPU) 601 and the graphics processing unit (GPU) 602, the above functions defined in the methods of the present application are executed.
[0147] It should be noted that the computer-readable medium described in this application can be a computer-readable signal medium, a computer-readable medium, or any combination of the two. The computer-readable medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor device, apparatus, or component, or any combination of the above. More specific examples of the computer-readable medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, the computer-readable medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution device, apparatus, or component. And in this application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution device, apparatus, or component. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination of the above.
[0148] The computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages - such as Java, Smalltalk, C++, and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or, alternatively, can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0149] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of apparatuses, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based apparatus that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0150] The modules described in the embodiments of the present application can be implemented in software or in hardware. The described modules can also be provided in a processor.
[0151] As another aspect, the present application also provides a computer-readable medium, which can be included in the electronic device described in the above embodiments; or can exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to: obtain a binocular image of a road collected by a binocular camera, correct the binocular image to obtain a corrected image, process the corrected image using a stereo matching algorithm to obtain a disparity map; enhance the features of the disparity map using a surface normal vector method to obtain a surface normal vector map; construct and train a drivable area segmentation model to obtain a trained drivable area segmentation model, where the drivable area segmentation model includes an encoder module, a multi-scale uncertainty perception module, and a decoder module connected in sequence, the encoder module is used to extract features from the surface normal vector map to obtain a first feature vector, the input multi-scale uncertainty perception module is used to reset the weights corresponding to each pixel point in the first feature vector to obtain a second feature vector, and the decoder module is used to decode the second feature vector to obtain a segmentation result of the drivable area; input the surface normal vector map into the trained drivable area segmentation model to obtain a segmentation result of the drivable area.
[0152] The above description is only a preferred embodiment of the present application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with (but not limited to) the technical features with similar functions disclosed in the present application.
Claims
1. A drivable area detection method based on enhanced disparity maps and multi-scale uncertainty perception, characterized in that, It includes the following steps: Obtain the binocular images of the road collected by the binocular camera, correct the binocular images to obtain the corrected images, and process the corrected images using a stereo matching algorithm to obtain a disparity map; Enhance the features of the disparity map using the surface normal vector method to obtain a surface normal vector map; Construct a drivable area segmentation model and train it to obtain a trained drivable area segmentation model. The drivable area segmentation model includes an encoder module, a multi-scale uncertainty perception module, and a decoder module connected in sequence. The encoder module is used to extract features from the surface normal vector map to obtain a first feature vector. The input to the multi-scale uncertainty perception module is used to reset the weights corresponding to each pixel point in the first feature vector to obtain a second feature vector. In the multi-scale uncertainty perception module, the first feature vector is convolved at three different scales to obtain three sets of convolution results. The three sets of convolution results are respectively upsampled and then passed through the Softplus activation function to obtain feature maps at three scales: {[F0] l ,[F1] l}, where l ∈ 1, 2, 3. Among them, F1 is the drivable area and F0 is the non-drivable area; Perform an averaging operation on the feature maps at three scales to obtain the average value of the feature maps: Obtain the Dirichlet intensity map S according to F1 and F0; Calculate the non-drivable area credibility B0, the drivable area credibility B1, and the uncertainty parameter U according to the Dirichlet intensity map S, F1, and F0; Optimize the non-drivable area credibility B0, the drivable area credibility B1, and the uncertainty parameter U according to the Dempster combination rule to obtain the optimized credibility b0, b1, and the optimized uncertainty parameter u; Calculate a new Dirichlet intensity map S' according to the optimized uncertainty parameter u; Calculate the non-drivable area attention parameter δ0 and the drivable area attention parameter δ1 according to the new Dirichlet intensity map S' and the optimized credibility b0, b1; Among them, represents element-wise multiplication, that is, the multiplication operation is performed pixel by pixel for each corresponding pixel; The drivable area probability map P is calculated based on the drivable area attention parameter δ1 and the new Dirichlet intensity map: Reassign the weight of each pixel point of the feature vector with the drivable area probability map P as the reference basis to obtain a second feature vector, and the decoder module is used to decode the second feature vector to obtain the segmentation result of the drivable area; Input the surface normal vector map into the trained drivable area segmentation model to obtain the segmentation result of the drivable area.
2. The drivable area detection method based on enhanced parallax map and multi-scale uncertainty perception according to claim 1, characterized in that The correction includes distortion correction and stereo correction.
3. The method for detecting drivable regions based on enhanced disparity maps and multi-scale uncertainty perception according to claim 1, wherein The enhancing the features of the disparity map using the surface normal vector method to obtain a surface normal vector map specifically includes: Perform a convolution operation on each pixel point in the disparity map using a horizontal image gradient filter and a vertical image gradient filter to obtain the corresponding horizontal normal vector and vertical normal vector: Where D represents the disparity map, and fx and fy respectively represent the camera focal lengths in the x direction and y direction; For any point A on the disparity map, take a number of adjacent pixel points and represent them as a set: N A =[G1, G2,......, G m T , and calculate the distance sets of a number of pixel points to point A in the X, Y, and Z directions: G m -A = [ΔX m , ΔY m , ΔZ m T , where m ∈ [1, 12]; Select one of the pixel points adjacent to point A as a reference, and combine it with the horizontal normal vector S X and the vertical normal vector S Y , to obtain an expression form of the normal vector S Z in the Z direction at point A: Convert S Z to spherical coordinates in the form of: Where α is the inclination angle and β is the azimuth angle; The inclination angle α is expressed as: Among them, i represents the i-th pixel point in the disparity map, where i = 1, 2, ……, n, and n is the total number of pixel points in the disparity map, S Zi represents the corresponding spherical coordinate in the Z direction; P i = S xi cosβ + S yi sinβ; S xi represents the corresponding spherical coordinate in the X direction; S yi represents the corresponding spherical coordinate in the Y direction; The azimuth angle β is expressed as: Repeat the above steps for each pixel point in the disparity map to obtain the corresponding surface normal vector map.
4. The drivable area detection method based on enhanced disparity map and multi-scale uncertainty perception according to claim 1, characterized in that, The encoder module is a pyramid structure composed of a fully connected layer, a batch normalization layer, a ReLu activation function layer, a max pooling layer, and four encoder layers connected in sequence. The encoder layer uses ResNet-50 as the Backbone, and each layer corresponds to the corresponding layer of ResNet.
5. The method for detecting drivable areas based on enhanced disparity maps and multi-scale uncertainty perception according to claim 1, wherein Perform convolutions on the first feature vector at three different scales respectively, specifically including: Perform convolution operations on the first feature vector using convolutional layers with kernel sizes of 1×1, 3×3, 3×3 and dilation rates of 0, 2, 5 respectively.
6. The drivable area detection method based on enhanced disparity map and multi-scale uncertainty perception according to claim 1, wherein The decoder module includes five decoder layers and a Softmax activation layer connected in sequence. The decoder layer includes a transposed convolution layer, a batch normalization layer, a skip connection layer, and an activation function layer.
7. An apparatus for detecting drivable areas based on an enhanced disparity map and multi-scale uncertainty perception, characterized in that, It includes: A parallax map acquisition module, configured to acquire binocular images of a road collected by a binocular camera, correct the binocular images to obtain corrected images, and process the corrected images using a stereo matching algorithm to obtain a parallax map; A feature enhancement module, configured to enhance the features of the parallax map using the surface normal vector method to obtain a surface normal vector map; The model construction module is configured to construct and train a drivable area segmentation model to obtain a trained drivable area segmentation model. The drivable area segmentation model includes an encoder module, a multi-scale uncertainty perception module, and a decoder module connected in sequence. The encoder module is used to extract features from the surface normal vector map to obtain a first feature vector, and the input multi-scale uncertainty perception module is used to reset the weights corresponding to each pixel point in the first feature vector to obtain a second feature vector. In the multi-scale uncertainty perception module, the first feature vector is convolved at three different scales to obtain three sets of convolution results, the three sets of convolution results are respectively upsampled, and then through the Softplus activation function, feature maps at three scales are obtained: {[F0] l ,[F1] l}, l ∈ 1, 2, 3, where F1 is the drivable area and F0 is the non-drivable area; Perform an averaging operation on the feature maps at three scales to obtain the average value of the feature maps: Obtain the Dirichlet intensity map S according to F1 and F0: Calculate the non-drivable area credibility B0, the drivable area credibility B1, and the uncertainty parameter U according to the Dirichlet intensity map S, F1, and F0; Optimize the non-drivable area credibility B0, the drivable area credibility B1, and the uncertainty parameter U according to the Dempster combination rule to obtain the optimized credibility b0, b1, and the optimized uncertainty parameter u; Calculate a new Dirichlet intensity map S' according to the optimized uncertainty parameter u; Calculate the non-drivable area attention parameter δ0 and the drivable area attention parameter δ1 according to the new Dirichlet intensity map S' and the optimized credibility b0, b1; Among them, represents element-wise multiplication, that is, the multiplication operation is performed pixel by pixel for each corresponding pixel. Calculate the drivable area probability map P according to the drivable area attention parameter δ1 and the new Dirichlet intensity map; Reassign the weight of each pixel point of the feature vector based on the drivable area probability map P as a reference to obtain a second feature vector, and the decoder module is used to decode the second feature vector to obtain the segmentation result of the drivable area; An execution module, configured to input the surface normal vector map into the trained drivable area segmentation model to obtain the segmentation result of the drivable area.
8. An electronic device, comprising: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method according to any one of claims 1-6.
Citation Information
Patent Citations
3D target detection algorithm based on camera and laser radar data fusion
CN113985445A
Underwater target three-dimensional reconstruction method based on deep learning
CN115147709A