Road surface contour estimation method and system of vehicle-mounted binocular system based on deep learning

By applying a deep learning-based pavement profile estimation method in the vehicle-mounted binocular system, the deep learning stereo matching model of the gantry loop unit and the real-time pitch angle of the inertial navigation unit is used to solve the problem of sparseness and long time-consuming binocular visual parallax map, and high-precision, continuous and real-time pavement profile estimation is achieved, which is suitable for vehicle suspension pre-image control.

CN119992496APending Publication Date: 2025-05-13NANJING TECH UNIV

Patent Information

Application Number
CN202510069295.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the prior art, binocular vision has sparse parallax maps, pavement profile estimates are too smooth and sparse, and binocular vision based on deep learning is too time-consuming and cannot meet the real-time requirements.

Method used

A pavement profile estimation method based on deep learning is proposed for vehicle-mounted binocular systems. By acquiring the original image, extracting feature maps and post-processing, the pixel coordinates and confidence information of road surface obstacles are obtained, and the local image of the obstacle area is obtained, and the deep learning stereo matching model based on the gate-type cycle unit is input to generate a local disparity map, and the real-time pitch angle provided by the inertial navigation unit is fused to achieve millimeter-level pavement profile estimation.

Benefits of technology

While maintaining high inference speed, it achieves high accuracy and continuity of road profile, which is suitable for vehicle suspension pre-aim control, improving the effectiveness of vehicle suspension control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992496A_ABST
    Figure CN119992496A_ABST
Patent Text Reader

Abstract

The invention discloses a road contour estimation method and system of a vehicle-mounted binocular system based on deep learning. The method comprises the steps of obtaining an original image, obtaining a feature map through the original image, performing post-processing on the feature map, and obtaining road obstacle pixel coordinates and confidence information; processing the pixel coordinates and the confidence information of the road surface obstacle to obtain a local image of an obstacle area of the original image; inputting the local image of the barrier region of the original image into a deep learning stereo matching model based on a gate-type cycle unit for processing to obtain a local disparity map corresponding to the barrier region; millimeter-level pavement contour estimation is realized by fusing a real-time pitch angle provided by an inertial navigation unit through a local disparity map of a corresponding obstacle area. According to the method, important features of a road surface obstacle area can be concerned, irrelevant features are inhibited, and high-precision, high-efficiency, high-continuity and high-robustness road surface contour estimation is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of vehicle suspension control, and in particular to a road profile estimation method and system for a vehicle-mounted binocular system based on deep learning. Background Art

[0002] Vehicle suspension control is an intelligent control technology that combines sensors, control algorithms and actuators to achieve real-time monitoring and dynamic adjustment of the vehicle suspension system to optimize the performance, comfort and safety of the vehicle. Vehicle control algorithms can be divided into feedback control and feedforward control according to the control strategy, and are widely used in vehicle assisted driving and autonomous driving systems.

[0003] However, since feedback control detects the real-time state of the vehicle (such as body acceleration, wheel displacement, etc.) and adjusts the parameters of the suspension system based on this information, it has problems such as information lag, high computational complexity, and poor robustness. In order to better meet actual engineering needs, researchers focus on predicting the future state of the vehicle in advance to achieve better control effects. At present, researchers are studying feedforward control strategies based on high-precision sensors.

[0004] Feedforward control is a control strategy that uses high-precision sensors (such as lidar, binocular cameras, monocular cameras, etc.) to predict the future state of the vehicle (such as road roughness, vehicle speed, etc.) and adjust the suspension system parameters in advance. However, since lidar is easily affected by environmental factors, there is a lot of noise in the point cloud data. At the same time, the sparse radar point cloud data limits the continuous road profile estimation, reducing the effectiveness of vehicle suspension control; secondly, binocular traditional vision is easily affected by natural factors such as lighting, resulting in a large number of noise points in the disparity map, which limits the estimation of the road profile. The stereo matching algorithm based on deep learning uses a large number of data sets for pre-training, is highly robust to the environment, and can obtain dense disparity maps. As a road profile estimation algorithm, it can obtain dense and continuous road elevation curves.

[0005] However, the existing deep learning-based stereo matching algorithm is too time-consuming, that is, the vehicle preview control only needs to focus on the road obstacle area, and the remaining flat areas can be satisfied through the vehicle's own passive control, which limits its application in engineering. Summary of the invention

[0006] In order to overcome the problems of sparse disparity maps in traditional binocular vision, overly smooth and sparse road contour estimation, and excessive time-consuming and inability to meet real-time requirements of binocular vision based on deep learning, the present invention proposes a road contour estimation method and system for a vehicle-mounted binocular system based on deep learning, which achieves binocular vision while maintaining a high reasoning speed while still ensuring good accuracy and continuity of the road contour.

[0007] On the one hand, to achieve the above-mentioned purpose, the present invention provides a road profile estimation method of a vehicle-mounted binocular system based on deep learning, comprising:

[0008] Acquire an original image, obtain a feature map through the original image, and post-process the feature map to obtain pixel coordinates and confidence information of road obstacles, wherein the original image is a corrected road image;

[0009] Processing the pixel coordinates and confidence information of the road obstacle to obtain a local image of the obstacle area of ​​the original image;

[0010] Inputting the local image of the obstacle area of ​​the original image into a deep learning stereo matching model based on a gated recurrent unit for processing to obtain a local disparity map corresponding to the obstacle area;

[0011] The local disparity map of the corresponding obstacle area is integrated with the real-time pitch angle provided by the inertial navigation unit to achieve millimeter-level road profile estimation.

[0012] Preferably, obtaining the feature map comprises:

[0013] After padding and scaling the original image, the image is input into the convolutional neural network Yolov5s for processing, and the feature map is output;

[0014] The convolutional neural network Yolov5s includes thirteen convolutional units, each of which is connected to a batch normalization layer and an activation function.

[0015] Preferably, the thirteen convolution units include:

[0016] The first convolution unit includes the convolution layer Conv1_1, with a convolution kernel size of 6, a stride and padding of 2, and an output channel of 32;

[0017] The second convolution unit includes convolution layer Conv2_1 and C3 module 2, with a convolution kernel size of 3, a stride of 2, a padding of 1, and an output channel of 64;

[0018] The third convolution unit includes convolution layer Conv3_1 and C3 module 3, with a convolution kernel size of 3, a stride of 2, a padding of 1, and an output channel of 128;

[0019] The fourth convolution unit includes convolution layer Conv4_1 and C3 module 4, with a convolution kernel size of 3, a stride of 2, a padding of 1, and an output channel of 256;

[0020] The fifth convolution unit includes convolution layer Conv5_1 and C3 module five, with a convolution kernel size of 3, a stride of 2, a padding of 1, and an output channel of 512;

[0021] The sixth convolution unit includes an SPPF module with an output channel of 512;

[0022] The seventh convolution unit includes the convolution layer Conv7_1, the upsampling layer seven, the connection layer Concat1 and the C3 module seven. The convolution kernel size is 1, the step size is 1, the padding is 0, and the output channel is 256;

[0023] The eighth convolution unit includes a convolution layer Conv8_1, an upsampling layer eight, a connection layer Concat2 and a C3 module eight, the convolution kernel size is 1, the step size is 1, the padding is 0, and the output channel is 128;

[0024] The ninth convolution unit includes the convolution layer Conv9_1, the connection layer Concat3 and the C3 module nine, with a convolution kernel size of 3, a stride of 2, a padding of 1, and an output channel of 128;

[0025] The tenth convolution unit includes the convolution layer Conv10_1, the connection layer Concat4 and the C3 module ten, with a convolution kernel size of 3, a stride of 2, a padding of 1, and an output channel of 256;

[0026] The eleventh convolution unit includes a convolution layer Conv11_1, with a convolution kernel size of 1, a stride of 1, a padding of 0, and an output channel of (5+n cls )×3, where n cls To detect the type;

[0027] The twelfth convolution unit includes the convolution layer Conv12_1, the convolution kernel size is 1, the stride is 1, the padding is 0, and the output channel is (5+n cls )×3;

[0028] The thirteenth convolution unit includes a convolution layer Conv13_1, a convolution kernel size of 1, a stride of 1, a padding of 0, and an output channel of (5+n cls )×3.

[0029] Preferably, each C3 module in the convolution unit consists of a convolution layer Conv1, a convolution layer Conv2, a bottleneck layer Bottleneck and a convolution layer Conv3, and each convolution layer is connected to a batch normalization layer and an activation function layer;

[0030] Among them, the convolution kernel size used by the convolution layer Conv1 and the convolution layer Conv2 is 1, and the output channels are half of the output channels of the convolution layer Conv1_1, the convolution layer Conv2_1, the convolution layer Conv3_1, the convolution layer Conv4_1 and the convolution layer Conv5_1 respectively; the convolution kernel size used by the convolution layer Conv3 is 1, and the output channel is the sum of the output channels of the convolution layer Conv1 and the convolution layer Conv2.

[0031] Preferably, the convolution layer in the bottleneck layer Bottleneck consists of a convolution layer Conv4 and a convolution layer Conv5, wherein the convolution kernel sizes used by the convolution layer Conv4 and the convolution layer Conv5 are 1 and 3 respectively, and the output channels of the convolution layer Conv4 and the convolution layer Conv5 are the same as the output channels of the convolution layer Conv2.

[0032] Preferably, the SPPF module in the sixth convolution unit is composed of a convolution layer ConvP_1, a maximum pooling layer Maxpool1, a maximum pooling layer Maxpool2, a maximum pooling layer Maxpool3 and a convolution layer ConvP_2, and each convolution layer is connected to a batch normalization layer and an activation function layer;

[0033] Among them, the convolution kernel size used by the convolution layer ConvP_1 is 1, and the output channel is 256; the convolution kernel size used by the maximum pooling layer Maxpool1, the maximum pooling layer Maxpool2, and the maximum pooling layer Maxpool3 is 5, the step size is 1, and the padding is 2; the convolution kernel size used by the convolution layer ConvP_2 is 1, the step size is 1, the padding is 0, and the output channel is 512.

[0034] Preferably, obtaining a local image of an obstacle area in the original image includes:

[0035] The pixel coordinates of the road obstacle, the confidence information, and the vehicle driving trajectory are subjected to cropping dynamic stereo matching CDSM to obtain a local image of the obstacle area of ​​the original image.

[0036] Preferably, obtaining the local disparity map corresponding to the obstacle area includes:

[0037] Filling the edge of the local image of the obstacle area of ​​the original image based on a preset condition, processing the filled local image through the deep learning stereo matching model based on the gated recurrent unit, and obtaining a local disparity map of the corresponding obstacle area;

[0038] The deep learning stereo matching model based on gated recurrent units includes a feature map extraction module, a parameterized full-pair correlation pyramid construction module, an iterative convolution module based on gated recurrent units, and an uncertainty perception refinement module;

[0039] The feature map extraction module is used to extract feature maps of different scales;

[0040] The parameterized all-pairs correlation pyramid construction module is used to construct a parameterized cost space using multi-Gaussian distribution theory;

[0041] The iterative convolution module based on the gated recurrent unit is used to process the candidate hidden states by resetting the gates, updating the gates, and obtaining the hidden state output;

[0042] The uncertainty-aware refinement module is used to estimate the uncertainty map and guide the fusion of the residual map and the disparity map.

[0043] Preferably, achieving the millimeter-level road profile estimation includes:

[0044] For each pixel point (u i ,v i ) corresponds to the disparity value d i Perform stereo calibration to obtain the focal length f,f of the binocular camera x ,f y , the baseline length b, and the principal point of the image (u 0 ,v 0 );

[0045] Calculate the coordinates of each pixel in the camera coordinate system through triangulation And according to the attitude position of the camera and the real-time relative pitch angle of the vehicle measured by the inertial navigation unit, the spatial coordinates in the wheel coordinate system are calculated.

[0046] On the other hand, to achieve the above-mentioned purpose, the present invention also provides a road profile estimation system of a vehicle-mounted binocular system based on deep learning, comprising:

[0047] A first image processing module, an obstacle area local image acquisition module, a second image processing module and a millimeter-level road surface profile calculation module;

[0048] The first image processing module is used to obtain an original image, obtain a feature map through the original image, and post-process the feature map to obtain pixel coordinates and confidence information of road obstacles;

[0049] The obstacle area local image acquisition module is used to process the pixel coordinates and confidence information of the road obstacle to obtain the local image of the obstacle area of ​​the original image;

[0050] The second image processing module is used to input the local image of the obstacle area of ​​the original image into the deep learning stereo matching model based on the gated recurrent unit for processing, so as to obtain a local disparity map corresponding to the obstacle area;

[0051] The millimeter-level road surface profile calculation module is used to achieve millimeter-level road surface profile estimation by fusing the local disparity map of the corresponding obstacle area with the real-time pitch angle provided by the inertial navigation unit.

[0052] Compared with the prior art, the present invention has the following advantages and technical effects:

[0053] The present invention inputs the original image into the classic convolutional neural network Yolov5s, and obtains a feature map after passing through a convolution layer, a pooling layer, a normalization layer and an activation function; obtains the pixel coordinates and confidence of the obstacle in the image by using a feature map post-processing method; uses CDSM to determine whether it is a key frame, and selectively crops the original image to obtain a local image of the obstacle area; inputs it into a deep learning stereo matching model based on a gated recurrent unit (GRU), and obtains a feature map after passing through a series of convolution layers, pooling layers, batch normalization layers, gated recurrent layers, upsampling layers and activation functions; uses an uncertain perception module to restore the disparity map of the original size; uses the real-time pitch angle information and disparity map information provided by the inertial navigation unit to calculate the road surface profile information in the wheel coordinate system. The present invention can focus on the important features of the road obstacle area, suppress irrelevant features, and realize high-precision, high-efficiency, high-continuity, and high-robustness road surface profile estimation, which is suitable for vehicle suspension preview control. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0055] Figure 1 This is a flow chart of a road profile estimation method of a vehicle-mounted binocular system based on deep learning according to an embodiment of the present invention;

[0056] Figure 2 Schematic diagram of a cropped dynamic stereo matching structure according to an embodiment of the present invention. DETAILED DESCRIPTION

[0057] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0058] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0059] The present invention proposes a road profile estimation method for a vehicle-mounted binocular system based on deep learning. Figure 1 ,include:

[0060] Acquire an original image, obtain a feature map through the original image, and post-process the feature map to obtain pixel coordinates and confidence information of road obstacles;

[0061] Processing the pixel coordinates and confidence information of the road obstacle to obtain a local image of the obstacle area of ​​the original image;

[0062] Inputting the local image of the obstacle area of ​​the original image into a deep learning stereo matching model based on a gated recurrent unit for processing to obtain a local disparity map corresponding to the obstacle area;

[0063] The local disparity map of the corresponding obstacle area is integrated with the real-time pitch angle provided by the inertial navigation unit to achieve millimeter-level road profile estimation.

[0064] Specifically, this embodiment inputs the original image into the classic convolutional neural network Yolov5s, and obtains a feature map after passing through a convolution layer, a pooling layer, a normalization layer and an activation function; obtains the pixel coordinates and confidence of the obstacles in the image by using a feature map post-processing method; uses CDSM to determine whether it is a key frame, and selectively crops the original image to obtain a local image of the obstacle area; inputs it into a deep learning stereo matching model based on a gated recurrent unit (GRU), and obtains a feature map after passing through a series of convolution layers, pooling layers, batch normalization layers, gated recurrent layers, upsampling layers and activation functions; uses an uncertain perception module to restore a disparity map of the original size; and uses the real-time pitch angle information and disparity map information provided by an inertial navigation unit to calculate the road surface contour information in the wheel coordinate system.

[0065] Furthermore, a feature map is obtained, including:

[0066] After padding and scaling the original image, the image is input into the convolutional neural network Yolov5s for processing, and the feature map is output;

[0067] The convolutional neural network Yolov5s includes thirteen convolutional units, each of which is connected to a batch normalization layer and an activation function.

[0068] Specifically, the thirteen convolution units include:

[0069] The first convolution unit includes the convolution layer Conv1_1, with a convolution kernel size of 6, a stride and padding of 2, and an output channel of 32;

[0070] The second convolution unit includes convolution layer Conv2_1 and C3 module 2, with a convolution kernel size of 3, a stride of 2, a padding of 1, and an output channel of 64;

[0071] The third convolution unit includes convolution layer Conv3_1 and C3 module 3, with a convolution kernel size of 3, a stride of 2, a padding of 1, and an output channel of 128;

[0072] The fourth convolution unit includes convolution layer Conv4_1 and C3 module 4, with a convolution kernel size of 3, a stride of 2, a padding of 1, and an output channel of 256;

[0073] The fifth convolution unit includes convolution layer Conv5_1 and C3 module five, with a convolution kernel size of 3, a stride of 2, a padding of 1, and an output channel of 512;

[0074] The sixth convolution unit includes an SPPF module with an output channel of 512;

[0075] The seventh convolution unit includes the convolution layer Conv7_1, the upsampling layer seven, the connection layer Concat1 and the C3 module seven. The convolution kernel size is 1, the step size is 1, the padding is 0, and the output channel is 256;

[0076] The eighth convolution unit includes a convolution layer Conv8_1, an upsampling layer eight, a connection layer Concat2 and a C3 module eight, the convolution kernel size is 1, the step size is 1, the padding is 0, and the output channel is 128;

[0077] The ninth convolution unit includes the convolution layer Conv9_1, the connection layer Concat3 and the C3 module nine, with a convolution kernel size of 3, a stride of 2, a padding of 1, and an output channel of 128;

[0078] The tenth convolution unit includes the convolution layer Conv10_1, the connection layer Concat4 and the C3 module ten, with a convolution kernel size of 3, a stride of 2, a padding of 1, and an output channel of 256;

[0079] The eleventh convolution unit includes a convolution layer Conv11_1, with a convolution kernel size of 1, a stride of 1, a padding of 0, and an output channel of (5+n cls )×3, where n cls To detect the type;

[0080] The twelfth convolution unit includes the convolution layer Conv12_1, the convolution kernel size is 1, the stride is 1, the padding is 0, and the output channel is (5+n cls )×3;

[0081] The thirteenth convolution unit includes a convolution layer Conv13_1, a convolution kernel size of 1, a stride of 1, a padding of 0, and an output channel of (5+n cls )×3.

[0082] Furthermore, each C3 module in the convolution unit consists of a convolution layer Conv1, a convolution layer Conv2, a bottleneck layer Bottleneck and a convolution layer Conv3, and each convolution layer is connected to a batch normalization layer and an activation function layer.

[0083] Among them, the convolution kernel size used by the convolution layer Conv1 and the convolution layer Conv2 is 1, and the output channels are half of the output channels of the convolution layer Conv1_1, the convolution layer Conv2_1, the convolution layer Conv3_1, the convolution layer Conv4_1 and the convolution layer Conv5_1 respectively; the convolution kernel size used by the convolution layer Conv3 is 1, and the output channel is the sum of the output channels of the convolution layer Conv1 and the convolution layer Conv2.

[0084] Specifically, the feature map passes through the convolution layer Conv1 and the convolution layer Conv2 respectively, and Conv2 is followed by the bottleneck layer Bottleneck, and then the feature map is spliced, and finally the convolution layer Conv3; the convolution layer Conv1 and the convolution layer Conv2 in the C3 module use the convolution kernel size of 1, and their output channels are half of the output channels of the preceding convolution layers Conv1_1, Conv2_1, Conv3_1, Conv4_1 and Conv5_1; the convolution layer Conv3 in the C3 module uses the convolution kernel size of 1, and the output channel is the sum of the output channels of the convolution layers Conv1 and Conv2. The convolution layer in Bottleneck consists of the convolution layer Conv4 and the convolution layer Conv5, where the convolution layers Conv4 and Conv5 in Bottleneck use the convolution kernel size of 1 and 3 respectively, and the output channels are the same as the output channels of the convolution layer Conv2.

[0085] Furthermore, the SPPF module in the sixth convolutional unit is composed of a convolutional layer ConvP_1, a maximum pooling layer Maxpool1, a maximum pooling layer Maxpool2, a maximum pooling layer Maxpool3 and a convolutional layer ConvP_2, and each convolutional layer is connected to a batch normalization layer and an activation function layer;

[0086] Among them, the convolution kernel size used by the convolution layer ConvP_1 is 1, and the output channel is 256; the convolution kernel size used by the maximum pooling layer Maxpool1, the maximum pooling layer Maxpool2, and the maximum pooling layer Maxpool3 is 5, the step size is 1, and the padding is 2; the convolution kernel size used by the convolution layer ConvP_2 is 1, the step size is 1, the padding is 0, and the output channel is 512.

[0087] Specifically, the SPPF module consists of a convolution layer ConvP_1, a maximum pooling layer Maxpool1, a maximum pooling layer Maxpool2, a maximum pooling layer Maxpool3 and a convolution layer ConvP_2, and each convolution layer is followed by a batch normalization layer and an activation function layer; the specific process is that the feature map passes through the convolution layer ConvP_1 and the maximum pooling layer Maxpool1, the maximum pooling layer Maxpool2, and the maximum pooling layer Maxpool3 in sequence to form a multi-scale feature map, and finally the multi-scale feature map is Concat-operated and then passes through the convolution layer ConvP_2; the convolution layer ConvP_1 in the SPPF module uses a convolution kernel size of 1 and an output channel of 256; the maximum pooling layer Maxpool1, the maximum pooling layer Maxpool2, and the maximum pooling layer Maxpool3 in the SPPF module use a convolution kernel size of 5, a step size of 1, and a padding of 2; the convolution layer ConvP_2 uses a convolution kernel size of 1, a step size of 1, a padding of 0, and an output channel of 512.

[0088] The activation function uses the SiLU function f(x)=x·sigmoid(x), which means the output is between 0 and 1.

[0089] Furthermore, a local image of the obstacle area of ​​the original image is obtained, including:

[0090] The pixel coordinates of road obstacles, confidence information, and vehicle driving trajectory are clipped and dynamically stereo matched by CDSM to obtain the local image of the obstacle area in the original local image.

[0091] Specifically, performing cropping dynamic stereo matching CDSM includes:

[0092] First, the confidence of the obstacle is judged. If the confidence is greater than the threshold set by the system, the pixel coordinates of the obstacle are judged. If the wheel trajectory passes through the obstacle, the pixel coordinates of the obstacle are compared with the pre-set clipping threshold to clip out the obstacle area of ​​the appropriate height.

[0093] Furthermore, obtaining a local disparity map of the corresponding obstacle area includes:

[0094] Filling the edge of the local image of the obstacle area of ​​the original image based on a preset condition, processing the filled local image through the deep learning stereo matching model based on the gated recurrent unit, and obtaining a local disparity map of the corresponding obstacle area;

[0095] Among them, the deep learning stereo matching model based on gated recurrent units includes a feature map extraction module, a parameterized full-pair correlation pyramid construction module, an iterative convolution module based on gated recurrent units and an uncertainty perception refinement module.

[0096] Specifically, the feature map extraction module uses a series of residual block structures to extract feature maps of different scales to obtain feature vectors with sizes of 1 / 4, 1 / 8, and 1 / 16 of the original image.

[0097] The parameterized all-pair correlation pyramid construction module uses the multi-Gaussian distribution theory and the following formula to construct a parameterized cost space:

[0098]

[0099] In the formula, Represents the parameter settings of the multivariate Gaussian distribution, including weights Mean and standard deviation M represents the number of Gaussian distributions. ~ means sampling from one of the distributions, α i Received constraints.

[0100] The iterative convolution module based on the gated recurrent unit consists of multiple convolutional layers to form a network containing a reset gate, an update gate and candidate hidden states. Assume that the input sequence is x t , the hidden state is h t , then for the reset gate, we can get:

[0101] r t =σ(W r ·[h t-1 ,x t ]+b r );

[0102] In the formula, r t is the output of the reset gate, W r is the weight matrix, h t-1 is the hidden state at the previous moment, x t is the input at the current moment, b r is the bias term, and σ is the sigmoid function.

[0103] The output of the update gate can be calculated as follows:

[0104] z t =σ(W z ·[h t-1 ,x t ]+b z );

[0105] In the formula, z t is the output of the update gate, W z is the weight matrix, h t-1 is the hidden state at the previous moment, x t is the input at the current moment, bz is the bias term, and σ is the sigmoid function.

[0106] The candidate hidden state can be calculated as follows:

[0107]

[0108] in, is a candidate hidden state, W h is the weight matrix, r t ⊙h t-1 Represents the element-wise multiplication of the reset gate output and the hidden state at the previous moment, x t is the input at the current moment, b h is the bias term, and tanh is the hyperbolic tangent function.

[0109] Finally, we can get the hidden state output at the current moment:

[0110]

[0111] In the formula, h t is the hidden state at the current moment, z t ⊙h t-1 represents the element-wise multiplication of the update gate output and the hidden state at the previous moment, Represents the element-wise multiplication of the update gate output and the candidate hidden state.

[0112] It is worth noting that all weights and biases in the above formulas are for one convolutional layer.

[0113] The uncertainty-aware refinement module consists of a series of convolutional layers. First, the mean μ, variance σ, and weight α are input into a series of convolutional blocks, where the last convolutional block contains a sigmoid activation function to estimate the uncertainty map.

[0114] Then the uncertainty map, disparity map and left feature map are concatenated and the residual map is predicted through a series of convolutional layers. Except for the last layer, each layer uses the Leaky-ReLU function.

[0115] Finally, the uncertainty map is used to guide the fusion of the residual map and the disparity map. The specific calculation method is as follows:

[0116]

[0117] In the formula, is the disparity map after fusion, is the disparity map, R is the residual map, and U is the uncertainty map.

[0118] Furthermore, the method for obtaining a disparity map having the same size as the original local image is as follows:

[0119] Fill the edges of the local image of the obstacle area to ensure that the length and width are multiples of 32;

[0120] The filled local image is passed through a deep learning stereo matching model based on a gated recurrent unit to obtain a local disparity map of the obstacle area.

[0121] The deep learning stereo matching convolutional neural network based on gated recurrent units in this embodiment includes multiple layers of convolutional layers, and each layer of convolutional layers is followed by a normalization layer and an activation function. Convolution is divided into multiple stages. The first convolution stage is a feature extractor composed of multiple layers of convolutional layers, and the feature extractor includes multiple layers of convolutional layers and a Bottleneck layer. The normalization layer used in the feature extraction layer stage is a batch normalization layer; the second convolution stage is a parameterized cost space builder composed of multiple layers of convolutional layers and pooling layers. After parameterizing and decomposing the 1 / 4 dimensional feature map to form a parameterized cost space, a series of maximum pooling layers are used to obtain multi-level features to form a multi-layer full-pair correlation pyramid; the third convolution stage is a convolution layer containing multiple layers of gated recurrent units, and the semantic information of multiple layers of images is obtained by fusing the correlation pyramid and the feature extractor, and a coarse disparity map of 1 / 4 scale is iteratively updated from 0; the fourth convolution stage is to use an uncertain refinement perception module to upsample the coarse disparity map to obtain a fine disparity map of the original image size.

[0122] Furthermore, millimeter-level road profile estimation is achieved, including:

[0123] Each pixel point (u i ,v i ) corresponds to the disparity value d i Perform stereo calibration to obtain the focal length f,f of the binocular camera x ,f y , the baseline length b, and the principal point of the image (u 0 ,v 0 );

[0124] Calculate the coordinates of each pixel in the camera coordinate system through triangulation And according to the attitude position of the camera and the real-time relative pitch angle of the vehicle measured by the inertial navigation unit, the spatial coordinates in the wheel coordinate system are calculated.

[0125] Specifically, the local disparity map is fused with the real-time pitch angle provided by the inertial navigation unit to perform spatial coordinate conversion, including:

[0126] Based on the local disparity map, each pixel point (u i,v i ) corresponds to the disparity value d i Through stereo correction, the focal length f,f of the binocular camera can be obtained x ,f y , the baseline length b, and the principal point of the image (u 0 ,v 0 ), through triangulation, the coordinates of each pixel in the camera coordinate system can be calculated

[0127]

[0128] Then, the original image taken by the left camera is subjected to stereo correction to obtain each pixel point in the corrected left camera image, which can be converted from pixel coordinates to camera coordinates, taking into account the camera posture (such as camera installation height Δh, pitch angle and the horizontal and front-to-back distance Δx from the left and right wheels L(R) ,Δz), and the real-time relative pitch angle of the vehicle measured by the inertial navigation unit Calculate the spatial coordinate X in the wheel coordinate system L(R) ,Y L(R) ,Z L(R) :

[0129]

[0130] In the formula, θ p is the real-time pitch angle of the camera during driving. It is the coordinate of the target point in the wheel coordinate system when it has not been dynamically pitched, where the Z axis is along the vehicle's driving direction and the Y axis is the vertical direction of the vehicle.

[0131] Figure 2 The present invention is an overall processing flow of a road profile estimation method for a binocular system based on deep learning, wherein the above represents the first stage processing steps in this process, determines the key frame input, first scales and fills the original image to meet the input requirements of the target detection network, and then obtains the pixel coordinate information and confidence information of the obstacle in the image through the target detection network and post-processing steps; compares the predetermined vehicle driving path with the pixel coordinates of the obstacle through binocular camera calibration, specifically, when the vehicle driving path passes through an obstacle and when the obstacle confidence is higher than the threshold set in this embodiment, it is used as a key frame input, and the local original image of the obstacle area is selectively cropped using the set cropping threshold, and then input into a deep learning stereo matching convolutional neural network model based on a gated recurrent unit to obtain a disparity map of the obstacle area; uses the external and internal parameters of the inertial measurement unit and the binocular camera to convert the spatial coordinate system, and calculates the road profile information of the obstacle area in the wheel coordinate system.

[0132] This embodiment also proposes a road profile estimation system based on a vehicle-mounted binocular system based on deep learning, including:

[0133] A first image processing module, an obstacle area local image acquisition module, a second image processing module and a millimeter-level road surface profile calculation module;

[0134] The first image processing module is used to obtain the original image, obtain the feature map through the original image, and post-process the feature map to obtain the pixel coordinates and confidence information of the road obstacle;

[0135] Obstacle area local image acquisition module: used to process the pixel coordinates and confidence information of the road obstacle to obtain the local image of the obstacle area of ​​the original image;

[0136] The second image processing module is used to input the local image of the obstacle area of ​​the original image into the deep learning stereo matching model based on the gated recurrent unit for processing, so as to obtain a local disparity map corresponding to the obstacle area;

[0137] Millimeter-level road surface profile calculation module: used to achieve millimeter-level road surface profile estimation by integrating the local disparity map of the corresponding obstacle area with the real-time pitch angle provided by the inertial navigation unit.

[0138] The above are only preferred specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A road profile estimation method for a vehicle-mounted binocular system based on deep learning, characterized in that: include: Acquire an original image, obtain a feature map through the original image, and post-process the feature map to obtain pixel coordinates and confidence information of road obstacles, wherein the original image is a corrected road image; Processing the pixel coordinates and confidence information of the road obstacle to obtain a local image of the obstacle area of ​​the original image; Inputting the local image of the obstacle area of ​​the original image into a deep learning stereo matching model based on a gated recurrent unit for processing to obtain a local disparity map corresponding to the obstacle area; The local disparity map of the corresponding obstacle area is integrated with the real-time pitch angle provided by the inertial navigation unit to achieve millimeter-level road profile estimation.

2. The road profile estimation method of the vehicle-mounted binocular system based on deep learning according to claim 1 is characterized in that: Obtaining the feature map includes: After padding and scaling the original image, the image is input into the convolutional neural network Yolov5s for processing, and the feature map is output; The convolutional neural network Yolov5s includes thirteen convolutional units, each of which is connected to a batch normalization layer and an activation function.

3. The road profile estimation method of the vehicle-mounted binocular system based on deep learning according to claim 2 is characterized in that: The thirteen convolution units include: The first convolution unit includes the convolution layer Conv1_1, with a convolution kernel size of 6, a stride and padding of 2, and an output channel of 32; The second convolution unit includes convolution layer Conv2_1 and C3 module 2, with a convolution kernel size of 3, a stride of 2, a padding of 1, and an output channel of 64; The third convolution unit includes convolution layer Conv3_1 and C3 module 3, with a convolution kernel size of 3, a stride of 2, a padding of 1, and an output channel of 128; The fourth convolution unit includes convolution layer Conv4_1 and C3 module 4, with a convolution kernel size of 3, a stride of 2, a padding of 1, and an output channel of 256; The fifth convolution unit includes convolution layer Conv5_1 and C3 module five, with a convolution kernel size of 3, a stride of 2, a padding of 1, and an output channel of 512; The sixth convolution unit includes an SPPF module with an output channel of 512; The seventh convolution unit includes the convolution layer Conv7_1, the upsampling layer seven, the connection layer Concat1 and the C3 module seven. The convolution kernel size is 1, the step size is 1, the padding is 0, and the output channel is 256; The eighth convolution unit includes a convolution layer Conv8_1, an upsampling layer eight, a connection layer Concat2 and a C3 module eight, the convolution kernel size is 1, the step size is 1, the padding is 0, and the output channel is 128; The ninth convolution unit includes the convolution layer Conv9_1, the connection layer Concat3 and the C3 module nine, with a convolution kernel size of 3, a stride of 2, a padding of 1, and an output channel of 128; The tenth convolution unit includes the convolution layer Conv10_1, the connection layer Concat4 and the C3 module ten, with a convolution kernel size of 3, a stride of 2, a padding of 1, and an output channel of 256; The eleventh convolution unit includes a convolution layer Conv11_1, a convolution kernel size of 1, a stride of 1, a padding of 0, and an output channel of (5+ncls)×3, where n cls To detect the type; The twelfth convolution unit includes the convolution layer Conv12_1, the convolution kernel size is 1, the stride is 1, the padding is 0, and the output channel is (5+n cls )×3; The thirteenth convolution unit includes a convolution layer Conv13_1, a convolution kernel size of 1, a stride of 1, a padding of 0, and an output channel of (5+n cls )×3.

4. The road profile estimation method of the vehicle-mounted binocular system based on deep learning according to claim 3 is characterized in that: Each C3 module in the convolution unit consists of a convolution layer Conv1, a convolution layer Conv2, a bottleneck layer Bottleneck and a convolution layer Conv3. Each convolution layer is followed by a batch normalization layer and an activation function layer. Among them, the convolution kernel size used by the convolution layer Conv1 and the convolution layer Conv2 is 1, and the output channels are half of the output channels of the convolution layer Conv1_1, the convolution layer Conv2_1, the convolution layer Conv3_1, the convolution layer Conv4_1 and the convolution layer Conv5_1 respectively; the convolution kernel size used by the convolution layer Conv3 is 1, and the output channel is the sum of the output channels of the convolution layer Conv1 and the convolution layer Conv2.

5. The road profile estimation method of the vehicle-mounted binocular system based on deep learning according to claim 4 is characterized in that: The convolution layer in the bottleneck layer Bottleneck consists of a convolution layer Conv4 and a convolution layer Conv5, wherein the convolution kernel sizes used by the convolution layer Conv4 and the convolution layer Conv5 are 1 and 3 respectively, and the output channels of the convolution layer Conv4 and the convolution layer Conv5 are the same as the output channels of the convolution layer Conv2.

6. The road profile estimation method of the vehicle-mounted binocular system based on deep learning according to claim 3 is characterized in that: The SPPF module in the sixth convolution unit is composed of a convolution layer ConvP_1, a maximum pooling layer Maxpool1, a maximum pooling layer Maxpool2, a maximum pooling layer Maxpool3 and a convolution layer ConvP_2, and each convolution layer is connected to a batch normalization layer and an activation function layer; Among them, the convolution kernel size used by the convolution layer ConvP_1 is 1, and the output channel is 256; the convolution kernel size used by the maximum pooling layer Maxpool1, the maximum pooling layer Maxpool2, and the maximum pooling layer Maxpool3 is 5, the step size is 1, and the padding is 2; the convolution kernel size used by the convolution layer ConvP_2 is 1, the step size is 1, the padding is 0, and the output channel is 512.

7. The road profile estimation method of the vehicle-mounted binocular system based on deep learning according to claim 1, characterized in that: Obtain the local image of the obstacle area of ​​the original image, including: The pixel coordinates of the road obstacle, the confidence information, and the vehicle driving trajectory are subjected to cropping dynamic stereo matching CDSM to obtain a local image of the obstacle area of ​​the original image.

8. The road profile estimation method of the vehicle-mounted binocular system based on deep learning according to claim 1, characterized in that: Obtaining a local disparity map of the corresponding obstacle area includes: Filling the edge of the local image of the obstacle area of ​​the original image based on a preset condition, processing the filled local image through the deep learning stereo matching model based on the gated recurrent unit, and obtaining a local disparity map of the corresponding obstacle area; The deep learning stereo matching model based on gated recurrent units includes a feature map extraction module, a parameterized full-pair correlation pyramid construction module, an iterative convolution module based on gated recurrent units, and an uncertainty perception refinement module; The feature map extraction module is used to extract feature maps of different scales; The parameterized all-pair correlation pyramid construction module is used to construct a parameterized cost space using multi-Gaussian distribution theory; The iterative convolution module based on the gated recurrent unit is used to process the candidate hidden states by resetting the gates, updating the gates, and obtaining the hidden state output; The uncertainty-aware refinement module is used to estimate the uncertainty map and guide the fusion of the residual map and the disparity map.

9. The road profile estimation method of the vehicle-mounted binocular system based on deep learning according to claim 1, characterized in that: Implementing the millimeter-level road profile estimation includes: For each pixel point (u i ,v i ) corresponds to the disparity value d i Perform stereo calibration to obtain the focal length f,f of the binocular camera x ,f y , baseline length b, and the principal point of the image (u0, v0); Calculate the coordinates of each pixel in the camera coordinate system through triangulation And according to the attitude position of the camera and the real-time relative pitch angle of the vehicle measured by the inertial navigation unit, the spatial coordinates in the wheel coordinate system are calculated.

10. A road profile estimation system based on a vehicle-mounted binocular system based on deep learning, characterized in that: include: A first image processing module, an obstacle area local image acquisition module, a second image processing module and a millimeter-level road surface profile calculation module; The first image processing module is used to obtain an original image, obtain a feature map through the original image, and post-process the feature map to obtain pixel coordinates and confidence information of road obstacles; The obstacle area local image acquisition module is used to process the pixel coordinates and confidence information of the road obstacle to obtain the local image of the obstacle area of ​​the original image; The second image processing module is used to input the local image of the obstacle area of ​​the original image into the deep learning stereo matching model based on the gated recurrent unit for processing, so as to obtain a local disparity map corresponding to the obstacle area; The millimeter-level road surface profile calculation module is used to achieve millimeter-level road surface profile estimation by fusing the local disparity map of the corresponding obstacle area with the real-time pitch angle provided by the inertial navigation unit.

Citation Information

Patent Citations

  • Pavement-learning-based obstacle detection method and apparatus

    CN107909009A

  • Fusion target identification method based on UV parallax detection and YOLOv5

    CN115457508A

  • Road obstacle distance measuring method combining deep learning and binocular vision

    CN116051649A

  • Digital inscription character recognition method based on deep learning

    CN117197815A

  • Rapid binocular stereo matching method

    CN117649436A

Cited By

  • Real height measurement method and system based on monocular camera, and medium

    CN121274849A

  • Methods, systems, and media for measuring true height using a monocular camera

    CN121274849B