Road surface driving quality index measurement method based on binocular stereo matching

By applying the deep learning binocular stereo matching algorithm in road maintenance management, the existing road driving quality index measurement methods are solved, and efficient and accurate road driving quality index measurement is achieved, providing an important basis for road maintenance.

CN120163764APending Publication Date: 2025-06-17ANHUI UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510138387.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-06-17

Smart Images

  • Figure CN120163764A_ABST
    Figure CN120163764A_ABST
Patent Text Reader

Abstract

The invention provides a pavement driving quality index measurement method based on binocular stereo matching. The method comprises the following steps: capturing a road image through a binocular camera, extracting preliminary depth information, generating a fine feature map through multi-scale feature extraction and spatial pyramid pooling operation, constructing a cost matching space for parallax regression calculation, and generating a parallax map. And calculating the depth of a central point by combining camera parameters, and outputting a road surface driving quality index (RQI) and a driving track. The method also records depth information and Beidou position information in the driving process. Experiments prove that the light-weight GSConv convolution and the LWC residual block are adopted, the feature extraction process is optimized, the matching speed and precision are improved, accurate measurement of the international flatness index is achieved, and a quantitative index is provided for road maintenance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of traffic engineering and computer vision, and particularly relates to a method for measuring pavement ride quality index based on binocular stereo matching. Background Technique

[0002] With the continuous advancement of the construction of a transportation power in China, the detection of pavement ride quality has become increasingly important. The detection of pavement ride quality is not only a key foundation for ensuring road driving safety and improving driving comfort, but also a necessary prerequisite for road maintenance and repair. The pavement ride quality index (RQI) is an important indicator for evaluating pavement smoothness, which is calculated based on the internationally common international roughness index (IRI). The value of RQI directly affects the driving safety, comfort, load-bearing capacity and service life of the road. A high RQI value indicates good pavement smoothness and can provide a safer and more comfortable driving experience; while a low RQI value indicates uneven pavement, increasing vehicle vibration and driving resistance, affecting driving safety and accelerating pavement damage. Therefore, accurate RQI measurement is crucial for ensuring traffic safety and improving driving comfort.

[0003] China started relatively late in the field of pavement ride quality measurement, and the detection technology and equipment are relatively lagging behind. Many road maintenance and management departments still rely on traditional manual detection methods. Existing RQI measurement methods include manual detection, bump integrator, precise level, and lidar sensors, etc., but each has its limitations: manual detection has low efficiency and large human errors, and is not suitable for long-distance and high-precision detection; the bump integrator is affected by vehicle speed and vehicle vibration damping performance, and needs to be frequently calibrated, with poor applicability; the precise level method has high accuracy, but the detection speed is slow and it is difficult to meet the long-distance measurement requirements; although the lidar sensor technology has high accuracy, the equipment cost and maintenance cost are high.

[0004] To improve the efficiency and accuracy of RQI measurement, it is urgent to develop more advanced measurement technologies. Applying artificial intelligence and deep learning algorithms to achieve automatic measurement of pavement smoothness, or researching and developing new measurement equipment with low cost and high efficiency to meet the needs of modern road maintenance and management has important practical significance. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for measuring pavement ride quality index based on binocular stereo matching, which solves the problems of slow speed, high cost and insufficient accuracy of traditional RQI calculation methods.

[0006] The technical solution adopted by the present invention is a method for measuring pavement ride quality index based on binocular stereo matching, including the following steps:

[0007] Step S1, install and calibrate a binocular camera at the rear of the vehicle to obtain the left view I Land the right view I R , as well as the focal length f and baseline b parameters of the binocular camera;

[0008] Step S2, capture the initial depth information of the road;

[0009] Step S3, perform multi-scale feature extraction to generate a more refined feature map;

[0010] Step S4, extract the spatial information of different scales of the more refined feature map, and splice to obtain a two-dimensional fusion feature map;

[0011] Step S5, construct a cost matching space;

[0012] Step S6, capture the information and spatial relationship related to the disparity, generate a 4D tensor, and perform disparity regression calculation;

[0013] Step S7, repeat steps S2 - S6, and output the disparity values m1, m2,..., m of the central pixel points of the disparity map at each moment k , combine the focal length f and baseline b parameters of the binocular camera, calculate the vertical distance from the road surface corresponding to the pixel point to the plane where the binocular camera is located, and then calculate the road surface driving quality index.

[0014] Further, in the step S2, the left view I L and the right view I R first pass through a convolution respectively, and then pass through two Fused - MBConv convolutional layers. Batch normalization and ReLU activation functions are added after the Fused - MBConv convolutional layer, and finally the initial feature maps of the left view I L and the right view I R are obtained. The specific process is as follows:

[0015] F stemL = ReLU(BatchNorm(Fused - MBConv(ReLU(BatchNorm(Fused - MBConv(Conv(I L )))))))

[0016] F stemR = ReLU(BatchNorm(Fused - MBConv(ReLU(BatchNorm(Fused - MBConv(conv(I R )))))))

[0017] where, I L represents the left view obtained by the binocular camera, I RDenote the right view obtained by the binocular camera, Conv(·) denotes ordinary convolution, BatchNorm(·) denotes batch normalization operation, ReLU(·) denotes ReLU activation function, Fused-MBConv(·) denotes Fused-MBConv module, F stemL denotes the preliminary feature map of the left view, F stemR denotes the preliminary feature map of the right view.

[0018] Furthermore, in step S3, through two cascaded LWC residual layers, multi-scale feature extraction is performed on the preliminary feature maps of the left and right views, and finally a more refined feature map is generated. The specific formula is as follows:

[0019] F 2L = GSConv(LWC2(LWC1(F stemL )))

[0020] F 2R = GSConv(LWC2(LWC1(F stemR )))

[0021] where, F 2L denotes the more refined feature map of the left view, F 2R denotes the more refined feature map of the right view, F stemL denotes the preliminary feature map of the left view, F stemR denotes the preliminary feature map of the right view, LWC1 denotes the first layer of LWC residual layer, LWC2 denotes the second layer of LWC residual layer, GSConp(·) denotes GSConv convolution operation.

[0022] Furthermore, in step S4, spatial pyramid pooling operation is performed on the more refined feature maps of the left and right views to extract spatial information of different scales of the more refined feature maps, and a two-dimensional fusion feature map is obtained by splicing. The specific process includes:

[0023] S41, through spatial pyramid pooling operation, extract spatial information of different scales. The more refined feature maps F 2L 、F 2R of the left and right views respectively pass through Fused-MBConv with convolution kernel sizes of 1×1, 3×3, 5×5, 7×7 and strides of 1, 2, 3, 4 in sequence to obtain multi-scale features of F 2L 、F 2R . The specific formula is as follows:

[0024] F multi-scaleL = concat(Fused-MBConv(F 2L , k j , s i )|kj = 1, 3, 5, 7, s i = 1, 2, 3, 4)

[0025] F multi-scaleR = concat(Fused - MBConv(F 2R , k j , s i ) | k j = 1, 3, 5, 7, s i = 1, 2, 3, 4)

[0026] where Fused - MBConv(F 2L , k j , s i ) represents performing a Fused - MBConv convolution operation on the more refined feature map F of the left view with a convolution kernel size of k j and a stride of s i ; Fused - MBConv(F 2L , k 2R , s j , s i ) represents performing a Fused - MBConv convolution operation on the more refined feature map F of the right view with a convolution kernel size of k j and a stride of s i ; concat(·) represents a concatenation operation, F 2R represents the multi - scale features extracted from the left view, and F multi-scaleL represents the multi - scale features extracted from the right view; multi-scaleR

[0027] S42, for the more refined feature maps of the left and right, the multi - scale features F multi-scaleL , F multi-scaleR , perform bilinear interpolation and concatenation operations to obtain the two - dimensional fused feature maps F L , F R of the left and right views. The specific formula is as follows:

[0028] F L = concat(bilinear(LWC2(LWC1(F stemL ))), size(F 2L ))), F 2L , bilinear(F multi-scsleL , size(F 2L )))

[0029] F R = concat(bilinear(LWC2(LWC1(F stemR ))), size(F​2R )) , F 2R , bilinear(F multi-scaleR , size(F 2R )))

[0030] Among them, F L represents the two-dimensional fusion feature map of the left view, F R represents the two-dimensional fusion feature map of the right view, F stemL represents the preliminary feature map of the left view, LWC1 represents the first layer of LWC residual layer, LWC2 represents the second layer of LWC residual layer, F 2L represents a more refined feature map of the left view, F 2R represents a more refined feature map of the right view, F multi-scaleL represents the multi-scale features extracted through the left view, F multi-scaleR represents the multi-scale features extracted through the right view, bilinear(·) represents the bilinear interpolation operation for adjusting the size of the feature map, size(·) represents the size of the extracted feature map, concat represents the concatenation operation, H′ represents the height of the feature map, W′ represents the width of the feature map, C represents the number of channels, represents the set of real numbers.

[0031] Furthermore, the step S5 specifically includes:

[0032] S51, shifting the two-dimensional fusion feature map F R of the right view to the right by different disparities d, and the specific formula is as follows:

[0033]

[0034] where d ∈ [0, D), D represents the disparity dimension parameter, h represents the height index of the feature map, w represents the width index of the feature map, c represents the channel index, W′ represents the width of the feature map, F R represents the two-dimensional fusion feature map of the right view, represents the shifted feature map of the right view;

[0035] S52, performing feature concatenation on the two-dimensional fusion feature map F L of the left view and the shifted feature map of the right view, and the specific formula is as follows:

[0036]

[0037] where Cost represents the four-dimensional matching cost space, h represents the height index of the feature map, w represents the width index of the feature map, : represents taking all elements of a certain dimension, Concat(·) represents the concatenation operation in the channel dimension, FL The 2D fusion feature map representing the left view, The right view translation feature map, D represents the disparity dimension, H' represents the height of the feature map, W' represents the width of the feature map, C represents the number of channels, represents the set of real numbers.

[0038] Further, in step S6, the matching cost space Cost is input into a 3D convolutional network to capture disparity-related information and spatial relationships, generating a 4D tensor Perform disparity regression calculation to obtain a disparity map, specifically including:

[0039] S61, in the disparity dimension D, obtain the predicted value of the disparity through the 4D tensor C', and calculate the final disparity according to its probability at each candidate disparity position. The specific formula is as follows:

[0040]

[0041] where P(d, h, w) represents the probability that the disparity at (h, w) is d, h represents the height index of the feature map, w represents the width index of the feature map, C'(d, h, w) represents the value of the 4D tensor C' generated by the matching cost space at position (d, h, w), e represents the natural constant, and D represents the disparity dimension;

[0042] S62, perform disparity regression calculation to obtain a disparity map. The specific formula is as follows:

[0043]

[0044] where, represents the final predicted disparity value at position (h, w), h represents the height index of the feature map, w represents the width index of the feature map, P(d, h, w) represents the probability that the disparity at position (h, w) is d, and D represents the disparity dimension;

[0045] S63, calculate the loss function during the training process. The specific formula is as follows;

[0046]

[0047] where Smooth L1 (·) represents the output value of the loss, x represents the predicted disparity value and the error between the actual disparity d, that is

[0048] Further, step S7 specifically includes:

[0049] S71. Using the disparity value of the central pixel point, the focal length f of the camera, and the baseline b parameter, calculate the vertical distance from the road surface corresponding to each pixel point to the plane where the binocular camera is located, denoted as h1, h2, …, h k , and the specific calculation formula is as follows:

[0050]

[0051] where f represents the focal length of the binocular camera, b represents the baseline of the binocular camera, m i represents the disparity value of the central pixel at the i-th moment, and h i represents the vertical distance from the central pixel point to the camera plane at the i-th moment, i = 1, 2, …, k, and k represents the total number of recorded moments;

[0052] S72. Based on the vertical distance from the central pixel point to the camera plane obtained in S71, calculate the international roughness index, and the specific formula is as follows:

[0053] Δd i = |h i+1 - h i |

[0054]

[0055] where Δd i represents the longitudinal displacement at the i-th moment, h i+1 represents the vertical distance from the pixel center point to the camera plane at the (i + 1)-th moment, h i represents the vertical distance from the pixel center point to the camera plane at the i-th moment, S represents the cumulative longitudinal displacement on the vehicle driving trajectory, k represents the total number of recorded moments, Dist represents the total length of the entire driving trajectory, and IRI represents the international roughness index;

[0056] S73. Calculate the pavement ride quality index, and the specific formula is as follows:

[0057]

[0058] where IRI represents the international roughness index, a0 and a1 represent constants, a0 is 0.026, a1 is 0.65, e is the natural constant, and RQI represents the pavement ride quality index.

[0059] The beneficial effects of the present invention are as follows:

[0060] 1. The present invention adopts a binocular stereo matching algorithm based on deep learning, which can automatically extract road surface texture features from the images collected by the binocular camera, greatly improving the speed and accuracy of matching.

[0061] 2. By introducing lightweight GSConv convolution and LWC residual blocks, the present invention optimizes the feature extraction process while maintaining depth and rich features, effectively reducing the computational complexity.

[0062] 3. The design of spatial pyramid pooling operation and Fused-MBConv convolutional layer in the present invention enhances the ability to capture spatial information of different scales and improves the accuracy of feature fusion.

[0063] 4. By combining the disparity regression technology, the present invention realizes the accurate measurement of the international roughness index, and through the calculation of the riding quality index RQI of the road surface, provides a quantitative index for road maintenance and evaluation, significantly improves the automation level of the road surface quality detection process, and provides an important basis for road maintenance decision-making. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0065] Figure 1 It is the overall flowchart of the method of the present invention.

[0066] Figure 2 It is the structural block diagram of the riding quality index measuring device of the present invention.

[0067] Figure 3 It is the installation schematic diagram of the riding quality index measuring device of the present invention.

[0068] Figure 4 It is the structural diagram of the improved binocular stereo matching algorithm of the present invention.

[0069] Figure 5 It is the detailed schematic diagram of the LWC residual block.

[0070] Figure 6 It is the detailed schematic diagram of the feature fusion of the SPPC module.

[0071] Figure 7 It is the flowchart of the disparity reading of the method of the present invention.

[0072] Figure 8 It is the flowchart of calculating the riding quality index of the road surface of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0073] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0074] Embodiment 1

[0075] An embodiment of the present invention provides a method for measuring the pavement driving quality index based on binocular stereo matching, as Figures 1 to 8 shown, and the specific steps include:

[0076] Step S1, install a binocular camera at the rear of the vehicle, and calibrate the binocular camera by the Zhang Zhengyou calibration method to obtain the left view I L and the right view I R , as well as the focal length f and baseline b parameters of the binocular camera. And match the Beidou positioning module and the data processing module to collect data, calculate the driving quality index through the binocular stereo matching algorithm, and then display the calculated result through the display module. The specific modules include:

[0077] G1: Beidou positioning module.

[0078] This module is based on the Chinese Beidou Navigation Satellite System (BDS) and provides positioning, navigation, and timing services. It can provide users with high-precision position, speed, and time information all-weather and all-time globally. The present invention realizes the collection of vehicle positions and the calculation of driving trajectories by receiving Beidou satellite signals.

[0079] G2: Data processing module.

[0080] This module is mainly composed of an NVIDIA Jetson Orin Nano small computer. Jetson Orin Nano is an NVIDIA embedded edge computing AI application platform. The present invention mainly uses this device to calculate the binocular stereo matching algorithm and the driving quality index.

[0081] G3: Display module.

[0082] This module uses an ordinary LCD display screen to display the results of the binocular stereo matching algorithm and the calculation of the driving quality index.

[0083] G4: Binocular camera.

[0084] Collect information using a binocular camera of Intel RealSense D435i. This type of camera is a high-precision device designed for 3D vision applications and can stably provide depth information and image data under various indoor and outdoor lighting conditions. The camera integrates binocular vision technology and an Inertial Measurement Unit (IMU), and performs excellently in aspects such as visual positioning and navigation, three-dimensional object recognition, and spatial motion tracking. When installing, adjust the angle and height of the binocular camera so that the binocular camera is perpendicular to the road surface, and the height of the binocular camera from the ground is h. Optimize the camera to receive the video stream, set the resolution to h×w, and select the central point pixel.

[0085] The above devices and modules all belong to the prior art. The computer can download the corresponding driver program from the official website of the Intel RealSense camera and install it. Insert the Intel RealSense D435i into the USB 3.0 port of the computer, and use the OpenCV library to obtain the left and right video streams of the project camera. Connect the Beidou positioning module to the computer, select the appropriate connection cable according to the interface type of the module, and install the driver program of the Beidou positioning module on the computer. Open the driver program, find the port corresponding to the Beidou positioning module, and set the parameters of the baud rate, parity bit, and data bit.

[0086] Step S2, based on S1, initially extract the feature maps of the left view I L and the right view I R to capture the initial depth information of the road.

[0087] The left view I L and the right view I R first pass through a common convolution with a convolution kernel size of 3×3 and a stride of 2 respectively, and then pass through two Fused-MBConv convolution layers with a convolution kernel size of 3×3. A Batch Normalization Layer and a ReLU activation function are added after the Fused-MBConv convolution layer, and finally the initial feature maps of the left view I L and the right view I R are obtained. The specific process is as follows:

[0088] F stemL = ReLU(BatchNorm(Fused-MBConv(ReLU(BatchNorm(Fused-MBConv(Conv(I L )))))))

[0089] F stemR = ReLU(BatchNorm(Fused-MBConv(ReLU(BatchNorm(Fused-MBConv(Conv(IR )))))))

[0090] Among them, I L represents the left view obtained by the binocular camera, and I R represents the right view obtained by the binocular camera. Conv(·) represents ordinary convolution, BatchNorm(·) represents batch normalization operation, ReLU(·) represents ReLU activation function, and Fused-MBConv(·) represents Fused-MBConv module. F stemL represents the preliminary feature map of the left view, and F stemR represents the preliminary feature map of the right view.

[0091] Step S3: Based on the preliminary feature maps of the left view and the right view preliminarily extracted in S2, multi-scale feature extraction is performed on them through two cascaded LWC residual layers, and finally a more refined feature map is generated. The specific formula is as follows:

[0092] F 2L = GSConv(LWC2(LWC1(F stemL )))

[0093] F 2R = GSConv(LWC2(LWC1(F stemR )))

[0094] Among them, F 2L represents the more refined feature map of the left view, and F 2R represents the more refined feature map of the right view, F stemL represents the preliminary feature map of the left view, and F stemR represents the preliminary feature map of the right view. LWC1 represents the first layer of LWC residual layer, LWC2 represents the second layer of LWC residual layer, and GSConp(·) represents GSConv convolution operation.

[0095] The residual layer is stacked by residual blocks. The first residual layer is composed of M residual blocks, and the second residual layer is composed of N residual blocks. The numbers of M and N are adjustable, and in the present invention, they are preferably set to 3 and 16.

[0096] The residual block is constructed by GSConv, and the structure of the residual block is as Figure 5 shown. If the input feature map of the residual block is C1, then C1 respectively passes through GSConv_1 convolution and convolution downsampling to generate feature maps C2 and C3. The feature map C2 passes through GSConv_2 to obtain the feature map C4. The feature map C3 and the feature map C4 are concatenated in the channel dimension, and finally the final feature map C5 is output through the ReLU activation function.

[0097] Step S4: Based on the more refined feature maps of the left view and the more refined feature maps of the right view obtained in S3, perform spatial pyramid pooling operation (SPPC, Spatial Pyramid Pooling Concat) to extract the spatial information of different scales of the more refined feature maps, and splice them to obtain a two-dimensional fusion feature map.

[0098] S41: Through the spatial pyramid pooling operation, extract the spatial information of different scales, as Figure 6 shown, the more refined feature map F 2L of the left view, and the more refined feature map F 2R of the right view, respectively pass through Fused-MBConv with a convolution kernel size of 1×1, 3×3, 5×5, 7×7 and a stride of 1, 2, 3, 4 to obtain the multi-scale features of F 2L and F 2R . The specific formula is as follows:

[0099] F multi-scaleL = concat(Fused-MBConv(F 2L , k j , s i ) | k j = 1, 3, 5, 7, s i = 1, 2, 3, 4)

[0100] F multi-scaleR = concat(Fused-MBConv(F 2R , k j , s i ) | k j = 1, 3, 5, 7, s i = 1, 2, 3, 4)

[0101] Among them, Fused-MBConv(F 2L , k j , s i ) represents performing a Fused-MBConv convolution operation on the more refined feature map F j of the left view with a convolution kernel size of k i and a stride of s 2L ; Fused-MBConv(F 2R , k j , s i ) represents performing a Fused-MBConv convolution operation on the more refined feature map F j of the right view with a convolution kernel size of k i and a stride of s 2R ; concat(·) represents the splicing operation, F multi-scaleLDenote the multi-scale features extracted from the left view as F multi-scaleR Denote the multi-scale features extracted from the right view.

[0102] S42. For the more refined left and right feature maps F 2L and F 2R obtained in S3, and the multi-scale features F multi-scaleL and F multi-scaleR obtained in S41, perform bilinear interpolation and concatenation operations. The target size of the bilinear interpolation is equal to the tensor size of the more refined feature map output by S3. Finally, obtain the two-dimensional fusion feature maps F L and F R of the left and right views. The specific formula is as follows:

[0103] F L = concat(bilinear(LWC2(LWC1(F stemL )), size(F 2L )), F 2L , bilinear(F multi-scaleL , size(F 2L )))

[0104] F R = concat(bilinear(LWC2(LWC1(F stemR )), size(F 2R )), F 2R , bilinear(F multi-scaleR , size(F 2R )))

[0105] Among them, F L represents the two-dimensional fusion feature map of the left view, F R represents the two-dimensional fusion feature map of the right view, F stemL represents the preliminary feature map of the left view, LWC1 represents the first LWC residual layer, LWC2 represents the second LWC residual layer, F 2L represents the more refined feature map of the left view, F 2R represents the more refined feature map of the right view, F multi-scaleL represents the multi-scale features extracted from the left view, F multi-scaleR represents the multi-scale features extracted from the right view, bilinear(·) represents the bilinear interpolation operation for adjusting the size of the feature map, size(·) represents the size of the extracted feature map, and concat represents the concatenation operation; H' represents the height of the feature map, W' represents the width of the feature map, and C represents the number of channels. Denote the set of real numbers.

[0106] Step S5, based on S4, construct a cost matching space through the two-dimensional fusion feature maps of the left and right views.

[0107] S51, translate the two-dimensional fusion feature map F of the right view R to the right by different disparities d. The specific formula is as follows:

[0108]

[0109] where d ∈ [0, D), D represents the disparity dimension parameter, which is preferably 192 in the present invention, h represents the height index of the feature map, w represents the width index of the feature map, c represents the channel index, W' represents the width of the feature map, and F R represents the two-dimensional fusion feature map of the right view, represents the translated feature map of the right view.

[0110] S52, perform feature concatenation on the two-dimensional fusion feature map F of the left view L and the translated feature map of the right view The specific formula is as follows:

[0111]

[0112] where Cost represents the four-dimensional matching cost space, h represents the height index of the feature map, w represents the width index of the feature map, : represents taking all elements of a certain dimension, Concat(·) represents the concatenation operation in the channel dimension, F L represents the two-dimensional fusion feature map of the left view, represents the translated feature map of the right view, D represents the disparity dimension, H' represents the height of the feature map, W' represents the width of the feature map, C represents the number of channels, represents the set of real numbers.

[0113] Step S6, input the matching cost space Cost in S5 into a 3D convolutional network to capture disparity-related information and spatial relationships, generate a 4D tensor perform disparity regression calculation, and finally obtain the disparity map.

[0114] S61, in the disparity dimension D, adopt the Soft-argmin mechanism for C' to obtain the predicted value of the disparity, and calculate the final disparity according to its probability at each candidate disparity position. The specific formula is as follows:

[0115]

[0116] Among them, P(d, h, w) represents the probability that the disparity is d at (h, w), h represents the height index of the feature map, w represents the width index of the feature map, C'(d, h, w) represents the value of the 4D tensor C' generated by the matching cost space at the position (d, h, w), e represents the natural constant, and D represents the disparity dimension.

[0117] S62. Perform disparity regression calculation to obtain a disparity map. The specific formula is as follows:

[0118]

[0119] Among them, represents the final predicted disparity value at the position (h, w), h represents the height index of the feature map, w represents the width index of the feature map, P(d, h, w) represents the probability that the disparity is d at the position (h, w), and D represents the disparity dimension.

[0120] S63. Calculate the loss function during the training process. The specific formula is as follows;

[0121]

[0122] Among them, Smooth L1 (·) represents the output value of the loss, x represents the predicted disparity value and the error between the actual disparity d, that is

[0123] Step S7. Repeat steps S2 - S6, and output the disparity values m1, m2…, m of the central pixel points of the disparity map at each moment k , and combine the binocular camera focal length f and the baseline b parameters to calculate the vertical distance from the road surface corresponding to the pixel points to the plane where the binocular camera is located, and then calculate the road surface driving quality index.

[0124] S71. Use the disparity values of the central pixel points, the focal length f of the camera, and the baseline b parameters to calculate the vertical distance from the road surface corresponding to each pixel point to the plane where the binocular camera is located, denoted as h1, h2…, h k , and the specific calculation formula is as follows:

[0125]

[0126] Among them, f represents the binocular camera focal length, b represents the baseline of the binocular camera, m i represents the disparity value of the central pixel at the i-th moment, h i represents the vertical distance from the central pixel point to the camera plane at the i-th moment, i = 1, 2,…, k, and k represents the total number of recorded moments.

[0127] S72. Based on S71, obtain the vertical distance from the central pixel point to the camera plane at each moment, and calculate the international roughness index. The specific formula is as follows:

[0128] Δd i =|h i+1 -h i |

[0129]

[0130] where Δd i represents the longitudinal displacement at the i-th moment, with the unit of meter (m), h i+1 represents the vertical distance from the pixel center point to the camera plane at the (i + 1)-th moment, h i represents the vertical distance from the pixel center point to the camera plane at the i-th moment, S represents the cumulative longitudinal displacement on the vehicle driving trajectory, k represents the total number of recorded moments, Dist represents the total length of the entire driving trajectory, with the unit of kilometer (km), and IRI represents the international roughness index.

[0131] S73. Calculate the pavement ride quality index. The specific formula is as follows:

[0132]

[0133] where IRI represents the international roughness index, a0 and a1 represent constants, a0 is 0.026, a1 is 0.65, e is the natural constant, and RQI represents the pavement ride quality index.

[0134] The pavement ride quality index RQI is a comprehensive index that can be used to evaluate pavement smoothness and riding comfort, provide a basis for road management and maintenance, improve the road riding experience and safety, and can also accurately locate road problem areas, guide the location and method of maintenance operations, and improve the repair efficiency.

[0135] RQI is usually divided into five grades. RQI≥90 indicates excellent, 80≤RQI<90 indicates good, 70≤RQI<80 indicates fair, 60≤RQI<70 indicates poor, and RQI<60 indicates very poor. Different grades correspond to different maintenance requirements and repair strategies. Roads with excellent grades indicate good riding quality and meet the requirements, while roads with poor grades need to be repaired in a timely manner.

[0136] To evaluate the compatibility of the Fused-MBConv, LWC, and SPPC modules, ablation experiments were conducted on the proposed network structure. The results are shown in Table 1, where √ indicates the effect after using the corresponding module.

[0137] Table 1 Results of ablation experiments

[0138]

[0139] As can be seen from the data analysis in Table 1, compared with the baseline model PSMNet (i.e., without using Fused-MBConv, LWC, SPPC), the endpoint error (EPE) is reduced by 2.9%, that is, a 2.9% improvement in detection accuracy is achieved. In addition, this module also shows an optimization effect in terms of the number of parameters. The integration of the Dropout operation and the SE module in the Fused-MBConv module not only enhances the prediction accuracy of the model, but also effectively reduces the number of parameters in the network and improves the operation speed of the network. After introducing the improved LWC module, the model achieves a significant 45% reduction in the number of parameters (Parameters), and the detection accuracy is also improved by 2.8%. This result shows that by reducing the number of residual blocks, the number of parameters in the network is greatly reduced; by combining GSConv with grouped convolution and performing reasonable grouping and shuffling operations on the feature maps, the richness and accuracy of features can be ensured while reducing the number of parameters, thereby effectively improving the prediction accuracy of the model.

[0140] The ablation experiment for SPPC shows that the SPPC module improves the detection accuracy by 7.3% compared with the baseline model PSMNet without increasing the number of parameters, verifying that the present invention can effectively expand the receptive field, connect the context and obtain spatial information at different scales by using convolution kernels of different sizes and skip-layer splicing. Then, the three different modules are fused. The final model has a 50.19% decrease in the number of parameters and a 11.28% decrease in the error compared with the baseline model PSMNet, which proves the feasibility of the technical solution of the present invention and can effectively improve the network performance.

[0141] Table 2 Comparison between the present invention and other existing methods

[0142]

[0143]

[0144] The comparison of the endpoint error (EPE) between the method proposed in the present invention and other existing deep learning stereo matching methods on the SceneFlow test set is shown in Table 2. Among them, the endpoint error of the present invention is the lowest and the performance is the best, corresponding to the shallow Fused-MBConv + LWC residual block + SPPC module in Table 1. As can be seen from Table 2, among the other 7 methods, PSMNet has the best effect. However, compared with the PSMNet network, the method of the present invention reduces the number of parameters significantly while also reducing the endpoint error (EPE).

[0145] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiment.

[0146] The above description is only for the preferred embodiments of the present invention and is not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.

Claims

1. A road ride quality index measurement method based on binocular stereo matching, characterized in that: The following steps are involved: Step S1: Install and calibrate a binocular camera at the rear of the car to obtain a left view I L and right view I R , as well as the focal length f and baseline b parameters of the binocular camera; Step S2, capturing preliminary depth information of the road; Step S3, performing multi-scale feature extraction to generate a more refined feature map; Step S4, extracting spatial information of different scales of more refined feature maps, and splicing them to obtain a two-dimensional fused feature map; Step S5, constructing a cost matching space; Step S6, capturing disparity-related information and spatial relationships, generating a 4D tensor, and performing disparity regression calculation; Step S7, repeat steps S2-S6, and output the disparity values ​​m1, m2, m3, m4, m5, m6, m7, m8, m9, m10, m11, m22, m12, m13, m14, m15, m16, m17, m18, m29, m20, m21, m21, m30, m19, m21, m22, m31, m42, m k , combined with the focal length f of the binocular camera and the baseline b parameter, the vertical distance from the road surface corresponding to the pixel point to the plane where the binocular camera is located is calculated, and then the road driving quality index is calculated.

2. The method for measuring road ride quality index based on binocular stereo matching according to claim 1, characterized in that: In the step S2, the left view I L and right view I R They first pass through a convolution, and then pass through two Fused-MBConv convolution layers. The Fused-MBConv convolution layer is followed by batch normalization and ReLU activation function, and finally the left view I is obtained. L and right view I R The preliminary feature map of , the specific process is as follows: Among them, I L Represents the left view obtained by the binocular camera, I R represents the right view obtained by the binocular camera, Conv(·) represents ordinary convolution, BatchNorm(·) represents batch normalization operation, ReLU(·) represents ReLU activation function, Fused-MBConv(·) represents Fused-MBConv module, F stemL represents the preliminary feature map of the left view, F stemR Represents the preliminary feature map of the right view.

3. The method for measuring road ride quality index based on binocular stereo matching according to claim 1, characterized in that: In step S3, multi-scale feature extraction is performed on the preliminary feature maps of the left and right views through two layers of LWC residual layers connected in series, and finally a more refined feature map is generated. The specific formula is as follows: F 2L =GSConv(LWC2(LWC1(F stemL ))) F 2R =GSConv(LWC2(LWC1(F stemR ))) Among them, F 2L represents a more refined feature map of the left view, F 2R represents a more refined feature map of the right view, F stemL represents the preliminary feature map of the left view, F stemR represents the preliminary feature map of the right view, LWC1 represents the first LWC residual layer, LWC2 represents the second LWC residual layer, and GSConv(·) represents the GSConv convolution operation.

4. The method for measuring road ride quality index based on binocular stereo matching according to claim 1, characterized in that: In step S4, a spatial pyramid pooling operation is performed on the finer feature maps of the left and right views to extract spatial information of different scales of the finer feature maps, and the two-dimensional fused feature maps are obtained by splicing. The specific process includes: S41, through the spatial pyramid pooling operation, extracts spatial information of different scales, and the left and right views have more refined feature maps F 2L 、F 2R , respectively, through the Fused-MBConv with convolution kernel sizes of 1×1, 3×3, 5×5, 7×7 and step sizes of 1, 2, 3, and 4, and obtain F 2L 、F 2R The multi-scale features of are as follows: F multi-scaleL =concat(Fused-MBConv(F 2L ,k j ,s i )|k j =1,3,5,7,s i =1,2,3,4) F multi-scaleR =concat(Fused-MBConv(F 2R ,k j ,s i )|k j =1,3,5,7,s i =1,2,3,4) Among them, Fused-MBConv(F 2L ,k j ,s i ) indicates that when the convolution kernel size is k j and step length is s i Next, a more refined feature map F for the left view 2L Perform Fused-MBConv convolution operation; Fused-MBConv (F 2R ,k j ,s i ) indicates that when the convolution kernel size is k j and step length is s i Next, a more refined feature map F for the right view 2R Perform Fused-MBConv convolution operation; concat(·) represents concatenation operation, F multi-scaleL represents the multi-scale features extracted from the left view, F multi-scaleR represents the multi-scale features extracted from the right view; S42, for more refined feature maps on the left and right, multi-scale features F multi-scaleL ,F multi-scaleR , perform bilinear interpolation and splicing operations to obtain the two-dimensional fusion feature map F of the left and right views L 、F R , the specific formula is as follows: F L =concat(bilinear(LWC2(LWC1(F stemL )),size(F 2L )),F 2L ,bilinear(F multi-scaleL ,size(F 2L ))) F R =concat(bilinear(LWC2(LWC1(F stemR )),size(F 2R )),F 2R ,bilinear(F multi-scaleR ,size(F 2R ))) Among them, F L Represents the 2D fused feature map of the left view, F R Represents the 2D fused feature map of the right view, F stemL represents the preliminary feature map of the left view, LWC1 represents the first LWC residual layer, LWC2 represents the second LWC residual layer, and F 2L represents a more refined feature map of the left view, F 2R represents a more refined feature map of the right view, F multi-scaleL Denotes the multi-scale features extracted from the left view, F multi-scaleR represents the multi-scale features extracted from the right view, bilinear(·) represents the bilinear interpolation operation, which is used to adjust the size of the feature map, size(·) represents the size of the extracted feature map, and concat represents the concatenation operation. H' represents the height of the feature map, W' represents the width of the feature map, and C represents the number of channels. Represents the set of real numbers.

5. The method for measuring road ride quality index based on binocular stereo matching according to claim 1, characterized in that: The step S5 specifically includes: S51, the two-dimensional fusion feature map F of the right view R Shift rightward according to different parallax d. The specific formula is as follows: Where d∈[0,D), D represents the disparity dimension parameter, h represents the height index of the feature map, w represents the width index of the feature map, c represents the channel index, W' represents the width of the feature map, and F R Represents the 2D fused feature map of the right view, represents the right view translation feature map; S52, two-dimensional fusion feature map F of the left view L And the right view translation feature map Perform feature splicing, the specific formula is as follows: Among them, Cost represents the four-dimensional matching cost space, h represents the height index of the feature map, w represents the width index of the feature map, : represents taking all elements of a certain dimension, Concat(·) represents the concatenation operation on the channel dimension, and F L represents the 2D fused feature map of the left view, represents the right view translation feature map, D represents the disparity dimension, H' represents the height of the feature map, W' represents the width of the feature map, and C represents the number of channels. Represents the set of real numbers.

6. The method for measuring road ride quality index based on binocular stereo matching according to claim 1, characterized in that: In step S6, the matching cost space Cost is input into the 3D convolutional network to capture the disparity-related information and spatial relationship and generate a 4D tensor. Perform disparity regression calculation to obtain a disparity map, including: S61, in the disparity dimension D, the predicted value of the disparity is obtained through the 4D tensor C', and the final disparity is calculated according to its probability at each candidate disparity position. The specific formula is as follows: Among them, P(d,h,w) represents the probability that the disparity at (h,w) is d, h represents the height index of the feature map, w represents the width index of the feature map, C'(d,h,w) represents the value of the 4D tensor C' generated by the matching cost space at the position (d,h,w), e represents a natural constant, and D represents the disparity dimension; S62, performing disparity regression calculation to obtain a disparity map. The specific formula is as follows: in, Represents the final predicted disparity value at position (h, w), h represents the height index of the feature map, w represents the width index of the feature map, P(d, h, w) represents the probability that the disparity at position (h, w) is d, and D represents the disparity dimension; S63, calculating the loss function in the training process, the specific formula is as follows; Among them, Smooth L1 (·) represents the output value of the loss, and x represents the predicted disparity value The error between the actual disparity d is 7. The method for measuring road ride quality index based on binocular stereo matching according to claim 1, characterized in that: The step S7 specifically includes: S71, using the disparity value of the central pixel point and the focal length f and baseline b parameters of the camera, calculate the vertical distance from the road surface corresponding to each pixel point to the plane where the binocular camera is located, recorded as h1, h2…, h k , the specific calculation formula is as follows: Among them, f represents the focal length of the binocular camera, b represents the baseline of the binocular camera, and m i represents the disparity value of the central pixel at the i-th moment, h i represents the vertical distance from the center pixel to the camera plane at the i-th moment, i = 1, 2, ..., k, k represents the total time of recording; S72, based on S71, obtain the vertical distance from the central pixel point to the camera plane at each moment, and calculate the international flatness index. The specific formula is as follows: Δd i =|h i+1 -h i | Where, Δd i represents the longitudinal displacement at the i-th moment, h i+1 Indicates the vertical distance from the pixel center point to the camera plane at the i+1th moment, h i represents the vertical distance from the pixel center point to the camera plane at the i-th moment, S represents the accumulated longitudinal displacement on the vehicle's driving trajectory, k represents the total time of recording, Dist represents the total length of the entire driving trajectory, and IRI represents the international roughness index; S73, calculating the road ride quality index, the specific formula is as follows: Among them, IRI represents the International Roughness Index, a0 and a1 represent constants, a0 is 0.026, a1 is 0.65, e is a natural constant, and RQI represents the road ride quality index.