High-speed road obstacle identification method based on cascade connection of Haar features and depth estimation features
By cascading Haar features and depth estimation features, combined with shadow segmentation and three-dimensional model reconstruction, the problem of low obstacle recognition accuracy in the traffic field is solved, and the adaptability to lighting and dynamic scenes is improved.
Patent Information
- Application Number
- CN202510750908.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-06-06
AI Technical Summary
Existing target detection algorithms in the traffic field have low recognition accuracy for small or severely occluded obstacles, and lack adaptability to lighting and dynamic scenes, resulting in inaccurate obstacle recognition and difficulty in depth estimation in texture-missing areas.
A method based on the cascade of Haar features and depth estimation features is adopted to reconstruct the three-dimensional model through shadow segmentation, feature extraction and fusion combined with GPS to improve the obstacle recognition accuracy.
It improves the recognition accuracy of road obstacle target detection and the target recognition capability of dynamic objects, and effectively solves the obstacle recognition problems under the influence of lighting and dynamic scenes.
Smart Images

Figure CN120689839A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of depth estimation and computer target detection, and in particular relates to a high-speed road obstacle recognition method based on the cascade of Haar features and depth estimation features. Background Art
[0002] With advances in science and technology, the application of computer-based object detection in the transportation sector has gradually expanded to encompass intelligent transportation systems, autonomous driving, and traffic monitoring. Research in this area has primarily focused on the accurate detection and recognition of objects such as pedestrians, vehicles, traffic signs, and traffic lights. The development of deep learning methods and the abundance of datasets have provided greater opportunities and challenges for object detection research in the transportation sector. With technological advancements and increasing application demands, object detection research in the transportation sector will continue to grow and evolve.
[0003] Existing object detection algorithms still exhibit certain errors in target positioning, particularly for small or heavily occluded targets, resulting in relatively low positioning accuracy. Object detection algorithms also perform poorly in complex scenarios, such as those with densely packed objects, overlapping objects, and occlusions. Accurately detecting and locating objects in these scenarios is difficult. Some object detection algorithms also perform poorly for small objects because their visual features are weaker and easily overwhelmed by background noise, leading to inaccurate detection results.
[0004] Existing object detection algorithms rarely consider the effects of lighting when applied in the transportation sector, resulting in the identification of object shadows as targets. They also have limited ability to track and estimate objects in dynamic scenes. Furthermore, they perform poorly at estimating depth in areas with missing or unclear textures, resulting in an inability to fully describe object attributes and categories, further leading to confusion in object identification and false detections. Summary of the Invention
[0005] In order to address the deficiencies in the prior art, the present invention provides a high-speed road obstacle recognition method based on the cascade of Haar features and depth estimation features, which effectively solves the problems of low obstacle recognition accuracy caused by illumination influence in road obstacle target detection and poor target recognition ability of dynamic objects.
[0006] In order to achieve the above object, the technical solution adopted by the present invention is: In a first aspect, a method for identifying obstacles on an expressway is provided, comprising: acquiring a road image and corresponding depth data from a vehicle traveling on the expressway; acquiring depth features of the road image from the depth data, and performing shadow segmentation on the acquired road image in combination with the depth features of the road image to obtain a segmented image; extracting Haar features and depth estimation features from the segmented image, and performing parallel matching processing; introducing a similarity-based attention mechanism to fuse the matching results of the Haar features and the depth estimation features to obtain a fused feature vector; performing threshold processing on the fused feature vector to obtain a depth estimation result, reconstructing a three-dimensional model of the obstacle based on the depth estimation result, and combining with GPS to realize the identification of obstacles on the expressway.
[0007] Furthermore, shadow segmentation is performed on the collected road image, including: estimating the scene lighting conditions with the help of depth data, establishing a Shadow Mapping model based on the described lighting conditions to determine the shadow content of the image, and analyzing the shadow information to distinguish between objects and shadows; the shadow area has different depth characteristics, and the shadow area is segmented by combining the depth information with the image for segmentation; the Shadow Mapping model uses the shadow map to determine whether the pixel is in the shadow during actual rendering; the current pixel is converted from the world coordinate system to the clipping space coordinate system of the light source perspective, and the corresponding depth value is obtained in the shadow map through the corresponding coordinates; then, the depth value of the current pixel is compared with the depth value in the shadow map. If the depth value of the current pixel is greater than the depth value in the shadow map, it means that the pixel is in the shadow.
[0008] Furthermore, the scene lighting conditions are estimated with the help of depth data, including: using a depth sensor to obtain a depth image of the scene; introducing Gaussian smoothing to smooth the normals, and introducing an adaptive Gaussian kernel to dynamically adjust Parameters to balance noise suppression and edge preservation; a Gaussian kernel is applied to each pixel position of the depth image: , in, is the standard deviation of the basic Gaussian filter; k is the proportional coefficient, which controls sensitivity to gradients; is the magnitude of the image gradient, calculated by the Sobel operator; is the standard deviation of the Gaussian kernel at the pixel position (u, v); For each pixel, the three-dimensional gradient and normalization of the depth image are calculated: , in, is the normal vector after unitization; D is the depth image; 、 are the rates of change of depth values in the x and y directions, respectively; Calculate the plane fitting normal, combine the RANSAC algorithm with the least squares method to fit the plane, eliminate outliers, and estimate the normal; , in, is the weight, the closer the distance, the greater the weight; a, b, c, d are plane parameters, the normal is (a, b, c); N is the total number of points involved in plane fitting, 、 is the two-dimensional pixel coordinate of the i-th point, is the depth value of the i-th point; Use multi-feature fusion method to detect shadow areas in images, using feature vector Fusion of depth features, normal features, intensity features and RGB color features; , in, is the normal vector The direction angle of is the intensity information of pixel i; 、 、 are the RGB channel color information of pixel i respectively; is the residual; is the combined feature vector used to describe the multimodal features of pixel i; , in, is the unit normal vector at pixel i; L is the illumination direction; k is an empirical coefficient used to adjust the intensity of illumination compensation; It is ambient light; In the optimization goal of illumination estimation, in addition to considering the pixels in the non-shadow area, the eigenvectors of the shadow area are also considered; the eigenvectors of the shadow area are compared with the eigenvectors of the non-shadow area to improve the accuracy of illumination estimation, thereby deriving the optimization goal of illumination estimation: , in, is the mean of the eigenvectors of the non-shaded region.
[0009] Furthermore, the depth estimation feature is extracted, including: using binocular vision depth estimation, using different levels of features to extract depth information; low-level features: converting the binocular image into a grayscale image, extracting the grayscale value of each pixel as a low-level feature, , in, is the grayscale value at the pixel point (x, y); 、 、 They are x 、 y 、 z Light direction; is the modulus of the light direction vector; Mid-level features: Binary encoding is performed on each pixel in the image and its neighborhood to extract texture information as features. Rotation-invariant CLBP with multiple window sizes of 3×3, 5×5, and 7×7 is used to form multi-scale texture features: , , in, is the local standard deviation, which is used to enhance the robustness to noise; is a collection of multi-scale rotation-invariant local binary pattern features, In the entire grayscale image Above, the CLBP features are calculated using a window size of s×s, where s is the window size of CLBP. It is the local binary code value at the pixel (x, y), reflecting the contrast relationship between the point and the neighboring pixels. is a binarization function, is the grayscale value of the center pixel (x,y), is the grayscale value of the domain pixel (u,v); High-level features: The shadow segmented image is divided into multiple regions, and features such as shape, texture, and color of each region are extracted to represent high-level information.
[0010] Furthermore, Haar features are extracted, including: first, performing HOI cutting on the image to remove non-road areas on both sides; using the expanded Haar features, by sliding on the image, the pixels covered by the white area minus the pixels covered by the black area are its feature response value; in the recognition of obstacles, multiple Haar features of different positions and scales are weighted averaged to capture multiple local features of the obstacle and obtain the final feature fusion response value; the feature fusion response value is input into the AdaBoost classifier to obtain the candidate area of the obstacle.
[0011] Furthermore, the matching results of Haar features and depth estimation features are fused, including: independently converting the extracted Haar feature subset and depth estimation feature subset into vector form; the Haar low-level feature vector is represented as: , in, It is the feature vector formed by transforming the Haar feature subset; is the first Haar low-level eigenvalue, is the Nth Haar low-level eigenvalue; In depth estimation, low-level features are represented by pixel value vectors. The pixel values of the image are arranged into a vector in a set order, and each pixel value is used as a dimension of the vector. The low-level feature vector of depth estimation is expressed as: , in, It is a vector composed of low-level pixel features extracted by the depth estimation task; is the eigenvalue of the first pixel position, is the eigenvalue of the L-th pixel position; The intermediate features are represented by statistical feature vectors, which use local binary patterns to extract texture information, and the LBP histogram of each local area is used as a dimension of the vector; High-level features are represented by region descriptions, and then the descriptors are combined into a vector; for each region segmented from the image, the area, perimeter, circularity, color histogram, and color moment region descriptors of each region are calculated to represent its features; Define the full set of Haar features , a complete set of deep features : , The extracted low-level, medium-level, and high-level Haar features and the hierarchical features of depth estimation are divided into multiple subsets; each subset represents a specific feature; Will Divide into K subsets, each subset corresponds to a local feature or semantic area: , Similarly, Divide into K subsets , The dimensions of the Haar feature vector and the depth estimation layer feature vector are matched. If the dimensions of the two feature vectors are inconsistent, the dimensions are padded by adding extra zero elements to the vector with smaller dimensions. Perform parallel feature matching calculations on each subset; use the depth estimation results to assist the depth estimation feature and Haar feature matching process, using depth-nearest neighbor matching; , in, is the i-th deep feature subset; is the i-th Haar feature subset; Assume that the depth value of the depth estimation feature subset is D1, and the depth value of the Haar feature subset is D2. , Where C is the confidence of the depth estimate; Use triangulation to calculate the distance between matching points, set a threshold, and filter out matching point pairs (A, B) with a closer distance; for depth estimation feature A, select the one with the smallest distance as the nearest neighbor match.
[0012] Furthermore, the distance between matching points is calculated, including: first receiving the pre-estimated lighting condition data to perform lighting compensation and calibrate the camera; then calculating the disparity; finally, converting the disparity to depth by first converting the disparity and image coordinates to coordinates in the camera coordinate system, and then further converting them to coordinates in the world coordinate system to obtain the depth value in the world coordinate system; after lighting compensation, the specific formula is as follows: , , Where B is the baseline distance; L is the focal length; d is the parallax; is the illumination compensation factor, is the average grayscale value of the input image.
[0013] Furthermore, the matching results of Haar features and depth estimation features are fused, and the following steps are also included: inputting the depth feature vector and its corresponding Haar eigenvector , the dimensions are all K; use similarity measurement method to calculate and The similarity or distance between them; get a similarity value sim; convert the similarity value sim into an attention weight, indicating and The attention between them; use the softmax function to convert the similarity value into a probability distribution; set the attention weight to attention; and Perform weighted parallel fusion, use the attention weight as the weight coefficient, and obtain the fused feature vector ; Connect each feature sub-vector to obtain the feature vector after feature fusion, , in, is the jth similarity value, is the feature vector after feature fusion.
[0014] Furthermore, a three-dimensional model of the obstacle is reconstructed based on the depth estimation result, including: obtaining a depth map, first mapping the depth value from the object coordinate system or the camera coordinate system to the normalized device coordinate system; then mapping the normalized device coordinate system to the pixel coordinate system according to the internal parameters of the camera, , The range of label values is , and the range of depth map pixel values is , for each pixel (x, y) in the depth map, the corresponding classifier output label is , is the depth map pixel value after mapping; Convert each pixel position in the depth image to a normalized image plane coordinate system, and use the depth value and the corresponding normalized image plane coordinates to generate three-dimensional point cloud data; map the depth value and the corresponding image coordinates to the three-dimensional coordinate system to obtain the three-dimensional position of each point; form a point cloud dataset with the three-dimensional coordinates of each point; connect the points in the point cloud into triangular patches to form a continuous three-dimensional mesh model, optimize the generated three-dimensional mesh model, and complete the three-dimensional reconstruction.
[0015] In a second aspect, a high-speed road obstacle recognition system is provided, comprising a storage medium and a processor; the storage medium is used to store instructions; the processor is used to operate according to the instructions to execute the high-speed road obstacle recognition method described in the first aspect.
[0016] Compared with the prior art, the present invention achieves the following beneficial effects: the present invention obtains road images and corresponding depth data from vehicles traveling on expressways; obtains depth features of the road images from the depth data, and performs shadow segmentation on the acquired road images in combination with the depth features of the road images to obtain a segmented image; extracts Haar features and depth estimation features from the segmented image and performs parallel matching processing; introduces a similarity-based attention mechanism to fuse the matching results of Haar features and depth estimation features to obtain a fused feature vector; performs threshold processing on the fused feature vector to obtain a depth estimation result, reconstructs a three-dimensional model of the obstacle based on the depth estimation result, and realizes the recognition of expressway obstacles in combination with GPS, thereby improving the recognition accuracy of road obstacle target detection and the target recognition capability of dynamic objects, effectively compensating for the shortcomings of Haar target detection being affected by lighting and the target recognition capability of dynamic objects. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 This is a schematic diagram of the main implementation process of a method for identifying obstacles on expressways based on the cascade of Haar features and depth estimation features provided by an embodiment of the present invention; Figure 2 Schematic diagram of the ROI region in an embodiment of the present invention; Figure 3 Schematic diagram of an improved Haar feature classifier in an embodiment of the present invention. DETAILED DESCRIPTION
[0018] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention.
[0019] Example 1 like Figure 1 As shown, a method for identifying obstacles on expressways based on the cascade of Haar features and depth estimation features includes: acquiring a road image and corresponding depth data from a vehicle traveling on the expressway; acquiring depth features of the road image from the depth data, and performing shadow segmentation on the acquired road image in combination with the depth features of the road image to obtain a segmented image; extracting Haar features and depth estimation features from the segmented image and performing parallel matching processing; introducing a similarity-based attention mechanism to fuse the matching results of the Haar features and the depth estimation features to obtain a fused feature vector; performing threshold processing on the fused feature vector to obtain a depth estimation result, reconstructing a three-dimensional model of the obstacle based on the depth estimation result, and combining with GPS to realize the identification of expressway obstacles.
[0020] Step 1: Collect the vehicle image dataset D and the depth dataset E from the depth sensor.
[0021] The vehicle image dataset D needs to distinguish between positive samples (with obstacles) and negative samples (without obstacles), and obstacles in the positive samples must be labeled. To improve the model's generalization and sensitivity to lighting conditions, the dataset should include different types of obstacles and samples under different lighting conditions. If the data volume is insufficient, the dataset can be expanded through data augmentation techniques such as image flipping, scaling, and cropping. A depth dataset E is collected, and filters are used to denoise the depth data to improve data quality. Since subsequent steps require binocular stereo vision for depth estimation, camera calibration is required to obtain the camera's intrinsic and extrinsic parameters to ensure the accuracy of the depth data.
[0022] The color images in the vehicle image dataset D are converted into grayscale images, and then the Sobel operator is used to calculate the gradient values in the horizontal and vertical directions of the grayscale images.
[0023] Horizontal gradient value , Vertical gradient value , in, arrive Represents the grayscale values of the eight adjacent pixels centered on the current pixel.
[0024] Use OpenCV to calculate the autocorrelation matrix M for each pixel based on the gradient value within the window: , in, is the grayscale change intensity of the pixel in the horizontal direction, is the grayscale change intensity of the pixel in the vertical direction, is the total intensity of horizontal edges within the window, is the total strength of the vertical edges within the window.
[0025] For each pixel, calculate the eigenvalue of the autocorrelation matrix and .
[0026] For each corner point, pair its pixel coordinates (x, y) in the image with its real world coordinates (X, Y, Z) on the calibration plate.
[0027] Calculate corner response function , in, represents the determinant of the autocorrelation matrix M, trace(M) represents the trace of the autocorrelation matrix M, and k is a constant (usually a small value, 0.04 - 0.06).
[0028] For each pixel, the value of its corner response function is compared with a pre-set threshold. If it is greater than the threshold, it is considered a corner point. After gradually adjusting the threshold from small to large, 0.6 is obtained as the threshold.
[0029] Step 2: Obtain the depth features of the road image from the depth data, and perform shadow segmentation on the collected road image in combination with the depth features of the road image to obtain a segmented image.
[0030] Using a depth dataset E, we estimate the scene lighting conditions. Based on these conditions, we build a Shadow Mapping model to determine the image's shadow content and analyze the shadow information to distinguish between objects and shadows. Shadow regions have distinct depth characteristics, and by combining this depth information with the image for segmentation, we achieve shadow segmentation.
[0031] Using the depth dataset E, the specific steps for estimating scene lighting conditions are as follows: Use a depth sensor (binocular stereo vision) to obtain a depth image of the scene.
[0032] In order to reduce the influence of noise and the discontinuity of normals, Gaussian smoothing is introduced to smooth the normals, and an adaptive Gaussian kernel is introduced to dynamically adjust Parameters to balance noise suppression and edge preservation. Apply a Gaussian kernel to each pixel position of the image: , in, is the standard deviation of the basic Gaussian filter (basic value); k is the proportional coefficient, control sensitivity to gradients; is the magnitude of the image gradient, calculated by the Sobel operator; is the standard deviation of the Gaussian kernel at pixel location (u,v).
[0033] For each pixel, the three-dimensional gradient and normalization of the depth image are calculated: , in, is the normal vector after unitization; D is the depth image; 、 They are the rates of change of depth values in the x and y directions respectively.
[0034] Calculate the plane fitting normal, combine the RANSAC algorithm with the least squares method to fit the plane, eliminate outliers, and estimate the normal; , in, is the weight, the closer the distance, the greater the weight; a, b, c, d are plane parameters, the normal is (a, b, c); N is the total number of points involved in plane fitting (i.e. the number of points in the point cloud or local neighborhood), 、 is the two-dimensional pixel coordinate of the i-th point (image coordinate system), is the depth value of the i-th point.
[0035] Using depth information and normal information, the shadow area in the image is detected, and then the lighting conditions are inferred from the area outside the shadow area. A multi-feature fusion method is used to detect the shadow area in the image. Using feature vectors Fusion of depth features, normal features, intensity features and RGB color features; , in, is the normal vector The direction angle of is the intensity information of pixel i; 、 、 are the RGB channel color information of pixel i respectively; is the residual; is the combined feature vector used to describe the multimodal features of pixel i; , in, is the unit normal vector at pixel i; L is the illumination direction; k is an empirical coefficient used to adjust the intensity of illumination compensation; It's ambient light.
[0036] In the optimization goal of illumination estimation, in addition to considering the pixels in the non-shadow area, the eigenvectors of the shadow area are also considered; the eigenvectors of the shadow area are compared with the eigenvectors of the non-shadow area to improve the accuracy of illumination estimation, thereby deriving the optimization goal of illumination estimation: , in, is the mean of the eigenvectors of the non-shaded region.
[0037] A Shadow Mapping model is established based on the described lighting conditions to determine the shadow content of the image.
[0038] First, you need to construct a view matrix from the position and direction of the light source and the projection matrix The scene is then rendered from the light's perspective, using a depth map to store the depth from the light source to each pixel in the scene. Next, for each vertex V in the scene, its clip space coordinates from the light's perspective are calculated: , and will And converted to NDC coordinates . From the perspective of the main police officer, the scene is rendered normally. When performing actual rendering, a shadow map is needed to determine whether a pixel is in the shadow. The current pixel is converted from the world coordinate system to the clipping space coordinate system of the light source's perspective, and the corresponding depth value is obtained in the shadow map through the corresponding coordinates. Then, the depth value of the current pixel (depth data set E) is compared with the depth value in the shadow map. If the depth value of the current pixel is greater than the depth value in the shadow map, it means that the pixel is in the shadow.
[0039] Step 3: Extract Haar features and depth estimation features from the segmented image and perform parallel matching processing.
[0040] Representative and discriminative features are selected from the extracted Haar feature set, including low-level features (edges and lines), mid-level features (image texture and local patterns), and high-level features (obstacle outlines). Haar features effectively capture key information in the image. Furthermore, considering the differences in representation and information content between the hierarchical features used for depth estimation and Haar features, the depth estimation hierarchical features and Haar features are processed separately through parallel matching. This ensures comprehensive and accurate feature extraction, providing a foundation for subsequent feature fusion.
[0041] The details of the Haar feature algorithm for extracting image features are as follows.
[0042] First, perform HOI cutting on the image to remove non-road areas such as buildings or railings on both sides, roadside trees, and the sky. Figure 2 As shown, the red box is the ROI area.
[0043] In order to extract the features of various obstacles, the present invention uses the extended Haar feature. By sliding on the image, the pixels covered by the white area minus the pixels covered by the black area are the feature response values.
[0044] In the recognition of obstacles, multiple Haar features at different positions and scales are weighted averaged to capture multiple local features of the obstacle and obtain the final feature fusion response value.
[0045] Because the image extracted is the shadow segmentation, the AdaBoost classifier is improved. The specific AdaBoost classifier is as follows Figure 3 shown.
[0046] The improved AdaBoost classifier utilizes a hierarchical cascade structure and a multi-feature fusion strategy. It first introduces a scene classifier (G1) to quickly distinguish the scene categories of the input image, specifically moving vehicles on highways, road signs and markings, and shadows of buildings or vehicles. An obstacle recognition classifier is added to the shadow classifier, and a region of interest (ROI) is output to narrow the scope of subsequent processing. Based on this, multiple cascaded feature classifiers (G2,...,Gn) are designed to perform fine-grained classification for different semantic targets. Each classifier selects the most appropriate features based on the target characteristics and is trained independently to avoid mutual interference between features, thereby achieving feature diversity and targeted classifier optimization.
[0047] In terms of classifier optimization, dynamic feature weight adjustment is adopted to dynamically change feature importance based on the previous classification results. The shadow classifier reduces the weight of color features and increases the weight of texture features to adapt to complex scene changes. At the same time, a classifier cascade pruning mechanism is introduced to prevent irrelevant areas from entering subsequent calculations by rejecting low-confidence areas early. In terms of decision-making mechanism, a multi-level feedback and fusion strategy is adopted. During forward propagation, the cascade decision process is processed step by step. The results of the subsequent classifiers can be fed back to the previous classifier for optimization. Finally, the outputs of each classifier are fused through weighted voting, and non-maximum suppression (NMS) is used to solve the problem of overlapping detection of multiple classifiers.
[0048] In the improved AdaBoost classifier, the process of extracting obstacle candidates from shadow regions is divided into two main stages. First, the scene classifier (G1) performs coarse-grained classification on the input image, determines the scene type of the image, and extracts the region of interest (ROI). Next, the shadow classifier (Gx) detects shadow regions in the ROI and generates a shadow region mask using features such as color and texture. Simultaneously, the obstacle classifier (Gn) combines features such as geometry and depth information to detect obstacle candidates. Finally, the shadow regions are fused with the obstacle candidates through region intersection analysis and confidence fusion to remove redundant detections. Non-maximum suppression (NMS) is then used to determine the final obstacle candidates.
[0049] Utilize binocular stereo vision depth estimation and use features at different levels to extract depth information.
[0050] Low-level features (grayscale value): Convert the binocular image into a grayscale image and extract the grayscale value of each pixel as a low-level feature. , in, is the grayscale value at the pixel point (x, y); 、 、 They are x 、 y 、 z Light direction; is the modulus of the light direction vector.
[0051] Mid-level features (texture information): Binary encoding is performed on each pixel in the image and its neighborhood to extract texture information as features. Rotationally invariant CLBP with multiple window sizes (3×3, 5×5, 7×7) is used to form multi-scale texture features: , , in, is the local standard deviation, which is used to enhance the robustness to noise; is a collection of multi-scale rotation-invariant local binary pattern features, In the entire grayscale image Above, the CLBP features are calculated using a window size of s×s, where s is the window size of CLBP. It is the local binary code value at the pixel (x, y), reflecting the contrast relationship between the point and the neighboring pixels. is a binarization function, is the grayscale value of the center pixel (x,y), is the grayscale value of the neighborhood pixel (u,v).
[0052] High-level features (obstacle shape, key points): Using the data from step 2, the shadow-segmented image is segmented into multiple regions, and features such as shape, texture, and color of each region are extracted to represent high-level information.
[0053] Step 4: Introduce a similarity-based attention mechanism to fuse the matching results of Haar features and depth estimation features to obtain the fused feature vector.
[0054] The attention mechanism can highlight important features and suppress irrelevant ones, thereby improving the effectiveness of feature fusion. The fused feature vectors are paired with corresponding labels to form a training dataset. To ensure scale consistency across different features, the training data is normalized. The training dataset is then divided into a training set and a validation set. The fused object detection model is trained using the training set, and the model's performance is evaluated using the validation set to ensure robustness and accuracy across different scenarios.
[0055] The distance of objects in the scene estimated by binocular stereo vision is as follows.
[0056] The extracted Haar feature subset and depth estimation feature subset are independently converted into vector form.
[0057] Haar low-level feature vector representation (intermediate and advanced are the same): , in, It is the feature vector formed by transforming the Haar feature subset; is the first Haar low-level eigenvalue, is the Nth Haar low-level eigenvalue; In depth estimation, low-level features are represented by pixel value vectors. The pixel values of the image are arranged into a vector in a set order, and each pixel value serves as a dimension of the vector.
[0058] The low-level feature vector of depth estimation is expressed as: , in, It is a vector composed of low-level pixel features extracted by the depth estimation task; is the eigenvalue of the first pixel position, is the eigenvalue of the L-th pixel position; The intermediate features are represented by statistical feature vectors, which use local binary patterns to extract texture information, and the LBP histogram of each local area is used as a dimension of the vector; , Where s() is a symbolic function, if If yes, it is 1, otherwise it is 0; is the coordinate offset of the ith pixel in the neighborhood relative to the center pixel.
[0059] High-level features are represented by region descriptions, and then the descriptors are combined into a vector; for each region segmented from the image, the area, perimeter, circularity, color histogram, and color moment region descriptors of each region are calculated to represent its features. , in is the number of pixels with LBP value i.
[0060] Define the full set of Haar features , a complete set of deep features : , The extracted low-level, medium-level, and high-level Haar features and the hierarchical features of depth estimation are divided into multiple subsets; each subset represents a specific feature; Will Divide into K subsets, each subset corresponds to a local feature or semantic area: , Similarly, Divide into K subsets , The dimensions of the Haar feature vector and the depth estimation layer feature vector are matched. If the dimensions of the two feature vectors are inconsistent, the dimensions are padded by adding extra zero elements to the vector with smaller dimensions. , if ,but: , if ,but: , Each subset is assigned to a multi-core CPU system for parallel feature matching calculations. Each CPU core independently performs calculations on its own subset. The matching process uses depth estimation results to assist in matching depth estimation features with Haar features, using depth-nearest neighbor matching. , in, is the i-th deep feature subset; is the i-th Haar feature subset.
[0061] According to the pixel coordinates of the matched feature point pairs on the image and the corresponding depth information, the depth estimation feature subset and its corresponding subset are verified by calculating the depth value difference of the matched feature point pairs. Assume that the depth value of the depth estimation feature subset is D1, and the depth value of the Haar feature subset is D2. , Where C is the confidence of the depth estimate.
[0062] Depth information provides the relative position relationship of objects, which is used to constrain the spatial range of the nearest neighbor search. Triangulation is used to calculate the distance between matching points, and a threshold is set to filter out matching point pairs (A, B) with a closer distance. , Setting thresholds , filter out matching point pairs with closer distances: , For the depth estimation feature A, calculate the Euclidean distance between it and all candidate Haar feature vectors B, and select the one with the smallest distance as the nearest neighbor match.
[0063] , in, is the feature vector of depth estimation feature A; is the eigenvector of the candidate Haar eigenvector B.
[0064] Binocular stereo depth estimation uses pre-estimated lighting conditions to compensate for the illumination. The left and right cameras are calibrated. Parallax is then calculated: parallax = pixel coordinates in the left image minus the corresponding pixel coordinates in the right image. Finally, parallax is converted to depth. The parallax and image coordinates are first converted to coordinates in the camera coordinate system, and then further converted to coordinates in the world coordinate system to obtain the depth value in the world coordinate system. After illumination compensation, the specific formula is as follows: , , Among them, B is the baseline distance; L is the focal length; d is the parallax; H is the empirical coefficient used to adjust the intensity of illumination compensation. is the illumination compensation factor, is the average grayscale value of the input image.
[0065] Step 5: Threshold processing is performed on the fused feature vector to obtain the depth estimation result. The three-dimensional model of the obstacle is reconstructed based on the depth estimation result, and the GPS is combined to realize the identification of the expressway obstacles.
[0066] Perform threshold processing and calculate the Huber loss function to optimize the robustness and stability of the model. The Huber loss function is a loss function that combines the advantages of mean square error and absolute error, and shows good robustness in complex environments. Set to h⋅std (standard deviation), where h is an empirical adjustment factor that can be adjusted based on actual conditions. Finally, the obstacle is reconstructed into a 3D model based on the depth estimation results. This 3D reconstruction accurately describes and locates the location and specific information of obstacles on the expressway, providing reliable data support for subsequent decision-making and processing.
[0067] Input depth feature vector and its corresponding Haar eigenvector , the dimensions are all K; use similarity measurement method to calculate and The similarity or distance between them; get a similarity value sim; convert the similarity value sim into an attention weight, indicating and The attention between them; use the softmax function to convert the similarity value into a probability distribution; set the attention weight to attention; and Perform weighted parallel fusion, use the attention weight as the weight coefficient, and obtain the fused feature vector ; Connect each feature sub-vector to obtain the feature vector after feature fusion, , in, is the jth similarity value, is the feature vector after feature fusion.
[0068] Obtain a depth map. First, map the depth value from the object coordinate system (or camera coordinate system) to the normalized device coordinate system. Then, based on the camera's internal parameters, map the normalized device coordinate system to the pixel coordinate system. Orthogonal projection is used throughout to reduce computational complexity. Map the classification results to the depth map. Because the output is a label, label mapping is used as the mapping method.
[0069] , The range of label values is , and the range of depth map pixel values is , for each pixel (x, y) in the depth map, the corresponding classifier output label is , is the depth map pixel value after mapping.
[0070] Based on the depth estimation result after threshold processing, the three-dimensional model of the obstacle is reconstructed.
[0071] Convert each pixel position in the depth image to the normalized image plane coordinate system, , , Where (x, y) is the pixel coordinate in the depth image, are the principal point coordinates of the image, and are the horizontal and vertical focal lengths of the camera.
[0072] The depth values and the corresponding normalized image plane coordinates are used to generate 3D point cloud data. The depth values and the corresponding image coordinates are mapped into a 3D coordinate system to obtain the 3D position of each point.
[0073] , , , The three-dimensional coordinates (X, Y, Z) of each point are combined into a point cloud dataset. Because the depth estimation results have been thresholded, the operation of filtering to remove outliers is omitted.
[0074] Connect the points in the point cloud into triangular patches to form a continuous 3D mesh model. The generated 3D mesh model is optimized, including smoothing, to complete the 3D reconstruction.
[0075] Finally, with the help of the GPS system, the location and specific information of obstacles on the expressway are located, including the category and geometric data of the obstacles.
[0076] Example 2 Based on the high-speed road obstacle recognition method based on the cascade of Haar features and depth estimation features described in Example 1, this embodiment provides a high-speed road obstacle recognition system based on the cascade of Haar features and depth estimation features, including a storage medium and a processor; the storage medium is used to store instructions; the processor is used to operate according to the instructions to execute the high-speed road obstacle recognition method based on the cascade of Haar features and depth estimation features described in Example 1.
[0077] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A method for identifying obstacles on expressways, characterized in that: include: Acquire road images and corresponding depth data from vehicles traveling on the expressway; Acquire the depth features of the road image from the depth data, and perform shadow segmentation on the acquired road image in combination with the depth features of the road image to obtain a segmented image; Extract Haar features and depth estimation features from the segmented image and perform parallel matching; A similarity-based attention mechanism is introduced to fuse the matching results of Haar features and depth estimation features to obtain the fused feature vector; The fused feature vector is thresholded to obtain the depth estimation result. The three-dimensional model of the obstacle is reconstructed based on the depth estimation result, and the GPS is combined to realize the identification of expressway obstacles.
2. The method for identifying obstacles on expressways according to claim 1, wherein: Perform shadow segmentation on the collected road image, including: Using depth data to estimate scene lighting conditions, a Shadow Mapping model is built based on the described lighting conditions to determine the shadow content of the image and analyze shadow information to distinguish between objects and shadows. Shadow areas have different depth characteristics, and by combining depth information with image segmentation, shadow area segmentation is achieved. When the Shadow Mapping model performs actual rendering, it uses a shadow map to determine whether a pixel is in the shadow; the current pixel is converted from the world coordinate system to the clipping space coordinate system of the light source's perspective, and the corresponding depth value is obtained in the shadow map through the corresponding coordinates; then, the depth value of the current pixel is compared with the depth value in the shadow map. If the depth value of the current pixel is greater than the depth value in the shadow map, it means that the pixel is in the shadow.
3. The method for identifying obstacles on expressways according to claim 2, wherein: Estimate scene lighting conditions using depth data, including: Use a depth sensor to obtain a depth image of the scene; Introduce Gaussian smoothing to smooth the normals, and introduce adaptive Gaussian kernel to dynamically adjust Parameters to balance noise suppression and edge preservation; a Gaussian kernel is applied to each pixel position of the depth image: , in, is the standard deviation of the basic Gaussian filter; k is the proportional coefficient, which controls sensitivity to gradients; is the magnitude of the image gradient, calculated by the Sobel operator; is the standard deviation of the Gaussian kernel at the pixel position (u, v); For each pixel, the three-dimensional gradient and normalization of the depth image are calculated: , in, is the normal vector after unitization; D is the depth image; 、 are the rates of change of depth values in the x and y directions, respectively; Calculate the plane fitting normal, combine the RANSAC algorithm with the least squares method to fit the plane, eliminate outliers, and estimate the normal; , in, is the weight, the closer the distance, the greater the weight; a, b, c, d are plane parameters, the normal is (a, b, c); N is the total number of points involved in plane fitting, 、 is the two-dimensional pixel coordinate of the i-th point, is the depth value of the i-th point; Use multi-feature fusion method to detect shadow areas in images, using feature vector Fusion of depth features, normal features, intensity features and RGB color features; , in, is the normal vector The direction angle of is the intensity information of pixel i; 、 、 are the RGB channel color information of pixel i respectively; is the residual; is the combined feature vector used to describe the multimodal features of pixel i; , in, is the unit normal vector at pixel i; L is the illumination direction; k is an empirical coefficient used to adjust the intensity of illumination compensation; It is ambient light; In the optimization goal of illumination estimation, in addition to considering the pixels in the non-shadow area, the eigenvectors of the shadow area are also considered; the eigenvectors of the shadow area are compared with the eigenvectors of the non-shadow area to improve the accuracy of illumination estimation, thereby deriving the optimization goal of illumination estimation: , in, is the mean of the eigenvectors of the non-shaded region.
4. The method for identifying obstacles on expressways according to claim 3, characterized in that: Extract depth estimation features, including: Utilize binocular vision depth estimation and use features at different levels to extract depth information; Low-level features: Convert the binocular image into a grayscale image and extract the grayscale value of each pixel as a low-level feature. , in, is the grayscale value at the pixel point (x, y); 、 、 They are x 、 y 、 z Light direction; is the modulus of the light direction vector; Mid-level features: Binary encoding is performed on each pixel in the image and its neighborhood to extract texture information as features. Rotation-invariant CLBP with multiple window sizes of 3×3, 5×5, and 7×7 is used to form multi-scale texture features: , , in, is the local standard deviation, which is used to enhance the robustness to noise; is a collection of multi-scale rotation-invariant local binary pattern features, In the entire grayscale image Above, the CLBP features are calculated using a window size of s×s, where s is the window size of CLBP. It is the local binary code value at the pixel (x, y), reflecting the contrast relationship between the point and the neighboring pixels. is a binarization function, is the grayscale value of the center pixel (x,y), is the grayscale value of the domain pixel (u,v); High-level features: The shadow segmented image is divided into multiple regions, and features such as shape, texture, and color of each region are extracted to represent high-level information.
5. The method for identifying obstacles on expressways according to claim 4, characterized in that: Extract Haar features, including: First, perform HOI cutting on the image to remove the non-road areas on both sides; Using the extended Haar feature, by sliding on the image, the pixels covered by the white area minus the pixels covered by the black area are its feature response value; In the recognition of obstacles, multiple Haar features at different positions and scales are weighted averaged to capture multiple local features of the obstacle and obtain the final feature fusion response value; The feature fusion response value is input into the AdaBoost classifier to obtain the candidate area of the obstacle.
6. The method for identifying obstacles on expressways according to claim 5, characterized in that: The matching results of Haar features and depth estimation features are fused, including: The extracted Haar feature subset and the depth estimation feature subset are independently converted into vector form; The Haar low-level feature vector is expressed as: , in, It is the feature vector formed by transforming the Haar feature subset; is the first Haar low-level eigenvalue, is the Nth Haar low-level eigenvalue; In depth estimation, low-level features are represented by pixel value vectors. The pixel values of the image are arranged into a vector in a set order, and each pixel value is used as a dimension of the vector. The low-level feature vector of depth estimation is expressed as: , in, It is a vector composed of low-level pixel features extracted by the depth estimation task; is the eigenvalue of the first pixel position, is the eigenvalue of the L-th pixel position; The intermediate features are represented by statistical feature vectors, which use local binary patterns to extract texture information, and the LBP histogram of each local area is used as a dimension of the vector; High-level features are represented by region descriptions, and then the descriptors are combined into a vector; for each region segmented from the image, the area, perimeter, circularity, color histogram, and color moment region descriptors of each region are calculated to represent its features; Define the full set of Haar features , a complete set of deep features : , The extracted low-level, medium-level, and high-level Haar features and the hierarchical features of depth estimation are divided into multiple subsets; each subset represents a specific feature; Will Divide into K subsets, each subset corresponds to a local feature or semantic area: , Similarly, Divide into K subsets , The dimensions of the Haar feature vector and the depth estimation layer feature vector are matched. If the dimensions of the two feature vectors are inconsistent, the dimensions are padded by adding extra zero elements to the vector with smaller dimensions. Perform parallel feature matching calculations on each subset; use the depth estimation results to assist the depth estimation feature and Haar feature matching process, using depth-nearest neighbor matching; , in, is the i-th deep feature subset; is the i-th Haar feature subset; Assume that the depth value of the depth estimation feature subset is D1, and the depth value of the Haar feature subset is D2. , Where C is the confidence of the depth estimate; Use triangulation to calculate the distance between matching points, set a threshold, and filter out matching point pairs (A, B) with a closer distance; for depth estimation feature A, select the one with the smallest distance as the nearest neighbor match.
7. The method for identifying obstacles on expressways according to claim 6, characterized in that: Calculate the distance between matching points, including: First, we receive the estimated lighting conditions data to perform lighting compensation and calibrate the camera. Then, we calculate the parallax. Finally, we convert the parallax to depth by first converting the parallax and image coordinates to the camera coordinate system, and then further converting them to the world coordinate system to obtain the depth value in the world coordinate system. After lighting compensation, the specific formula is as follows: , , Where B is the baseline distance; L is the focal length; d is the parallax; is the illumination compensation factor, is the average grayscale value of the input image.
8. The method for identifying obstacles on expressways according to claim 7, characterized in that: The matching results of Haar features and depth estimation features are fused, and also include: Input depth feature vector and its corresponding Haar eigenvector , the dimensions are all K; use similarity measurement method to calculate and The similarity or distance between them; get a similarity value sim; convert the similarity value sim into an attention weight, indicating and The attention between them; use the softmax function to convert the similarity value into a probability distribution; set the attention weight to attention; and Perform weighted parallel fusion, use the attention weight as the weight coefficient, and obtain the fused feature vector ; Connect each feature sub-vector to obtain the feature vector after feature fusion, , in, is the jth similarity value, is the feature vector after feature fusion.
9. The method for identifying obstacles on expressways according to claim 8, characterized in that: Reconstruct the 3D model of the obstacle based on the depth estimation results, including: To obtain the depth map, first map the depth value from the object coordinate system or camera coordinate system to the normalized device coordinate system; then map the normalized device coordinate system to the pixel coordinate system based on the internal parameters of the camera. , The range of label values is , and the range of depth map pixel values is , for each pixel (x, y) in the depth map, the corresponding classifier output label is , is the depth map pixel value after mapping; Convert each pixel position in the depth image to a normalized image plane coordinate system, and use the depth value and the corresponding normalized image plane coordinates to generate three-dimensional point cloud data; map the depth value and the corresponding image coordinates to the three-dimensional coordinate system to obtain the three-dimensional position of each point; form a point cloud dataset with the three-dimensional coordinates of each point; connect the points in the point cloud into triangular patches to form a continuous three-dimensional mesh model, optimize the generated three-dimensional mesh model, and complete the three-dimensional reconstruction.
10. A high-speed road obstacle recognition system, characterized in that: including storage media and processors; The storage medium is used to store instructions; The processor is configured to operate according to the instruction to execute the expressway obstacle identification method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Wear surface damage depth estimation method and system based on multi-attention mechanism
CN114972882A
Damage degree identification method, system and equipment based on image data and medium
CN119445390A
Road obstacle recognition method based on machine vision
CN119478946A
Plastic waste detection method, device and equipment and storage medium
CN119810426A
Vehicle automatic driving method
CN119828678A
Cited By
Power transmission line external damage hidden danger early warning method based on visual device
CN121053547A
Obstacle scanning perception system for visual impairment aided navigation
CN121617073A
Obstacle scanning perception system for visually impaired assisted navigation
CN121617073B