Three-dimensional distance measurement method based on YOLOv8s-SPD and SGBM

Through the improved YOLOv8s-SPD network model and SGBM algorithm, the problem of missed detection of small targets in complex road environments is solved, the detection accuracy and computational efficiency are improved, and it is suitable for three-dimensional ranging in autonomous driving technology.

CN120765748APending Publication Date: 2025-10-10SHENZHEN POLYTECHNIC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510934813.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Existing target detection algorithms are prone to missing small targets in complex road environments, and have low computational efficiency and accuracy. The downsampling operation of traditional convolutional neural networks leads to the loss of small target features, and cross-view three-dimensional ranging technology faces challenges in computational efficiency and accuracy.

Method used

An improved YOLOv8s-SPD network model is adopted in combination with the spatial depth conversion convolution module (SPD-Conv) and the SGBM stereo matching algorithm. The spatial depth conversion convolution module (SPD-Conv) is introduced into the backbone and neck networks to enhance the model's detection ability for small targets, and the SGBM algorithm is used to calculate the depth information of the disparity map.

Benefits of technology

It significantly improves the accuracy and real-time performance of target detection and three-dimensional ranging in complex road environments, and can effectively cope with challenges such as small target detection, dynamic occlusion and lighting changes, providing strong support for the development of autonomous driving technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120765748A_ABST
    Figure CN120765748A_ABST
Patent Text Reader

Abstract

The invention discloses a three-dimensional distance measurement method based on YOLOv8s-SPD and SGBM, and relates to the three-dimensional distance measurement method based on the YOLOv8s-SPD and the SGBM. The objective of the invention is to solve the problems of small target leak detection and low calculation efficiency and precision caused by detection of targets such as pedestrians and vehicles in a complex road environment by using an existing distance measurement method. The method comprises the following steps: 1, acquiring a data set; 2, a YOLOv8s-SPD network model is constructed; 3, obtaining a trained network model; 4, the binocular camera obtains a left view and a right view; inputting the left view into the trained network model, and outputting a target category and a two-dimensional bounding box coordinate in the left view by the model; disparity maps of the left view and the right view are obtained based on an SGBM algorithm; generating a depth map; and calculating the distance of the target object based on the depth map, the target category in the left view output by the model and the two-dimensional bounding box coordinates. The method is applied to the field of three-dimensional distance measurement.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to a three-dimensional ranging method based on YOLOv8s-SPD and SGBM. BACKGROUND

[0002] With the rapid development of automatic driving technology, higher requirements are put forward for the accuracy and real-time performance of target detection and three-dimensional ranging in complex road environments. Traditional target detection algorithms (such as the YOLO series) have significant limitations in dealing with small targets, dynamic occlusion and dramatic changes in illumination. Specifically, the down-sampling operation of the conventional convolutional neural network (CNN) easily leads to the loss of small target features, resulting in missed detection. In addition, the three-dimensional ranging technology across different angles also faces challenges in terms of computational efficiency and accuracy. SUMMARY

[0003] The purpose of the application is to solve the problem that the existing ranging method is prone to missing detection of small targets such as pedestrians and vehicles in complex road environments, and to improve the computational efficiency and accuracy. Therefore, a three-dimensional ranging method based on YOLOv8s-SPD and SGBM is proposed.

[0004] The specific process of a three-dimensional ranging method based on YOLOv8s-SPD and SGBM is as follows:

[0005] Step one, obtain the data set; the specific process is as follows:

[0006] The data set is the VOC2007 data set; the VOC2007 data set is used as the training set;

[0007] Step two, build a YOLOv8s-SPD network model;

[0008] The YOLOv8s-SPD network model includes a backbone network Backbone, a neck network Neck and a detection head;

[0009] The working process of the YOLOv8s-SPD network model is as follows:

[0010] The picture is sequentially input into the backbone network Backbone, the neck network Neck and the detection head, and the detection head outputs the target class and the two-dimensional bounding box coordinates;

[0011] Step three, input the training set into the YOLOv8s-SPD network model, train the YOLOv8s-SPD network model, and obtain the trained improved YOLOv8s network model;

[0012] Step four, the binocular camera obtains the left view and the right view;

[0013] The left view is input into the trained YOLOv8s-SPD network model, and the trained YOLOv8s-SPD network model outputs the target category and two-dimensional bounding box coordinates in the left view;

[0014] Obtain the disparity map of the left view and the right view based on the SGBM algorithm;

[0015] According to the disparity map and the known binocular camera parameters, the depth value corresponding to each pixel in the disparity map is calculated to generate a depth map;

[0016] The distance of the target object is calculated based on the depth map, the target category and the two-dimensional bounding box coordinates in the left view output by the trained YOLOv8s-SPD network model.

[0017] The beneficial effects of the present invention are:

[0018] This paper proposes a two-dimensional target detection and SGBM three-dimensional ranging technology based on improved YOLOv8s to improve the detection capability of pedestrians, vehicles and other targets in complex road environments.

[0019] The present invention uses the VOC2007 dataset as the experimental data source. In terms of model architecture, the present invention improves the YOLOv8s network. Specifically, the space-to-depth conversion convolution module (Space-to-Depth Convolution, SPD-Conv) is combined in the backbone network Backbone and the neck network Neck of YOLOv8s. The space-to-depth conversion convolution module (SPD-Conv) consists of a space-to-depth (SPD) layer and a non-strided convolution non-strided layer, which can increase the number of channels of the feature map while maintaining spatial resolution, thereby enhancing the model's ability to detect small targets. In addition, by introducing more contextual information into the network, the model is made more robust when dealing with complex scenes such as dynamic occlusion and lighting changes.

[0020] In terms of three-dimensional distance measurement, the present invention adopts the SGBM (Semi-Global Block Matching) stereo matching algorithm.

[0021] The SGBM algorithm is a technique for calculating depth images in global matching stereo vision. It estimates the disparity between the left and right images by matching pixel blocks, and then calculates the depth information of the object. The advantage of the SGBM algorithm is that it exploits the correlation between pixel blocks and considers the information of surrounding pixels when calculating the disparity. Therefore, it can obtain a relatively accurate depth image and effectively cope with the complex conditions in real scenes. In addition, the SGBM algorithm adopts several optimization strategies to achieve a certain degree of improvement in calculation speed. This method improves computational efficiency and accuracy.

[0022] This invention significantly improves the accuracy and real-time performance of target detection and three-dimensional ranging in complex road environments. It can effectively cope with challenges such as small target detection, dynamic occlusion and lighting changes, and provides strong support for the development of autonomous driving technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 This is a flow chart of the two-dimensional detection and SGBM three-dimensional ranging technology method of the method of the present invention;

[0024] Figure 2 This is a schematic diagram of the network structure of the YOLOv8-SPD two-dimensional detection and SGBM three-dimensional ranging technology of the method of the present invention;

[0025] Figure 3 This is a schematic diagram of the principle of the space-to-depth transformation convolution module (SPD-Conv) in the method of the present invention, where Space-to-Depth Transformation is the space-to-depth transformation module and Non-step Convolution (Conv) is the non-strided convolution.

[0026] Figure 4a This is a comparison chart of mAP50 (B) of the method of the present invention, YOLOv8s-SPD and YOLOv8s;

[0027] Figure 4b Comparison chart of mAP50-95 (B) between YOLOv8s-SPD and YOLOv8s of the present invention;

[0028] Figure 5a Metrics / Precision(B) of YOLOv8s-SPD on the training set, where Metrics / Precision(B) is the accuracy indicator;

[0029] Figure 5b Metrics / recall(B) is the YOLOv8s-SPD on the training set, and Metrics / recall(B) is the recall rate indicator;

[0030] Figure 5c Metrics / mAP50(B) is the average precision of YOLOv8s-SPD on the training set. Metrics / mAP50(B) is the average precision under the condition of IoU (intersection over union) of 0.5.

[0031] Figure 5dMetrics / mAP50-95(B) of YOLOv8s-SPD on the training set, MetricsmAP50-95(B) is the average of mAP calculated at multiple IoU thresholds;

[0032] Figure 6a This is a performance indicator curve of val / box_loss of YOLOv8s-SPD on the validation set in the method of the present invention, where val / box_loss represents the loss function of the validation set;

[0033] Figure 6b This is a performance indicator curve of val / cls_loss of YOLOv8s-SPD on the validation set in the method of the present invention, where val / cls_loss represents the mean classification loss of the validation set;

[0034] Figure 6c val / dfl_loss is a performance indicator curve of YOLOv8s-SPD on the validation set in the method of the present invention, where val / dfl_loss represents the mean loss of the validation set;

[0035] Figure 7a is the original dataset Datesets;

[0036] Figure 7b This is a visualization of the detection results of YOLOv8s on the test set;

[0037] Figure 7c This is a visualization of the detection results of YOLOv8s-SPD on the test set in the inventive method;

[0038] Figure 8a This is a visualization of the detection results of YOLOv8s on the YOLOv8-SGBM algorithm;

[0039] Figure 8b This is a visualization of the detection results of YOLOv8s-SPD on the YOLOv8-SGBM algorithm in the method of the present invention. DETAILED DESCRIPTION

[0040] Specific implementation method 1: This implementation method is a three-dimensional ranging method based on YOLOv8s-SPD and SGBM. The specific process is as follows:

[0041] Step 1: Get the data set; the specific process is:

[0042] The dataset is the VOC2007 dataset; the VOC2007 dataset is used as the training set;

[0043] Step 2: Build the YOLOv8s-SPD network model;

[0044] The YOLOv8s-SPD network model includes the backbone network, the neck network, and the detection head;

[0045] The working process of the YOLOv8s-SPD network model is as follows:

[0046] The image is sequentially input into the backbone network, the neck network, and the detection head. The detection head outputs the target category and the two-dimensional bounding box coordinates.

[0047] Improvements to the YOLOv8s network architecture have been made, incorporating a spatial-to-depth convolution (SPD-Conv) module into the backbone and neck networks. This module, consisting of a space-to-depth (SPD) layer and a non-strided convolution (Conv) layer, can be applied to most CNN architectures. This module addresses challenges such as small object detection, dynamic occlusion, and illumination changes in traffic scenarios, effectively improving autonomous vehicles' perception of complex road environments.

[0048] Step 3: Input the training set into the YOLOv8s-SPD network model, train the YOLOv8s-SPD network model, and obtain a trained YOLOv8s-SPD network model; build a deep learning model suitable for binocular vision system detection in road scenes;

[0049] Step 4: The binocular camera obtains the left view and the right view;

[0050] The left view is input into the trained YOLOv8s-SPD network model, which outputs the target category and two-dimensional bounding box coordinates in the left view (for example, outputting two targets A and B in an image, as well as the center point position xy, width w, and height h of each target); targets such as vehicles and pedestrians;

[0051] Obtain the disparity map of the left view and the right view based on the SGBM algorithm;

[0052] According to the disparity map and the known binocular camera parameters, the depth value corresponding to each pixel in the disparity map is calculated to generate a depth map;

[0053] Based on the depth map, the target category and two-dimensional bounding box coordinates in the left view output by the trained YOLOv8s-SPD network model, the distance of the target object (the distance between targets A and B) is calculated.

[0054] Specific implementation method two: The difference between this implementation method and specific implementation method one is that the backbone network Backbone includes in sequence: input layer, first convolution layer, first spatial depth conversion convolution module SPD-Conv, first C2f module, second spatial depth conversion convolution module SPD-Conv, second C2f module, third spatial depth conversion convolution module SPD-Conv, third C2f module, fourth spatial depth conversion convolution module SPD-Conv, fourth C2f module, SPPF module.

[0055] Other steps and parameters are the same as those in the first embodiment.

[0056] Specific embodiment 3: This embodiment differs from specific embodiment 1 or 2 in that the working process of the backbone network Backbone in step 2 is:

[0057] The image is input into the first convolutional layer through the input layer, and the first convolutional layer outputs the feature map A;

[0058] Feature map A is sequentially input into the first spatial depth conversion convolution module and the first C2f module, and the first C2f module outputs feature map B;

[0059] Feature map B is sequentially input into the second spatial depth conversion convolution module and the second C2f module, and the second C2f module outputs feature map C;

[0060] The feature map C is sequentially input into the third spatial depth conversion convolution module and the third C2f module, and the third C2f module outputs the feature map D;

[0061] The feature map D is sequentially input into the fourth spatial depth conversion convolution module and the fourth C2f module, and the fourth C2f module outputs the feature map E;

[0062] The feature map E is input into the SPPF module, and the SPPF module outputs the feature map F.

[0063] Other steps and parameters are the same as those in the first or second embodiment.

[0064] Specific embodiment 4: This embodiment differs from any one of specific embodiments 1 to 3 in that the working process of the first spatial depth conversion convolution module is:

[0065] The size of feature map A is S×S×C; the first S is the length of feature map A, the second S is the width of feature map A, and C is the number of channels of feature map A;

[0066] Divide the feature map A into 4 feature maps, namely feature map A1, feature map A2, feature map A3, and feature map A4;

[0067] The size of feature map A1 is

[0068] The size of the feature map A2 is

[0069] The size of the feature map A3 is

[0070] The size of the feature map A4 is

[0071] The feature map A1, the feature map A2, the feature map A3 and the feature map A4 are spliced to obtain a feature map A', and the size of the feature map A' is

[0072] The feature map A' is input into a second convolutional layer, and the second convolutional layer outputs a feature map A'', and the size of the feature map A' is

[0073] The working process of each spatial-depth conversion convolutional module in the second spatial-depth conversion convolutional module, the third spatial-depth conversion convolutional module and the fourth spatial-depth conversion convolutional module is the same as that of the first spatial-depth conversion convolutional module.

[0074] The other steps and parameters are the same as those in one of the first to third embodiments.

[0075] Embodiment five: the difference between this embodiment and one of the first to fourth embodiments is that the neck network Neck in the step two comprises:

[0076] The first up-sampling layer, the first splicing layer, the fifth C2f, the second sampling layer, the second splicing layer, the sixth C2f, the third convolutional layer, the third splicing layer, the seventh C2f, the fourth convolutional layer, the fourth splicing layer and the eighth C2f.

[0077] The other steps and parameters are the same as those in one of the first to fourth embodiments.

[0078] Embodiment six: the difference between this embodiment and one of the first to fifth embodiments is that the working process of the neck network Neck is:

[0079] The feature map F output by the SPPF module is input into the first up-sampling layer, and the first up-sampling layer outputs a feature map G;

[0080] The feature map G and the feature map D output by the third C2f module are input into the first splicing layer for splicing, and the first splicing layer outputs a feature map H;

[0081] The feature map H is input into the fifth C2f, and the fifth C2f outputs a feature map I;

[0082] The feature map I is input into the second up-sampling layer, and the second up-sampling layer outputs a feature map J;

[0083] Feature map J and the second C2f module output feature map C are input into the second splicing layer for splicing, and the second splicing layer outputs feature map K;

[0084] The feature map K is input into the sixth C2f, and the sixth C2f outputs the feature map L;

[0085] The feature map L is input into the third convolutional layer, and the third convolutional layer outputs the feature map M;

[0086] The feature map M and the fifth C2f output feature map I are input into the third splicing layer, and the third splicing layer outputs the feature map N;

[0087] The feature map N is input into the seventh C2f, and the seventh C2f outputs the feature map O;

[0088] The feature map O is input into the fourth convolutional layer, and the fourth convolutional layer outputs the feature map P;

[0089] The feature map P and SPPF module output feature map F are input into the fourth splicing layer, and the fourth splicing layer outputs feature map Q;

[0090] The feature map Q is input into the eighth C2f, and the eighth C2f outputs the feature map R.

[0091] Other steps and parameters are the same as those in Specific Implementations 1 to 5-1.

[0092] Specific embodiment seven: This embodiment differs from any one of specific embodiments one to six in that: the detection head in step two includes a first detection head, a second detection head, and a third detection head;

[0093] The sixth C2f outputs a feature map L as the output result of the first detection head;

[0094] The seventh C2f outputs a feature map O as the output result of the second detection head;

[0095] The eighth C2f outputs the feature map R as the output result of the third detection head.

[0096] The other steps and parameters are the same as those in the first to sixth embodiments.

[0097] Specific embodiment eight: This embodiment differs from specific embodiments one to seven in that the VOC2007 dataset contains 20 categories, namely:

[0098] Airplane, bicycle, bird, boat, bottle, bus, car, cat, chair, cow, dining table, dog, horse, motorcycle, person, potted plant, sheep, sofa, train, TV.

[0099] Other steps and parameters are the same as those in Specific Embodiments 1 to 7-1.

[0100] Specific embodiment nine: This embodiment differs from any one of specific embodiments one to eight in that: in step three, the training set is input into the YOLOv8s-SPD network model (YOLOv8s-SPD), the YOLOv8s-SPD network model is trained to obtain a trained YOLOv8s-SPD network model; and a deep learning model suitable for binocular vision system detection in road scenes is constructed;

[0101] The specific process is:

[0102] The image size in the training set is 640×640;

[0103] During training, the input data size for each batch is 32, the number of training rounds is 300, and the learning rate is 0.01.

[0104] The other steps and parameters are the same as those in the specific implementation modes 1 to 8-1.

[0105] The following examples are used to verify the beneficial effects of the present invention:

[0106] Example 1:

[0107] The flowchart of two-dimensional detection and SGBM three-dimensional ranging technology based on YOLOv8s-SPD is as follows Figure 1 As shown, the improved YOLOv8s-SPD network structure is as follows Figure 2 As shown, the following steps are included:

[0108] Step 1: This paper uses the VOC2007 dataset as the experimental data source. The VOC2007 dataset is widely used for object detection and image segmentation tasks. It contains images from 20 categories and corresponding annotations, which can help the model effectively cope with the diverse conditions found on real roads. To effectively train, validate, and test the model, the entire dataset is divided into three parts: training, validation, and test sets in an 8:1:1 ratio. This division ensures sufficient training data while retaining sufficient test and validation samples to verify the effectiveness of the model.

[0109] The training set contains 17,202 images, representing approximately 80% of the entire dataset. These images are used for model learning and parameter optimization, including feature extraction, model training, and hyperparameter adjustment. Through the training set, the model learns from images how to distinguish between normal roads and roadside pedestrians and cars, understands the morphological characteristics of objects, and learns how to accurately detect objects.

[0110] Validation set: This set contains 2,150 images, representing approximately 10% of the entire dataset. This set is primarily used for hyperparameter tuning and overfitting monitoring during model training. By dynamically adjusting model training strategies (such as learning rate decay or early stopping) based on performance on the validation set, we ensure continuous optimization of feature extraction capabilities during the iterative process. Performance evaluation on the validation set plays a key role in balancing the model's fit to the training set and its generalization capabilities across diverse scenarios.

[0111] Test set: Contains 2,151 images, representing approximately 10% of the entire dataset. The test set is primarily used to evaluate the performance and generalization capabilities of a trained model. The parameters obtained on the training set can be used to test the model's performance on the test set to ensure its ability to accurately detect new data samples. The test set performance test results are crucial for assessing the model's generalization capabilities and its adaptability to diverse road environment samples.

[0112] The advantage of this data partitioning approach is that it creates a complete model development loop through the synergistic effect of the training, validation, and test sets, forming a comprehensive optimization mechanism from feature learning to performance verification. The training set, with an 80% data share, provides a sufficient sample size, enabling the model to fully learn multi-scale object features and complex scene patterns, such as capturing the appearance changes of vehicles under different lighting conditions in the VOC2007 dataset. The validation set, serving as the dynamic tuning hub, monitors the training process in real time at a 10% share. Based on periodic loss monitoring (such as the trigger threshold for the early stopping mechanism) and performance fluctuation analysis (such as the convergence trend of the validation set mean average performance), it adaptively adjusts the learning rate decay strategy and regularization strength, effectively preventing the model from becoming trapped in local optima or overfitting the training data. The test set, with 10% independent data, verifies the final model performance without any touch, ensuring objectivity and avoiding bias caused by parameter overfitting to the validation set. The three sets complement each other to form a hierarchical evaluation system, which not only strengthens the model's feature extraction capabilities but also ensures its generalization performance in diverse scenarios.

[0113] Therefore, a reasonable division ratio among the training set, validation set, and test set can achieve a balance between dynamic tuning and objective evaluation through the synergistic effect of the three, thereby comprehensively improving the practical application effect and robustness of the model in complex road conditions.

[0114] Table 1 Distribution of the number of categories in the VOC2007 dataset

[0115]

[0116]

[0117] Step 2: Improve the YOLOv8s network architecture by incorporating a Space-to-Depth Convolution (SPD-Conv) module into the backbone and neck network. This module consists of a Space-to-Depth (SPD) layer and a non-strided Convolution (Conv) layer and can be applied to most CNN architectures. This can improve the detection accuracy of small objects and objects in complex road conditions, and enhance the model's ability to detect road conditions.

[0118] The YOLOv8s network structure mainly includes the following parts:

[0119] Backbone network: CSPDarknet53 is used as the backbone network for extracting image features. The CSPDarknet53 network structure is lightweight and efficient, reducing computational effort while maintaining accuracy.

[0120] Neck Network: A Path Aggregation Network (PANet) is used as the Neck network to fuse feature maps of different scales. PANet connects feature maps of different scales through top-down and bottom-up paths, enhancing the semantic and spatial information of the features.

[0121] Both the backbone network and the neck network use the spatial depth conversion convolution module (SPD-Conv) for feature extraction.

[0122] The Spatial-Deepth Convolution (SPD-Conv) module converts feature maps from space to depth, reducing their size by half and increasing the number of channels by fourfold. This enables the model to more precisely capture multi-scale details in road scenes, significantly improving the detection accuracy of small objects such as pedestrians and cyclists. This module plays a key role in autonomous driving road condition detection. By enhancing target separation in complex backgrounds (such as vehicle occlusion and nighttime reflective scenes), it effectively addresses missed detections caused by object size differences or uneven lighting, while maintaining real-time inference efficiency.

[0123] SPD structure is as follows Figure 3 As shown in the figure, SPD extends the original image conversion technology to the downsampling process of the feature map inside and throughout the CNN, and slices the given intermediate feature map Y of arbitrary size S×S×C, that is, the obtained sub-feature map p x,y Y is downsampled by the ratio ratio, and then all sub-feature maps are spliced ​​by channel to obtain the new feature map Y′. The feature map has the following transformation:

[0124]

[0125] After the SPD layer, add a non-strided (i.e. stride = 1) convolution layer with C1 convolution kernel, where C1 <ratio 2 C, the intermediate feature map is transformed as follows:

[0126]

[0127] The reason for using non-strided convolution is to preserve the discriminative feature information as much as possible. Because strided convolution will lead to non-discriminative loss of information, for example, when stride = 2, asymmetric sampling will occur, in which the number of samples of even and odd rows (columns) is different.

[0128] Step 3: Input the training dataset into the improved YOLOv8s-SPD to conduct model training to build a deep learning model suitable for binocular vision system detection in road scenes. The specific process is as follows:

[0129] Modify the YOLOv8.cfg file and change the classes in the data's yaml file to the number of categories in the dataset.

[0130] Set the hyperparameters of the network model, including the input image size of 640×640 for training the dataset and testing the model performance, the input data size per batch (batch_size) of 32, the number of training rounds (epochs) of 300, and the learning rate (learningrate) of 0.01.

[0131] This paper uses the evaluation system of the target detection method to evaluate the model, including the precision AP, recall rate (Recall), F1-score, precision (Precision) and average precision mAP@0.5 (mean average precision, IoU threshold is 0.5) of each class. The calculation formula is as follows:

[0132]

[0133] TP represents the number of correctly identified positive samples, TN represents the number of correctly identified negative samples, FP represents the number of negative samples that are incorrectly identified as positive samples, and FN represents the number of positive samples that are incorrectly identified as negative samples. The F1 score can be further calculated using the precision and recall rates. TP Refers to the correctly identified positive samples, f TF is a correctly identified negative sample, f TNThe average precision (AP) of each target category can be calculated by dividing the area between the curve (PR curve) composed of precision and recall and the coordinate axis.

[0134] Since the purpose of the improved model proposed in this invention is to improve the target detection capability of the YOLOv8s algorithm, the main indicator for evaluating the model performance is mAP@0.5. Therefore, the average precision of the YOLOv8s model and the YOLOv8s-SPD proposed in this invention on the VOC2007 dataset is compared, which shows that the improved model has higher detection accuracy. Figure 4a 、 Figure 4b shown.

[0135] Use the training set to train the YOLOv8s-SPD model. The training results are as follows Figure 5a 、 5b , 5c, 5d and Figure 6a 、 6b , as shown in Figure 6c. val_box represents the bounding box of the validation set. Its value decreases and finally stabilizes after 150 training rounds. val_dfl represents the mean loss of the validation set. The mean loss stabilizes after 200 training rounds. val_cls represents the mean classification loss of the validation set. The mean loss converges after 200 training rounds. Precision represents the percentage of positive classes found in the validation set. The Precision value stabilizes after 250 training rounds. Recall, based on actual results, describes how many true positive examples are recalled by the binary classifier. The recall rate gradually converges after 180 training rounds.

[0136] Step 4: Use the validation set and test set to verify the performance of the trained YOLOv8s-SPD model, and comprehensively evaluate the detection effect of the model through multiple indicators; the specific process is as follows:

[0137] After the test set is fed into the YOLOv8s-SPD model, the model detects each image and outputs object information, including its category, location, and confidence level. We then compare the model's predictions with the true labels of the images in the test set and calculate key performance metrics such as precision, recall, and F1 score.

[0138] This comprehensive evaluation method not only measures the model's overall performance but also reveals its strengths and weaknesses when detecting different types of objects. Through detailed analysis of the test results, we gain a deeper understanding of the YOLOv8s-SPD model's effectiveness in real-world scenarios and provide technical support for subsequent optimization and deployment. This systematic verification process will help us objectively evaluate the model's actual performance and ensure its stability and effectiveness in security inspections.

[0139] Table 2 shows the comparison results of the detection performance of the proposed method and the YOLOv8s baseline model target detection algorithm. The comparison clearly shows that the proposed method exhibits higher detection accuracy and stronger robustness, surpassing the baseline algorithm. In addition, the proposed method has higher accuracy than the baseline model for eleven categories in the dataset.

[0140] Table 2 Accuracy comparison of YOLOv8s-SPD model and baseline model on VOC2007 dataset

[0141]

[0142]

[0143] The detection effect of the YOLOv8s-SPD model proposed in this paper on the test set is as follows Figure 7a 、 7b As shown in 7c, it can be seen that the YOLOv8s-SPD model can accurately detect pedestrians, vehicles and other targets in complex road environments.

[0144] Step 5: The binocular camera obtains the left view and the right view;

[0145] The left view is input into the trained improved YOLOv8s network model, which outputs the target category and two-dimensional bounding box coordinates in the left view (for example, outputting two targets A and B in an image, as well as the center point position xy, width w, and height h of each target); targets such as vehicles and pedestrians;

[0146] Obtain the disparity map of the left view and the right view based on the SGBM algorithm;

[0147] According to the disparity map and the known binocular camera parameters, the depth value corresponding to each pixel in the disparity map is calculated to generate a depth map;

[0148] Based on the depth map, the target category and two-dimensional bounding box coordinates in the left view output by the trained improved YOLOv8s network model, the distance of the target object (the distance between targets A and B) is calculated.

[0149] The SGBM stereo matching algorithm is a technology for depth image calculation in global matching stereo vision. It can estimate the disparity between the left and right images by matching pixel blocks in them, and then calculate the depth information of the object.

[0150] The advantage of the SGBM algorithm is that it exploits the correlation between pixel blocks and considers information from surrounding pixels when calculating disparity. This allows for relatively accurate depth images and effectively handles the complex conditions found in real-world scenarios. Furthermore, it employs several optimization strategies to significantly increase computational speed.

[0151] This invention selected a video of a road section with low traffic volume during the day for testing. In the traffic flow, there are pedestrians, electric vehicles, cars parked on the roadside and other objects, which conforms to a relatively real complex test environment. Under the operation of the algorithm system, these objects can be identified. The detection effect is shown in the figure below. Figure 8a 、 Figure 8b As shown in FIG. , the effectiveness and superiority of the present invention in road condition detection are demonstrated.

[0152] The present invention may have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art may make various corresponding changes and modifications based on the present invention, but these corresponding changes and modifications should all fall within the scope of protection of the claims attached to the present invention.

Claims

1. A three-dimensional ranging method based on YOLOv8s-SPD and SGBM, characterized by: The specific process of the method is: Step 1: Get the data set; the specific process is: The dataset is the VOC2007 dataset; the VOC2007 dataset is used as the training set; Step 2: Build the YOLOv8s-SPD network model; The YOLOv8s-SPD network model includes the backbone network, the neck network, and the detection head; The working process of the YOLOv8s-SPD network model is as follows: The image is sequentially input into the backbone network, the neck network, and the detection head. The detection head outputs the target category and the two-dimensional bounding box coordinates. Step 3: Input the training set into the YOLOv8s-SPD network model, train the YOLOv8s-SPD network model, and obtain the trained YOLOv8s-SPD network model; Step 4: The binocular camera obtains the left view and the right view; The left view is input into the trained YOLOv8s-SPD network model, and the trained YOLOv8s-SPD network model outputs the target category and two-dimensional bounding box coordinates in the left view; Obtain the disparity map of the left view and the right view based on the SGBM algorithm; According to the disparity map and the known binocular camera parameters, the depth value corresponding to each pixel in the disparity map is calculated to generate a depth map; The distance of the target object is calculated based on the depth map, the target category and the two-dimensional bounding box coordinates in the left view output by the trained YOLOv8s-SPD network model.

2. The three-dimensional ranging method based on YOLOv8s-SPD and SGBM according to claim 1, characterized in that: The backbone network Backbone includes in sequence: input layer, first convolution layer, first spatial depth conversion convolution module, first C2f module, second spatial depth conversion convolution module, second C2f module, third spatial depth conversion convolution module, third C2f module, fourth spatial depth conversion convolution module, fourth C2f module, SPPF module.

3. The three-dimensional ranging method based on YOLOv8s-SPD and SGBM according to claim 2, characterized in that: The working process of the backbone network in step 2 is as follows: The image is input into the first convolutional layer through the input layer, and the first convolutional layer outputs the feature map A; Feature map A is sequentially input into the first spatial depth conversion convolution module and the first C2f module, and the first C2f module outputs feature map B; Feature map B is sequentially input into the second spatial depth conversion convolution module and the second C2f module, and the second C2f module outputs feature map C; The feature map C is sequentially input into the third spatial depth conversion convolution module and the third C2f module, and the third C2f module outputs the feature map D; The feature map D is sequentially input into the fourth spatial depth conversion convolution module and the fourth C2f module, and the fourth C2f module outputs the feature map E; The feature map E is input into the SPPF module, and the SPPF module outputs the feature map F.

4. The three-dimensional ranging method based on YOLOv8s-SPD and SGBM according to claim 3, characterized in that: The working process of the first spatial depth conversion convolution module is as follows: The size of feature map A is S×S×C; the first S is the length of feature map A, the second S is the width of feature map A, and C is the number of channels of feature map A; Divide the feature map A into 4 feature maps, namely feature map A1, feature map A2, feature map A3, and feature map A4; The size of feature map A1 is The size of feature map A2 is The size of feature map A3 is The size of feature map A4 is The feature maps A1, A2, A3, and A4 are spliced ​​together to obtain the feature map A′. The size of the feature map A′ is The feature map A′ is input into the second convolutional layer, and the second convolutional layer outputs the feature map A″. The size of the feature map A″ is 5. The three-dimensional ranging method based on YOLOv8s-SPD and SGBM according to claim 4, characterized in that: The neck network Neck in step 2 includes: First upsampling layer, first splicing layer, fifth C2f, second sampling layer, second splicing layer, sixth C2f, third convolutional layer, third splicing layer, seventh C2f, fourth convolutional layer, fourth splicing layer, eighth C2f.

6. The three-dimensional ranging method based on YOLOv8s-SPD and SGBM according to claim 5, characterized in that: The working process of the neck network Neck is: The feature map F output by the SPPF module is input into the first upsampling layer, and the first upsampling layer outputs the feature map G; The feature map G and the feature map D output by the third C2f module are input into the first splicing layer for splicing, and the first splicing layer outputs the feature map H; The feature map H is input into the fifth C2f, and the fifth C2f outputs the feature map I; Feature map I is input into the second upsampling layer, and the second upsampling layer outputs feature map J; Feature map J and the second C2f module output feature map C are input into the second splicing layer for splicing, and the second splicing layer outputs feature map K; The feature map K is input into the sixth C2f, and the sixth C2f outputs the feature map L; The feature map L is input into the third convolutional layer, and the third convolutional layer outputs the feature map M; The feature map M and the fifth C2f output feature map I are input into the third splicing layer, and the third splicing layer outputs the feature map N; The feature map N is input into the seventh C2f, and the seventh C2f outputs the feature map O; The feature map O is input into the fourth convolutional layer, and the fourth convolutional layer outputs the feature map P; The feature map P and SPPF module output feature map F are input into the fourth splicing layer, and the fourth splicing layer outputs feature map Q; The feature map Q is input into the eighth C2f, and the eighth C2f outputs the feature map R.

7. The three-dimensional ranging method based on YOLOv8s-SPD and SGBM according to claim 6, characterized in that: The detection head in step 2 includes a first detection head, a second detection head, and a third detection head; The sixth C2f outputs a feature map L as the output result of the first detection head; The seventh C2f outputs a feature map O as the output result of the second detection head; The eighth C2f outputs the feature map R as the output result of the third detection head.

8. The three-dimensional ranging method based on YOLOv8s-SPD and SGBM according to claim 7, characterized in that: The VOC2007 dataset contains 20 categories, namely: Airplane, bicycle, bird, boat, bottle, bus, car, cat, chair, cow, dining table, dog, horse, motorcycle, person, potted plant, sheep, sofa, train, TV.

9. The three-dimensional ranging method based on YOLOv8s-SPD and SGBM according to claim 8, characterized in that: In step 3, the training set is input into the YOLOv8s-SPD network model, and the YOLOv8s-SPD network model is trained to obtain a trained YOLOv8s-SPD network model; the specific process is: The image size in the training set is 640×640; During training, the input data size for each batch is 32, the number of training rounds is 300, and the learning rate is 0.01.