Binocular distance measurement method based on YOLOv8-BiFPN network model
Through the YOLOv8-BiFPN network model, using BiFPN and data enhancement technology, the problem of feature loss in scenarios where the feature pyramid fixed weight fusion mechanism is difficult to adapt to multi-scale target detection and dynamic occlusion is solved. This achieves a balance between high-precision depth perception and real-time detection, and improves the robustness and efficiency of the binocular vision system.
Patent Information
- Application Number
- CN202510934814.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-10-10
AI Technical Summary
In the existing technology, the fixed weight fusion mechanism of the feature pyramid is difficult to adapt to multi-scale target detection, features are lost in dynamic occlusion scenarios, and the contradiction between algorithm computational complexity and real-time requirements is difficult to balance, especially in binocular vision systems where high-precision depth perception and real-time detection are difficult to balance.
A YOLOv8-BiFPN network model is adopted. By introducing BiFPN to replace the traditional FPN, its cross-scale bidirectional fusion mechanism with learnable weights is utilized, combined with data enhancement technologies such as three-dimensional spatial perturbation and Mosaic splicing, to optimize multi-scale feature fusion and feature extraction capabilities under dynamic occlusion.
It significantly improves the robustness and accuracy of large and small target detection, reduces computational complexity, resolves the contradiction between high-precision depth perception and real-time performance, and meets the three-dimensional perception needs in harsh scenarios.
Smart Images

Figure CN120765749A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and target detection, and particularly relates to a binocular ranging method based on the YOLOv8-BiFPN network model. Background Art
[0002] In fields with stringent requirements for three-dimensional perception, such as autonomous driving and robotic navigation, acquiring distance information of target objects in real time, robustly, and accurately is a core requirement. Binocular vision technology, with its passive depth perception capabilities, has become an important tool. However, traditional feature matching algorithms lack robustness in scenes with complex lighting, low textures, and occlusions, and are prone to matching failures or parallax calculation errors. Furthermore, high-precision algorithms are often computationally intensive, making it difficult to achieve real-time performance.
[0003] To address these challenges, the academic community has achieved numerous breakthroughs in object detection in recent years. In stereo matching algorithms, traditional methods such as Block Matching (BM), Graph Cut (GC), and Semi-Global Block Matching (SGBM) have significantly improved the computational efficiency and accuracy of disparity maps by optimizing local or global matching strategies. At the same time, end-to-end models based on deep learning have gradually become mainstream. For example, the YOLO series balances speed and accuracy through a single-stage detection architecture, while the Feature Pyramid Network (FPN) enhances the detection of small objects through multi-scale feature fusion. Furthermore, loss function optimization techniques effectively improve algorithm stability in complex scenarios by dynamically adjusting matching costs. Innovations in data augmentation techniques have also provided new insights for industrial scene modeling. For example, three-dimensional spatial perturbations simulate distance changes and mosaic stitching enhances occlusion robustness, significantly alleviating the problem of insufficient model generalization. Transfer learning and lightweight model design have further promoted the implementation of algorithms on mobile devices. For example, by reducing computational complexity through model pruning and quantization, they can meet industrial real-time requirements.
[0004] However, existing technologies still have significant limitations in practical applications. The fixed-weight fusion mechanism of the traditional feature pyramid struggles to adapt to the dramatic changes in multi-scale object detection, and the problem of partial feature loss in dynamic occlusion scenarios remains unresolved. Furthermore, the contradiction between the algorithm's computational complexity and real-time requirements is becoming increasingly prominent, making it difficult to achieve both the high accuracy of a two-stage detector and the low latency of a single-stage model. These bottlenecks collectively hinder the comprehensive upgrade of intelligent logistics systems. Therefore, an innovative solution is urgently needed that combines the high efficiency of monocular detection with the precise depth perception advantages of binocular vision. Summary of the Invention
[0005] The purpose of this invention is to solve the problems in the existing technology, such as the difficulty of adapting the fixed weight fusion mechanism of feature pyramid to multi-scale target detection, feature loss in dynamic occlusion scenarios, and the contradiction between algorithm computational complexity and real-time requirements. In particular, to address the defect that high-precision depth perception and real-time detection are difficult to achieve in binocular vision systems, a binocular ranging method based on the YOLOv8-BiFPN network model is proposed to improve the robustness and efficiency of three-dimensional perception in complex scenarios.
[0006] A binocular ranging method based on the YOLOv8-BiFPN network model has the following specific steps:
[0007] Step 1: Obtain a data set, perform data preprocessing and data enhancement on the acquired data set to obtain a training set;
[0008] Step 2: Build the YOLOv8-BiFPN network model;
[0009] Step 3: Input the training set into the YOLOv8-BiFPN network model, train the YOLOv8-BiFPN network model, and obtain the trained YOLOv8-BiFPN network model;
[0010] Step 4: Calculate the distance of the target object based on the binocular camera and the trained YOLOv8-BiFPN network model.
[0011] The beneficial effects of the present invention are:
[0012] The present invention introduces BiFPN to replace the traditional FPN, and utilizes its cross-scale bidirectional fusion mechanism with learnable weights to adaptively optimize multi-scale feature fusion, significantly improving the robustness and accuracy of large and small target detection, and effectively overcoming the problem that fixed weights are difficult to adapt to drastic changes in scale. Innovative data enhancement technologies such as three-dimensional scale space perturbation and Mosaic splicing are used to simulate distance changes and complex occlusion scenes, greatly enhancing the model's feature extraction capabilities under dynamic occlusion and reducing feature loss. A comparison of the YOLOv8-BiFPN model of the present invention with the traditional YOLOv8s model shows that the size of the YOLOv8-BiFPN model of the present invention is 5.37MB, and the size of the traditional YOLOv8s model is 5.39MB. The inference delay FPS of the YOLOv8-BiFPN model of the present invention is 121.95, and the inference delay FPS of the traditional YOLOv8s model is 100.25.
[0013] While ensuring the accuracy of binocular ranging, the computational complexity is significantly reduced, successfully resolving the contradiction between high-precision depth perception and real-time performance, and meeting the three-dimensional perception needs in harsh scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 This is a flow chart of the general object detection method of YOLOv8-BiFPN in the method of the present invention;
[0015] Figure 2 Schematic diagram of the general object detection network structure of YOLOv8-BiFPN in the method of the present invention. Conv3×3 / 2,64 indicates a convolution kernel size of 3×3, a stride of 2, and 64 output channels; C2f×3,128 indicates a stride of 3 and 128 output channels; Upsample2× indicates 2x upsampling.
[0016] Figure 3 This is a training index result diagram of YOLOv8-BiFPN and YOLOv8s in the method of the present invention on the training set;
[0017] (a) shows the train / box_loss results for YOLOv8-BiFPN and YOLOv8s on the training set. train / box_loss is the training bounding box loss, which measures the difference or error between the predicted object bounding box during training and the ground-truth bounding box. Smaller values indicate more accurate predictions.
[0018] (b) shows the train / cls_loss results for YOLOv8-BiFPN and YOLOv8s on the training set. train / cls_loss is the training classification loss, which evaluates the accuracy of the model's object classification predictions during training. Smaller values indicate better classification performance.
[0019] (c) is the train / dfl_loss result diagram of YOLOv8-BiFPN and YOLOv8s on the training set. train / dfl_loss is the training distribution focus loss, which is a specific loss function (usually distribution focus loss, Distribution Focal Loss) used to optimize bounding box prediction during the training phase, aiming to more accurately locate objects.
[0020] (d) shows the metrics / precision (B) results for YOLOv8-BiFPN and YOLOv8s on the training set. Metrics / precision (B) represents the accuracy on the test set, and the function represents the precision calculated on the test set. This value reflects the proportion of samples predicted as positive by the model that are actually positive. A higher value indicates a lower false positive rate.
[0021] (e) shows the metrics / recall (B) results for YOLOv8-BiFPN and YOLOv8s on the training set. Metrics / recall (B) is the recall rate on the test set. This function is the recall rate calculated on the test set. It reflects the proportion of all true positive examples correctly predicted by the model. A higher value indicates fewer missed detections by the model.
[0022] (f) shows the val / box_loss results for YOLOv8-BiFPN and YOLOv8s on the training set. val / box_loss is the validation bounding box loss, which measures the difference or error between the predicted object bounding box during the validation phase and the ground-truth bounding box. Smaller values indicate better localization on unseen data and are used to monitor overfitting during training.
[0023] (g) shows the val / cls_loss results for YOLOv8-BiFPN and YOLOv8s on the training set. val / cls_loss is the validation classification loss, which evaluates the accuracy of the model's object classification predictions during the validation phase. Smaller values indicate better classification performance on unseen data and are used to monitor overfitting during training.
[0024] (h) shows the valdfl_loss results for YOLOv8-BiFPN and YOLOv8s on the training set. valdfl_loss is the validation distribution-focused loss, a specific loss function (distribution-focused loss) calculated during the validation phase. It is used to monitor the performance of this loss on unseen data and determine the generalization ability of positioning optimization.
[0025] (i) shows the metrics / mAP50(B) results for YOLOv8-BiFPN and YOLOv8s on the training set. Metrics / mAP50(B) is the mean average precision (IoU=0.5) on the test set. mAP is the mean average precision (mAP) calculated on the test set, with an IoU threshold of 0.5. This is a core comprehensive metric for object detection, measuring the overall detection accuracy (combining precision and recall) of the model under a relaxed IoU threshold of 0.5.
[0026] (j) is the metrics / mAP50-95(B) result of YOLOv8-BiFPN and YOLOv8s on the training set. Metrics / mAP50-95(B) is the mean average precision of the test set (IoU = 0.5:0.95). The function is the average precision (mAP) calculated on the test set. The evaluation standard is the average of the IoU threshold from 0.5 to 0.95 (the step size is usually 0.05). This is a more rigorous and comprehensive core comprehensive indicator that measures the overall detection performance of the model under different positioning accuracy requirements (IoU from loose to strict);
[0027] Figure 4 The figure shows the average precision comparison of YOLOv8-BiFPN and YOLOv8s on the training set in the method of the present invention. (a) is the mAP50 performance comparison of the YOLOv8-BiFPN and YOLOv8s model training process. (B) is the loss value comparison of the YOLOv8-BiFPN and YOLOv8s model training process.
[0028] Figure 5 This is the detection result diagram of YOLOv8-BiFPN and YOLOv8s in the method of the present invention on the test set. DETAILED DESCRIPTION
[0029] Specific implementation method 1: This implementation method is a binocular ranging method based on the YOLOv8-BiFPN network model. The specific process is as follows:
[0030] Step 1: Obtain a data set, perform data preprocessing and data enhancement on the acquired data set to obtain a training set;
[0031] Step 2: Build the YOLOv8-BiFPN network model;
[0032] Combined with the Bidirectional Feature Pyramid Network (BiFPN) module, the YOLOv8 model is improved to build the YOLOv8-BiFPN network;
[0033] Step 3: Input the training set into the YOLOv8-BiFPN network model, train the YOLOv8-BiFPN network model, and obtain the trained YOLOv8-BiFPN network model;
[0034] Step 4: Calculate the distance of the target object based on the binocular camera and the trained YOLOv8-BiFPN network model.
[0035] Specific embodiment 2: This embodiment differs from specific embodiment 1 in that in step 1, a data set is obtained, and data preprocessing and data enhancement are performed on the data set to obtain a training set; the specific process is as follows:
[0036] Step 1: Get the VOC2007 dataset;
[0037] Step 1 and 2: Create the VOCdevkit folder;
[0038] The VOCdevkit folder contains the subfolders Annotations, JPEGImages, and ImageSets / Main.
[0039] The subfolder Annotations is used to store annotation files in TXT format;
[0040] The subfolder JPEGImages is used to store images in the VOC2007 dataset;
[0041] The subfolder ImageSets / Main is used to store the text list of the training set (the images in the training set are arranged in columns);
[0042] Step 13: Use the labeling tool LabelImg to draw bounding boxes and label categories on the images in the VOC2007 dataset and generate a labeling file in XML format;
[0043] Convert the XML annotation file to a TXT annotation file suitable for YOLO, and store the TXT annotation file in the Annotations folder.
[0044] Step 14: Perform data enhancement on the labeled file in TXT format to obtain the training set.
[0045] Data augmentation includes random flipping, random cropping and scaling, color jittering, random grayscale conversion, and random lighting transformation;
[0046] Random flipping is: flipping the image horizontally or vertically;
[0047] Random cropping and scaling involves randomly extracting a local region from the image (random cropping (keeping ≥80% of the target)) and scaling the cropped region to the target size at a random ratio to further simulate the morphological changes of objects at different distances and enhance the model's adaptability to scale changes.
[0048] Color dithering adjusts image color attributes such as brightness, contrast, and saturation to simulate lighting changes and differences in shooting equipment, thereby improving the model's robustness to color interference.
[0049] Random grayscale conversion is to convert a color image into a grayscale image with a certain probability, reducing the model's over-reliance on color information and is suitable for color-independent classification tasks (such as shape recognition).
[0050] Random illumination transformation: simulates image effects (such as shadows and overexposure) under different lighting conditions, enhancing the model's generalization ability in complex lighting scenes;
[0051] Other steps and parameters are the same as those in the first embodiment.
[0052] Specific embodiment three: This embodiment differs from specific embodiment one or two in that a YOLOv8-BiFPN network model is constructed in step two;
[0053] Combined with the Bidirectional Feature Pyramid Network (BiFPN) module, the YOLOv8 model is improved to build the YOLOv8-BiFPN network;
[0054] The specific process is:
[0055] The YOLOv8-BiFPN network model includes the backbone network, the neck network, and the detection head;
[0056] The working process of the YOLOv8-BiFPN network model is as follows:
[0057] The image is sequentially input into the backbone network, the neck network, and the detection head, which outputs the target category and two-dimensional bounding box coordinates.
[0058] Other steps and parameters are the same as those in the first or second embodiment.
[0059] Specific embodiment 4: This embodiment differs from any one of specific embodiments 1 to 3 in that the backbone network Backbone includes, in sequence: a first convolutional layer, a second convolutional layer, a first C2f module, a third convolutional layer, a second C2f module, a fourth convolutional layer, a third C2f module, a fifth convolutional layer, a fourth C2f module, and an SPPF module;
[0060] The SPPF module includes: a sixth convolutional layer, a first maximum pooling layer, a second maximum pooling layer, a third maximum pooling layer, a first concatenation layer Concat, and a seventh convolutional layer.
[0061] The other steps and parameters are the same as those in the first to third embodiments.
[0062] Specific embodiment 5: This embodiment differs from specific embodiments 1 to 4 in that the working process of the backbone network Backbone is as follows:
[0063] The image is sequentially input into the first convolutional layer, the second convolutional layer, the first C2f module, and the third convolutional layer. The third convolutional layer outputs the feature map A.
[0064] Feature map A is sequentially input into the second C2f module and the fourth convolutional layer, and the fourth convolutional layer outputs feature map B;
[0065] Feature map B is sequentially input into the third C2f module, the fifth convolutional layer, and the fourth C2f module, and the fourth C2f module outputs feature map C;
[0066] Feature map C is input into the SPPF module, and the SPPF module outputs feature map D.
[0067] The other steps and parameters are the same as those in the first to fourth embodiments.
[0068] Specific embodiment 6: This embodiment differs from any one of specific embodiments 1 to 5 in that the feature map C is input into the SPPF module, and the SPPF module outputs the feature map D; the specific process is as follows:
[0069] Feature map C is input into the sixth convolutional layer, and the sixth convolutional layer outputs feature map E;
[0070] Feature map E is input into the first maximum pooling layer, and the first maximum pooling layer outputs feature map F;
[0071] The feature map F is input into the second maximum pooling layer, and the second maximum pooling layer outputs the feature map G;
[0072] Feature map F and feature map G are input into the third maximum pooling layer, and the third maximum pooling layer outputs feature map H;
[0073] Feature map E and feature map H are input into the first concatenation layer Concat, and the first concatenation layer Concat outputs feature map I;
[0074] The feature map I is input into the seventh convolutional layer, and the seventh convolutional layer outputs the feature map D.
[0075] The other steps and parameters are the same as those in the first to fifth embodiments.
[0076] Specific embodiment seven: This embodiment is different from any one of specific embodiments one to six in that the neck network Neck includes: a BiFPN module, a first upsampling layer, a fifth splicing layer Concat, a fifth C2f module, a second upsampling layer, a sixth splicing layer Concat, a sixth C2f module, a fourteenth convolutional layer, a seventh splicing layer Concat, a seventh C2f module, a fifteenth convolutional layer, an eighth splicing layer Concat, and an eighth C2f module;
[0077] The BiFPN module comprises an eighth convolutional layer, a ninth convolutional layer, a tenth convolutional layer, an eleventh convolutional layer, a twelfth convolutional layer, a thirteenth convolutional layer, a second splicing layer Concat, a third splicing layer Concat, and a fourth splicing layer Concat;
[0078] The specific working process of the neck network Neck is as follows:
[0079] The feature map A, the feature map B, and the feature map D input the BiFPN module, and the BiFPN module outputs a feature map A', a feature map B', and a feature map D';
[0080] The feature map D' inputs a first upsampling layer, and the first upsampling layer outputs a feature map D";
[0081] The feature map D" and the feature map B' input a fifth splicing layer Concat, and the fifth splicing layer Concat outputs a feature map D"';
[0082] The feature map D"' inputs a fifth C2f module, and the fifth C2f module outputs a feature map
[0083] The feature map and the feature map B' are spliced to obtain a feature map B";
[0084] The feature map B" inputs a second upsampling layer, and the second upsampling layer outputs a feature map B"';
[0085] The feature map B"', the feature map B", and the feature map A' input a sixth splicing layer Concat, and the sixth splicing layer Concat outputs a feature map
[0086] The feature map inputs a sixth C2f module, and the sixth C2f module outputs a feature map G';
[0087] The feature map G' inputs a fourteenth convolutional layer, and the fourteenth convolutional layer outputs a feature map A";
[0088] The feature map A" and the feature map input a seventh splicing layer Concat, and the seventh splicing layer Concat outputs a feature map A"';
[0089] The feature map A"' inputs a seventh C2f module, and the seventh C2f module outputs a feature map H';
[0090] The feature map H' inputs a fifteenth convolutional layer, and the fifteenth convolutional layer outputs a feature map
[0091] The feature map The feature map D' is input into an eighth concatenation layer Concat, and the eighth concatenation layer Concat outputs a feature map E';
[0092] The feature map E' is input into an eighth C2f module, and the eighth C2f module outputs a feature map F'.
[0093] The other steps and parameters are the same as one of the first to sixth embodiments.
[0094] The eighth embodiment is different from one of the first to seventh embodiments in that the feature map A, the feature map B and the feature map D are input into a BiFPN module, and the BiFPN module outputs the feature map A', the feature map B' and the feature map D'.
[0095] The feature map A is input into an eighth convolutional layer, and the eighth convolutional layer outputs a feature map A1;
[0096] The feature map A1 is input into a ninth convolutional layer, and the ninth convolutional layer outputs a feature map A2;
[0097] The feature map B is input into a tenth convolutional layer, and the tenth convolutional layer outputs a feature map B1;
[0098] The feature map B1 and the feature map A2 are input into an eleventh convolutional layer, and the eleventh convolutional layer outputs a feature map B2;
[0099] The feature map D is input into a twelfth convolutional layer, and the twelfth convolutional layer outputs a feature map D1;
[0100] The feature map D1 and the feature map B2 are input into a thirteenth convolutional layer, and the thirteenth convolutional layer outputs a feature map D2;
[0101] The feature map D1, the feature map D2 and the feature map B2 are input into a second concatenation layer Concat for splicing to obtain a feature map D';
[0102] The feature map B1, the feature map B2 and the feature map D' are input into a third concatenation layer Concat for splicing to obtain a feature map B';
[0103] The feature map A1, the feature map A2 and the feature map B' are input into a fourth concatenation layer Concat for splicing to obtain a feature map A'.
[0104] The other steps and parameters are the same as one of the first to seventh embodiments.
[0105] The ninth embodiment is different from one of the first to eighth embodiments in that the detection head includes a detection head 1, a detection head 2 and a detection head 3.
[0106] The sixth C2f module outputs the feature map G' as an output result of the detection head 1, and outputs a target category and a two-dimensional bounding box coordinate;
[0107] The seventh C2f module outputs a feature map H' as an output result of the detection head 2, and outputs a target category and a two-dimensional bounding box coordinate;
[0108] The eighth C2f module outputs a feature map F' as an output result of the detection head 3, and outputs a target category and a two-dimensional bounding box coordinate.
[0109] The other steps and parameters are the same as one of the first to eighth embodiments.
[0110] The tenth embodiment is different from one of the first to ninth embodiments in that the distance of the target object is calculated based on the binocular camera and the trained YOLOv8-BiFPN network model in the fourth step; and the specific process is as follows:
[0111] The binocular camera obtains a left view and a right view;
[0112] The left view is input into the trained YOLOv8s-SPD network model, and the trained YOLOv8s-SPD network model outputs a target category and a two-dimensional bounding box coordinate in the left view (for example, outputs two targets A and B in a picture, and the center point positions xy, width w, and height h of each target); the target is, for example, a vehicle or a pedestrian;
[0113] The disparity map of the left view and the right view is obtained based on the SGBM algorithm;
[0114] According to the disparity map and the known binocular camera parameters, the depth value corresponding to each pixel point in the disparity map is calculated to generate a depth map;
[0115] The distance of the target object (the distance of the two targets A and B) is calculated based on the depth map and the target category and the two-dimensional bounding box coordinate in the left view output by the trained YOLOv8s-SPD network model.
[0116] The other steps and parameters are the same as one of the first to ninth embodiments.
[0117] The beneficial effects of the present application are verified by using the following embodiments:
[0118] Embodiment one:
[0119] Step one, obtain a data set, and perform data preprocessing and data enhancement on the obtained data set; the specific process is as follows:
[0120] Step one, obtain a VOC2007 data set;
[0121] Step two, create a VOCdevkit folder;
[0122] The VOCdevkit folder contains the subfolders Annotations, JPEGImages, and ImageSets / Main.
[0123] The subfolder Annotations is used to store annotation files in TXT format;
[0124] The subfolder JPEGImages is used to store images in the VOC2007 dataset;
[0125] The subfolder ImageSets / Main is used to store the text list of the dataset (arranging the images in the dataset in columns);
[0126] Step 13: Use the labeling tool LabelImg to draw bounding boxes and label categories on the images in the VOC2007 dataset and generate a labeling file in XML format;
[0127] Convert the XML annotation file to a TXT annotation file suitable for YOLO, and store the TXT annotation file in the Annotations folder.
[0128] Step 14: Perform data enhancement on the labeled file in TXT format to obtain the dataset.
[0129] Data augmentation includes random flipping, random cropping and scaling, color jittering, random grayscale conversion, and random lighting transformation;
[0130] Random flipping is: flipping the image horizontally or vertically;
[0131] Random cropping and scaling involves randomly extracting a local region from the image (random cropping (keeping ≥80% of the target)) and scaling the cropped region to the target size at a random ratio to further simulate the morphological changes of objects at different distances and enhance the model's adaptability to scale changes.
[0132] Color dithering adjusts image color attributes such as brightness, contrast, and saturation to simulate lighting changes and differences in shooting equipment, thereby improving the model's robustness to color interference.
[0133] Random grayscale conversion is to convert a color image into a grayscale image with a certain probability, reducing the model's over-reliance on color information and is suitable for color-independent classification tasks (such as shape recognition).
[0134] Random illumination transformation: simulates image effects (such as shadows and overexposure) under different lighting conditions, enhancing the model's generalization ability in complex lighting scenes;
[0135] VOC2007 dataset: VOC2007 dataset is a large-scale labeled image dataset for target detection, image classification and segmentation scenarios, containing 21504 images, covering 20 categories (airplane, bicycle, bird, boat, bottle, bus, car, cat, chair, cow, dining table, dog, horse, motorcycle, person, potted plant, sheep, sofa, train, TV) in daily items, totaling 62199 instances. This dataset is designed to simulate the multi-task, complex environment and diversified target detection challenges in real scenarios, and to improve model robustness through diversified object shapes, multi-scale and multi-angle objects and complex background interference. The training set of VOC2007 contains 16552 images (80%), and the test set contains 4952 images (20%).
[0136] Table 1 VOC2007 dataset structure division table
[0137]
[0138] Table 2 Data distribution of VOC2007 dataset
[0139]
[0140] Step two, build YOLOv8-BiFPN network model;
[0141] Combined with the bidirectional feature pyramid (Bidirectional Feature Pyramid Network, BiFPN) module, the YOLOv8 model is improved to build the YOLOv8-BiFPN network;
[0142] The specific process is as follows:
[0143] The YOLOv8-BiFPN network model includes a backbone network Backbone, a neck network Neck and a detection head;
[0144] The working process of the YOLOv8-BiFPN network model is as follows:
[0145] The picture is sequentially input into the backbone network Backbone, the neck network Neck and the detection head, and the detection head outputs the target class and the two-dimensional bounding box coordinates.
[0146] The backbone network Backbone includes in sequence: a first convolutional layer, a second convolutional layer, a first C2f module, a third convolutional layer, a second C2f module, a fourth convolutional layer, a third C2f module, a fifth convolutional layer, a fourth C2f module, and an SPPF module.
[0147] The SPPF module includes: a sixth convolutional layer, a first maximum pooling layer, a second maximum pooling layer, a third maximum pooling layer, a first concatenation layer Concat, and a seventh convolutional layer.
[0148] The working process of the backbone network Backbone is as follows:
[0149] The image is sequentially input into the first convolutional layer, the second convolutional layer, the first C2f module, and the third convolutional layer. The third convolutional layer outputs the feature map A.
[0150] Feature map A is sequentially input into the second C2f module and the fourth convolutional layer, and the fourth convolutional layer outputs feature map B;
[0151] Feature map B is sequentially input into the third C2f module, the fifth convolutional layer, and the fourth C2f module, and the fourth C2f module outputs feature map C;
[0152] Feature map C is input into the SPPF module, and the SPPF module outputs feature map D.
[0153] The feature map C is input into the SPPF module, and the SPPF module outputs the feature map D. The specific process is as follows:
[0154] Feature map C is input into the sixth convolutional layer, and the sixth convolutional layer outputs feature map E;
[0155] Feature map E is input into the first maximum pooling layer, and the first maximum pooling layer outputs feature map F;
[0156] The feature map F is input into the second maximum pooling layer, and the second maximum pooling layer outputs the feature map G;
[0157] Feature map F and feature map G are input into the third maximum pooling layer, and the third maximum pooling layer outputs feature map H;
[0158] Feature map E and feature map H are input into the first concatenation layer Concat, and the first concatenation layer Concat outputs feature map I;
[0159] The feature map I is input into the seventh convolutional layer, and the seventh convolutional layer outputs the feature map D.
[0160] The neck network Neck includes: a BiFPN module, a first upsampling layer, a fifth splicing layer Concat, a fifth C2f module, a second upsampling layer, a sixth splicing layer Concat, a sixth C2f module, a fourteenth convolutional layer, a seventh splicing layer Concat, a seventh C2f module, a fifteenth convolutional layer, an eighth splicing layer Concat, and an eighth C2f module;
[0161] The BiFPN module includes: an eighth convolutional layer, a ninth convolutional layer, a tenth convolutional layer, an eleventh convolutional layer, a twelfth convolutional layer, a thirteenth convolutional layer, a second concatenation layer Concat, a third concatenation layer Concat, and a fourth concatenation layer Concat;
[0162] The specific working process of the neck network Neck is:
[0163] Feature maps A, B, and D are input into the BiFPN module, and the BiFPN module outputs feature maps A′, B′, and D′;
[0164] The feature map D′ is input into the first upsampling layer, and the first upsampling layer outputs the feature map D″;
[0165] The feature map D″ and the feature map B′ are input into the fifth concatenation layer Concat, and the fifth concatenation layer Concat outputs the feature map D″′;
[0166] The feature map D″′ is input into the fifth C2f module, and the fifth C2f module outputs the feature map
[0167] Feature Map and concatenate it with the feature map B′ to obtain the feature map B";
[0168] The feature map B" is input into the second upsampling layer, and the second upsampling layer outputs the feature map B"';
[0169] Feature map B″′, feature map B" and feature map A′ are input to the sixth concatenation layer Concat, and the sixth concatenation layer Concat outputs feature map
[0170] Feature Map Input the sixth C2f module, and the sixth C2f module outputs the feature map G′;
[0171] The feature map G′ is input into the fourteenth convolutional layer, and the fourteenth convolutional layer outputs the feature map A";
[0172] Feature map A″ and feature map Input the seventh concatenation layer Concat, and the seventh concatenation layer Concat outputs the feature map A″′;
[0173] The feature map A″′ is input into the seventh C2f module, and the seventh C2f module outputs the feature map H′;
[0174] The feature map H′ is input into the fifteenth convolutional layer, and the fifteenth convolutional layer outputs the feature map
[0175] Feature Map And the feature map D′ is input into the eighth concatenation layer Concat, and the eighth concatenation layer Concat outputs the feature map E′;
[0176] The feature map E′ is input into the eighth C2f module, and the eighth C2f module outputs the feature map F′.
[0177] The feature map A, feature map B, and feature map D are input into the BiFPN module, and the BiFPN module outputs feature map A′, feature map B′, and feature map D′;
[0178] The specific process is:
[0179] Feature map A is input into the eighth convolutional layer, and the eighth convolutional layer outputs feature map A1;
[0180] Feature map A1 is input into the ninth convolutional layer, and the ninth convolutional layer outputs feature map A2;
[0181] Feature map B is input into the tenth convolutional layer, and the tenth convolutional layer outputs feature map B1;
[0182] Input feature map B1 and feature map A2 into the eleventh convolutional layer, and the eleventh convolutional layer outputs feature map B2;
[0183] The feature map D is input into the twelfth convolutional layer, and the twelfth convolutional layer outputs the feature map D1;
[0184] Feature map D1 and feature map B2 are input into the thirteenth convolutional layer, and the thirteenth convolutional layer outputs feature map D2;
[0185] Input the feature map D1, feature map D2 and feature map B2 into the second concatenation layer Concat for concatenation to obtain feature map D′;
[0186] Input the feature map B1, feature map B2 and feature map D′ into the third concatenation layer Concat for concatenation to obtain feature map B′;
[0187] The feature map A1, the feature map A2 and the feature map B′ are input into the fourth concatenation layer Concat for concatenation to obtain the feature map A′.
[0188] The detection head includes detection head 1, detection head 2, and detection head 3;
[0189] The sixth C2f module outputs the feature map G′ as the output result of the detection head 1, outputting the target category and the two-dimensional bounding box coordinates;
[0190] The seventh C2f module outputs the feature map H′ as the output result of the detection head 2, outputting the target category and the two-dimensional bounding box coordinates;
[0191] The eighth C2f module outputs the feature map F′ as the output result of the detection head 3, and outputs the target category and two-dimensional bounding box coordinates.
[0192] Step 3: Input the training set into the YOLOv8-BiFPN network model, train the YOLOv8-BiFPN network model, and obtain the trained YOLOv8-BiFPN network model;
[0193] The model is evaluated, including the precision AP, recall rate (Recall), F1-score, accuracy (Precision) and average accuracy mAP@0.5 (mean average precision, IoU threshold is 0.5) of each category; the calculation formula is as follows:
[0194]
[0195] Among them, TP represents the number of correctly identified positive samples, TN represents the number of correctly identified negative samples, FP is the number of negative samples that are mistakenly identified as positive samples, and FN is the number of positive samples that are mistakenly identified as negative samples.
[0196] The F1 score can be further calculated using the precision and recall rates.
[0197] f TP Refers to the correctly identified positive samples, f TN is a correctly identified negative sample, f TN The average precision (AP) of each target category can be calculated by dividing the area between the curve (PR curve) composed of precision and recall and the coordinate axis.
[0198] The YOLOv8-BiFPN model is trained using the data-enhanced training set, as shown in the following three loss curves and three metric curves: Figure 3 shown.
[0199] The average precision comparison results of the YOLOv8-BiFPN model and YOLOv8s on the training set are as follows Figure 4 As shown in the figure, val / box_loss represents the bounding box loss of the validation set. Its value decreases and finally stabilizes after 200 training rounds. val / cls_loss represents the mean classification loss of the validation set. The mean loss converges after 150 training rounds. val / dfl_loss represents the mean object detection loss of the validation set. The mean loss stabilizes after 150 training rounds.
[0200] Use the test set to test the trained YOLOv8-BiFPN model.
[0201] In this specific embodiment, Table 3 shows the comparison results of YOLOv8-BiFPN and YOLOv8s proposed in the present invention on the VOC2007 dataset.
[0202] Table 3 Comparison of detection accuracy of different detection methods on the VOC2007 dataset
[0203]
[0204]
[0205] The detection effect of YOLOv8-BiFPN proposed in this paper on the VOC2007 test set is as follows Figure 5 As shown, it can be seen that the model can accurately detect the items of the specified category in the dataset.
[0206] Step 4: Calculate the distance of the target object based on the binocular camera and the trained YOLOv8-BiFPN network model; the specific process is as follows:
[0207] The binocular camera obtains left and right views;
[0208] The left view is input into the trained YOLOv8s-SPD network model, which outputs the target category and two-dimensional bounding box coordinates in the left view (for example, outputting two targets A and B in an image, as well as the center point position xy, width w, and height h of each target); targets such as vehicles and pedestrians;
[0209] Obtain the disparity map of the left view and the right view based on the SGBM algorithm;
[0210] According to the disparity map and the known binocular camera parameters, the depth value corresponding to each pixel in the disparity map is calculated to generate a depth map;
[0211] Based on the depth map, the target category and two-dimensional bounding box coordinates in the left view output by the trained YOLOv8s-SPD network model, the distance of the target object (the distance between targets A and B) is calculated.
[0212] The present invention may have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art may make various corresponding changes and modifications based on the present invention, but these corresponding changes and modifications should all fall within the scope of protection of the claims attached to the present invention.
Claims
1. A binocular ranging method based on the YOLOv8-BiFPN network model, characterized by: The specific process of the method is: Step 1: Obtain a data set, perform data preprocessing and data enhancement on the acquired data set to obtain a training set; Step 2: Build the YOLOv8-BiFPN network model; Step 3: Input the training set into the YOLOv8-BiFPN network model, train the YOLOv8-BiFPN network model, and obtain the trained YOLOv8-BiFPN network model; Step 4: Calculate the distance of the target object based on the binocular camera and the trained YOLOv8-BiFPN network model.
2. A binocular ranging method based on the YOLOv8-BiFPN network model according to claim 1, characterized in that: In step 1, a data set is obtained, and data preprocessing and data enhancement are performed on the data set to obtain a training set. The specific process is as follows: Step 1: Get the VOC2007 dataset; Step 1 and 2: Create the VOCdevkit folder; The VOCdevkit folder contains the subfolders Annotations, JPEGImages, and ImageSets / Main. The subfolder Annotations is used to store annotation files in TXT format; The subfolder JPEGImages is used to store images in the VOC2007 dataset; The subfolder ImageSets / Main is used to store the text list of the training set; Step 13: Use the labeling tool LabelImg to draw bounding boxes and label categories on the images in the VOC2007 dataset and generate a labeling file in XML format; Convert the annotation file in XML format to the annotation file in TXT format, and store the annotation file in the Annotations folder; Step 14: Perform data enhancement on the labeled file in TXT format to obtain the training set.
3. A binocular ranging method based on the YOLOv8-BiFPN network model according to claim 2, characterized in that: In the step 2, a YOLOv8-BiFPN network model is constructed; the specific process is as follows: The YOLOv8-BiFPN network model includes the backbone network, the neck network, and the detection head; The working process of the YOLOv8-BiFPN network model is as follows: The image is sequentially input into the backbone network, the neck network, and the detection head, which outputs the target category and two-dimensional bounding box coordinates.
4. A binocular ranging method based on the YOLOv8-BiFPN network model according to claim 3, characterized in that: The backbone network Backbone includes in sequence: a first convolutional layer, a second convolutional layer, a first C2f module, a third convolutional layer, a second C2f module, a fourth convolutional layer, a third C2f module, a fifth convolutional layer, a fourth C2f module, and an SPPF module; The SPPF module includes: a sixth convolutional layer, a first maximum pooling layer, a second maximum pooling layer, a third maximum pooling layer, a first concatenation layer Concat, and a seventh convolutional layer.
5. A binocular ranging method based on the YOLOv8-BiFPN network model according to claim 4, characterized in that: The working process of the backbone network Backbone is as follows: The image is sequentially input into the first convolutional layer, the second convolutional layer, the first C2f module, and the third convolutional layer. The third convolutional layer outputs the feature map A. Feature map A is sequentially input into the second C2f module and the fourth convolutional layer, and the fourth convolutional layer outputs feature map B; Feature map B is sequentially input into the third C2f module, the fifth convolutional layer, and the fourth C2f module, and the fourth C2f module outputs feature map C; Feature map C is input into the SPPF module, and the SPPF module outputs feature map D.
6. The binocular ranging method based on the YOLOv8-BiFPN network model according to claim 5, characterized in that: The feature map C is input into the SPPF module, and the SPPF module outputs the feature map D. The specific process is as follows: Feature map C is input into the sixth convolutional layer, and the sixth convolutional layer outputs feature map E; Feature map E is input into the first maximum pooling layer, and the first maximum pooling layer outputs feature map F; The feature map F is input into the second maximum pooling layer, and the second maximum pooling layer outputs the feature map G; Feature map F and feature map G are input into the third maximum pooling layer, and the third maximum pooling layer outputs feature map H; Feature map E and feature map H are input into the first concatenation layer Concat, and the first concatenation layer Concat outputs feature map I; The feature map I is input into the seventh convolutional layer, and the seventh convolutional layer outputs the feature map D.
7. The binocular ranging method based on the YOLOv8-BiFPN network model according to claim 6, characterized in that: The neck network Neck includes: a BiFPN module, a first upsampling layer, a fifth splicing layer Concat, a fifth C2f module, a second upsampling layer, a sixth splicing layer Concat, a sixth C2f module, a fourteenth convolutional layer, a seventh splicing layer Concat, a seventh C2f module, a fifteenth convolutional layer, an eighth splicing layer Concat, and an eighth C2f module; The BiFPN module includes: an eighth convolutional layer, a ninth convolutional layer, a tenth convolutional layer, an eleventh convolutional layer, a twelfth convolutional layer, a thirteenth convolutional layer, a second concatenation layer Concat, a third concatenation layer Concat, and a fourth concatenation layer Concat; The specific working process of the neck network Neck is: Feature maps A, B, and D are input into the BiFPN module, and the BiFPN module outputs feature maps A′, B′, and D′; The feature map D′ is input into the first upsampling layer, and the first upsampling layer outputs the feature map D″; The feature map D″ and the feature map B′ are input into the fifth concatenation layer Concat, and the fifth concatenation layer Concat outputs the feature map D″′; The feature map D″′ is input into the fifth C2f module, and the fifth C2f module outputs the feature map Feature Map and concatenate it with the feature map B′ to obtain the feature map B″; The feature map B″ is input into the second upsampling layer, and the second upsampling layer outputs the feature map B″′; Feature map B″′, feature map B″ and feature map A′ are input to the sixth concatenation layer Concat, and the sixth concatenation layer Concat outputs feature map Feature Map Input the sixth C2f module, and the sixth C2f module outputs the feature map G′; The feature map G′ is input into the fourteenth convolutional layer, and the fourteenth convolutional layer outputs the feature map A"; Feature map A" and feature map Input the seventh concatenation layer Concat, and the seventh concatenation layer Concat outputs the feature map A″′; The feature map A″′ is input into the seventh C2f module, and the seventh C2f module outputs the feature map H′; The feature map H′ is input into the fifteenth convolutional layer, and the fifteenth convolutional layer outputs the feature map Feature Map And the feature map D′ is input into the eighth concatenation layer Concat, and the eighth concatenation layer Concat outputs the feature map E′; The feature map E′ is input into the eighth C2f module, and the eighth C2f module outputs the feature map F′.
8. The binocular ranging method based on the YOLOv8-BiFPN network model according to claim 7, characterized in that: The feature maps A, B, and D are input into the BiFPN module, and the BiFPN module outputs feature maps A′, B′, and D′. The specific process is as follows: Feature map A is input into the eighth convolutional layer, and the eighth convolutional layer outputs feature map A1; Feature map A1 is input into the ninth convolutional layer, and the ninth convolutional layer outputs feature map A2; Feature map B is input into the tenth convolutional layer, and the tenth convolutional layer outputs feature map B1; Input feature map B1 and feature map A2 into the eleventh convolutional layer, and the eleventh convolutional layer outputs feature map B2; The feature map D is input into the twelfth convolutional layer, and the twelfth convolutional layer outputs the feature map D1; Feature map D1 and feature map B2 are input into the thirteenth convolutional layer, and the thirteenth convolutional layer outputs feature map D2; Input the feature map D1, feature map D2 and feature map B2 into the second concatenation layer Concat for concatenation to obtain feature map D′; Input the feature map B1, feature map B2 and feature map D′ into the third concatenation layer Concat for concatenation to obtain feature map B′; The feature map A1, the feature map A2 and the feature map B′ are input into the fourth concatenation layer Concat for concatenation to obtain the feature map A′.
9. The binocular ranging method based on the YOLOv8-BiFPN network model according to claim 8, characterized in that: The detection head includes detection head 1, detection head 2, and detection head 3; The sixth C2f module outputs the feature map G′ as the output result of the detection head 1, outputting the target category and the two-dimensional bounding box coordinates; The seventh C2f module outputs the feature map H′ as the output result of the detection head 2, outputting the target category and the two-dimensional bounding box coordinates; The eighth C2f module outputs the feature map F′ as the output result of the detection head 3, and outputs the target category and two-dimensional bounding box coordinates.
10. The binocular ranging method based on the YOLOv8-BiFPN network model according to claim 9, characterized in that: In step 4, the distance of the target object is calculated based on the binocular camera and the trained YOLOv8-BiFPN network model; the specific process is: The binocular camera obtains left and right views; The left view is input into the trained YOLOv8s-SPD network model, and the trained YOLOv8s-SPD network model outputs the target category and two-dimensional bounding box coordinates in the left view; Obtain the disparity map of the left view and the right view based on the SGBM algorithm; According to the disparity map and the known binocular camera parameters, the depth value corresponding to each pixel in the disparity map is calculated to generate a depth map; The distance of the target object is calculated based on the depth map, the target category and the two-dimensional bounding box coordinates in the left view output by the trained YOLOv8s-SPD network model.