Vehicle width detection method, device and equipment based on Yolov8 and storage medium

By introducing the front and rear detection head and specific network structure optimization in the Yolov8 model, the problem of inaccurate distance measurement in the existing technology is solved, and the safety and reliability of the autonomous driving system is significantly improved.

CN120107331APending Publication Date: 2025-06-06ZHIZI AUTOMOTIVE TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510453032.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-04-10
Filing Date
2025-04-11
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In the prior art, the distance is measured only by the peripheral frame of the target vehicle, which has low accuracy and is prone to errors, which leads to safety hazards.

Method used

Using the Yolov8-based vehicle width detection method, the Yolov8 model is optimized to obtain the width information of the target vehicle by adding the output of p2 layer in the Neck network and replacing the C2f module with the C2f-Att module, and introducing the head-detect in the head network, including the width proportional relationship between the front and rear frames and the 2D detection frame and the sigmoid activation function.

Benefits of technology

It significantly improves the accuracy of distance measurement and enhances the safety and reliability of the autonomous driving system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107331A_ABST
    Figure CN120107331A_ABST
Patent Text Reader

Abstract

The invention discloses a vehicle width detection method, device and equipment based on Yolov8 and a storage medium, relates to the technical field of target detection, and can solve the problems that in the prior art, the accuracy is low and potential safety hazards are likely to be caused when the distance is measured only through a peripheral frame of a target vehicle. According to the specific technical scheme, firstly, vehicle image data are collected to construct a training data set; then optimizing the Yolov8 model, including adding a p2 layer, replacing C2f with C2f-Att, and introducing a vehicle head and vehicle tail detection head; training the optimized model by adopting a training data set; and finally inputting the target vehicle picture into the model to obtain width information of the target vehicle. According to the invention, the Yolov8 model outputs the 2D detection frame and the headstock and tailstock frame at the same time, and the width information of the target vehicle can be accurately obtained by means of the headstock and tailstock frame, so that the precision of distance measurement is significantly improved, and the safety and reliability of an automatic driving system are further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target detection, and in particular to a vehicle width detection method, device, equipment and storage medium based on Yolov8. Background Art

[0002] With the acceleration of urbanization and the rapid development of transportation, vehicles have become an essential tool for every household. However, the continuous growth of traffic flow and vehicle density has also brought about problems such as traffic congestion and long driving. In order to reduce driving fatigue, industry and academia are gradually developing assisted driving systems. With the rapid development of deep learning, autonomous driving technology has also received widespread attention from academia and industry. For autonomous driving vehicles, it is crucial to accurately identify the vehicle in front and obtain the distance between it and the vehicle, which is not only related to driving safety, but also of great significance to pedestrian protection.

[0003] At present, existing target detection methods usually rely on the outer frame of the vehicle for detection. However, measuring the distance only by the outer frame of the target vehicle has low accuracy and is prone to errors, which in turn causes safety hazards. Therefore, how to detect the width information of the target vehicle to improve the accuracy of distance measurement has become one of the key issues that need to be solved in current autonomous driving technology. Summary of the invention

[0004] The present invention provides a vehicle width detection method, device, equipment and storage medium based on Yolov8, which can solve the problem that the prior art only measures the distance by the outer frame of the target vehicle, which has low accuracy and is prone to errors, thereby causing safety hazards. The technical solution is as follows:

[0005] According to a first aspect of the present invention, a vehicle width detection method based on Yolov8 is provided, the method comprising:

[0006] Collect vehicle image data to construct a training data set, and annotate the image data in the training data set, wherein the annotated content includes 2D detection frame coordinates and front and rear frame coordinates of the vehicle;

[0007] The Yolov8 model was optimized, including: adding the output of the p2 layer in the Neck network, replacing the C2f module with the C2f-Att module, and introducing the head-detect function in the Head network;

[0008] The optimized Yolov8 model is trained using the training data set to output a detection model;

[0009] Inputting the target vehicle image into the detection model to obtain the width information of the target vehicle;

[0010] The head-detect is introduced into the head network structure of the Yolov8 model, including: the width ratio relationship between the head-detect box and the 2D detection box and a sigmoid activation function with an output value between 0 and 1.

[0011] The vehicle width detection method based on Yolov8 provided by the present invention comprises the following steps: firstly, vehicle image data is collected to construct a training data set, and image data in the training data set is annotated, wherein the annotated content includes 2D detection frame coordinates and vehicle head and vehicle tail frame coordinates; the Yolov8 model is optimized, comprising: adding the output of the p2 layer in the Neck network, replacing the C2f module with the C2f-Att module, and introducing the vehicle head and vehicle tail detection head head-detect in the Head network; training the optimized Yolov8 model with the training data set, and outputting the detection model; inputting the target vehicle picture into the detection model, and obtaining the width information of the target vehicle; wherein the vehicle head and vehicle tail detection head head-detect is introduced into the head network structure of the Yolov8 model, comprising: the width ratio relationship between the vehicle head and vehicle tail frame and the 2D detection frame, and a sigmoid activation function with an output value between 0 and 1. The present invention learns the front and rear frames based on the 2D detection frame, so that the Yolov8 model can simultaneously output the 2D detection frame and the front and rear frames. With the help of the front and rear frames, the width information of the target vehicle can be accurately obtained, thereby significantly improving the accuracy of distance measurement and further improving the safety and reliability of the autonomous driving system.

[0012] As a further solution of the present invention: the C2f-Att module uses the EMA module to redistribute feature weights, specifically:

[0013] The image features output by the backbone are divided into two groups. One group enters the two paths of the 1×1 branch to extract horizontal and vertical features; the other group extracts global information through the 3×3 branch.

[0014] The outputs of the two branches are concatenated and cross-channel information is mixed through 1×1 convolution;

[0015] The sigmoid function is used to adjust the attention weights in the two-dimensional binomial distribution to generate an attention map.

[0016] The method of the present invention enables the model to obtain feature information of different directions and scales through a multi-directional feature extraction method, provides richer feature representation for subsequent target detection, and helps to improve the detection capability of targets of different directions and shapes.

[0017] As a further solution of the present invention: the width ratio relationship between the front and rear frames and the 2D detection frame is calculated by the first formula and the second formula:

[0018] scale 1 =(sub_x 1 -x 1 ) / (x 2 -x 1 );

[0019] scale 2 =(sub_x 2 -x 2 ) / (x 2 -x 1 );

[0020] Among them, scale 1 Corresponding relationship between the upper left corner coordinates; scale 2 is the coordinate correspondence of the lower right corner; x 1 is the horizontal coordinate of the upper left corner of the 2D detection box; x 2 is the horizontal coordinate of the lower right corner of the 2D detection box; sub_x 1 is the horizontal coordinate of the upper left corner of the front and rear frame; sub_x 2 It is the horizontal coordinate of the lower right corner of the front and rear frames;

[0021] The sigmoid activation function is calculated by the third formula:

[0022]

[0023] Among them, pred_scale is the ratio of the front and rear of the vehicle to the whole vehicle output by the Head network, and x is the output of the convolutional layer in the Head network.

[0024] The method of the present invention calculates the corresponding relationship between the upper left corner and the lower right corner coordinates through the first formula and the second formula respectively, and can accurately obtain the width ratio relationship between the front and rear frames and the 2D detection frame. This precise ratio relationship helps the model to more accurately locate the positions of the front and rear of the vehicle during the detection process, thereby improving the accuracy of vehicle width detection. In addition, the ratio relationship pred_scale between the front and rear of the vehicle and the whole vehicle output by the Head network is standardized through the third formula, so that the output result has a clear range, which is conducive to the subsequent processing and analysis of the detection results.

[0025] As a further solution of the present invention: the optimized Yolov8 model is trained using the training data set, and the output detection model includes:

[0026] Inputting the training data set into the optimized Yolov8 model;

[0027] The proportional loss and the category loss are calculated respectively to update the model parameters. When the loss value no longer decreases or the number of training rounds reaches the set number of training rounds, the training is stopped and the detection model is output.

[0028] The method of the present invention can optimize the performance of the model on different tasks in a targeted manner by calculating the two losses separately and updating the model parameters, thereby ensuring that the model can accurately detect the vehicle width and identify the vehicle category in practical applications.

[0029] As a further solution of the present invention: the proportional loss is calculated by the fourth formula:

[0030]

[0031] Among them, loss_scale is the proportional loss; N is the number of predicted detection boxes; scale is the ratio of the real front and rear of the vehicle to the whole vehicle; pred_scale is the ratio of the front and rear of the vehicle output by the Head network to the whole vehicle;

[0032] The class loss is calculated by the fifth formula:

[0033]

[0034] Where sub_cls_loss is the category loss; N is the number of predicted detection boxes; i is an integer from 1 to N; y i is the label of the front and rear of the vehicle; σ is the sigmoid activation function; p i The labels predicted by the network model.

[0035] The method of the present invention provides clear calculation methods for proportional loss and category loss through the fourth formula and the fifth formula, respectively, so that the model has a clear optimization goal during the training process. This helps to guide the update direction of the model parameters, accelerate the convergence speed of the model, and improve the training efficiency.

[0036] As a further solution of the present invention: the labeling of the image data in the training data set includes:

[0037] Use the labelimg labeling tool to label the image data in the training data set.

[0038] The method of the present invention labels image data through the labelimg labeling tool, so that the labeled data can be seamlessly connected to the training process of the Yolov8 model without complex format conversion, saving time and energy for data preprocessing and improving the efficiency of the entire development process.

[0039] According to a second aspect of the present invention, a vehicle width detection device based on Yolov8 is provided, comprising:

[0040] A construction module is used to collect vehicle image data to construct a training data set, and to annotate the image data in the training data set, wherein the annotated content includes 2D detection frame coordinates and front and rear frame coordinates of the vehicle;

[0041] An optimization module is used to optimize the Yolov8 model, specifically to add the output of the p2 layer in the Neck network, replace the C2f module with the C2f-Att module, and introduce the head-detect in the Head network, including the width ratio relationship between the head-detect box and the 2D detection box and the sigmoid activation function with an output value between 0 and 1;

[0042] A training module, used to train the optimized Yolov8 model using the training data set and output a detection model;

[0043] The detection module is used to input the target vehicle image into the detection model to obtain the width information of the target vehicle.

[0044] The vehicle width detection device based on Yolov8 provided by the present invention comprises a construction module, an optimization module, a training module and a detection module; the construction module collects vehicle image data to construct a training data set, and annotates the image data in the training data set, wherein the annotated content comprises 2D detection frame coordinates and front and rear frame coordinates; the optimization module optimizes the Yolov8 model, comprising: adding the output of the p2 layer in the Neck network, replacing the C2f module with the C2f-Att module, and introducing the front and rear detection head head-detect in the Head network, comprising the width ratio relationship between the front and rear frames and the 2D detection frame and a sigmoid activation function with an output value between 0 and 1; the training module trains the optimized Yolov8 model with the training data set, and outputs a detection model; the detection module inputs a target vehicle picture into the detection model, and obtains the width information of the target vehicle. The present invention learns the front and rear frames based on the 2D detection frame, so that the Yolov8 model can simultaneously output the 2D detection frame and the front and rear frames. With the help of the front and rear frames, the width information of the target vehicle can be accurately obtained, thereby significantly improving the accuracy of distance measurement and further improving the safety and reliability of the autonomous driving system.

[0045] According to a third aspect of the present invention, a vehicle width detection device based on Yolov8 is provided, wherein the vehicle width detection device based on Yolov8 comprises a processor and a memory, wherein the memory stores at least one computer instruction, and the instruction is loaded and executed by the processor to implement the steps performed in any one of the above-mentioned vehicle width detection methods based on Yolov8.

[0046] According to a fourth aspect of the present invention, a computer-readable storage medium is provided, wherein at least one computer instruction is stored in the storage medium, and the instruction is loaded and executed by a processor to implement the steps performed in any of the above-mentioned vehicle width detection methods based on Yolov8.

[0047] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0049] Figure 1 is a flow chart of a vehicle width detection method based on Yolov8 provided in an embodiment of the present invention;

[0050] Figure 2 It is a schematic diagram of the relationship between a 2D detection frame and a front and rear frame in a vehicle width detection method based on Yolov8 provided in an embodiment of the present invention;

[0051] Figure 3 It is a network structure of optimized Yolov8 in a vehicle width detection method based on Yolov8 provided in an embodiment of the present invention;

[0052] Figure 4 It is a structure of an efficient multi-scale attention mechanism in a vehicle width detection method based on Yolov8 provided by an embodiment of the present invention;

[0053] Figure 5 It is a structure of C2f-Att in a vehicle width detection method based on Yolov8 provided in an embodiment of the present invention;

[0054] Figure 6 It is a structural diagram of a vehicle width detection device based on Yolov8 provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0055] Here, exemplary embodiments will be described in detail, and examples thereof are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention.

[0056] The embodiment of the present invention provides a vehicle width detection method based on Yolov8, such as Figure 1 As shown, the vehicle width detection method based on Yolov8 includes the following steps:

[0057] Step 101: Collect vehicle image data to construct a training data set, and annotate the image data in the training data set, where the annotation content includes 2D detection frame coordinates and front and rear frame coordinates;

[0058] In this embodiment, Figure 2 As shown in the figure, the red box is the 2D detection box, which contains all the features of the vehicle. The green box is the front and rear box, which contains the information of the front and rear of the vehicle. Specifically, x 1 ,y 1 is the horizontal and vertical coordinates of the upper left corner of the 2D detection box; x 2 ,y 2 sub_x is the horizontal and vertical coordinates of the lower right corner of the 2D detection box; 1 、sub_y 1 sub_x is the horizontal and vertical coordinates of the upper left corner of the front and rear frames; 2 、sub_y 2 These are the horizontal and vertical coordinates of the lower right corner of the front and rear frames of the vehicle.

[0059] In one embodiment, labeling image data in a training data set includes:

[0060] Use the labelimg annotation tool to annotate the image data in the training dataset.

[0061] Specifically, Labelimg is an open source image annotation tool that supports the drawing of shapes such as rectangles and polygons. In this embodiment, it is used to prepare training data sets for deep learning models.

[0062] The method of the present invention labels image data through the labelimg labeling tool, so that the labeled data can be seamlessly connected to the training process of the Yolov8 model without complex format conversion, saving time and energy for data preprocessing and improving the efficiency of the entire development process.

[0063] Step 102, optimizing the Yolov8 model, including: adding the output of the p2 layer in the Neck network, replacing the C2f module with the C2f-Att module, and introducing the head-detect in the Head network;

[0064] Among them, the head-detection head of the vehicle head is introduced into the head network structure of the Yolov8 model, including: the width ratio relationship between the vehicle head and tail box and the 2D detection box and the sigmoid activation function with an output value between 0 and 1.

[0065] In actual use, such as Figure 2 As shown, it can be seen that the front and rear frames must be included in the 2D detection frame, and the front and rear frames can be considered as sub-frames of the 2D detection frame.

[0066] The method of the present invention can accurately locate the positions of the front and rear of the vehicle by introducing head-detect, thereby providing richer information for the accurate calculation of the vehicle width and significantly improving the accuracy of vehicle detection. In addition, the use of a sigmoid activation function with an output value between 0 and 1 can limit the output results of the model to a reasonable range, which helps to improve the stability and reliability of the entire detection system.

[0067] like Figure 3 As shown, in this embodiment, under the network architecture of Yolov8, the p2 layer network result is added to the Neck part. The normal YOLOv8 object detection model output layers are P3, P4, and P5. In order to improve the detection capability of small targets, a P2 layer is added to the network structure. The P2 layer performs fewer convolutions, and the size (resolution) of the feature map is larger, which is more conducive to small target recognition. In addition, the C2f-EMA module in the Neck layer enhances feature extraction by reallocating feature weights using the EMA attention mechanism. This mechanism gives priority to relevant features and spatial details of different channels in the image, such as Figure 4 , Figure 5 As shown in Figure 1, the input features are divided into groups, processed by parallel subnetworks, and integrated with advanced aggregation techniques. This enhancement significantly improves the detection capabilities of small and difficult objects and improves the efficiency of the neck network.

[0068] In one embodiment, the C2f-Att module uses the EMA module to redistribute feature weights, specifically:

[0069] The image features output by the backbone are divided into two groups. One group enters the two paths of the 1×1 branch to extract horizontal and vertical features, and the other group extracts global information through the 3×3 branch.

[0070] The outputs of the two branches are concatenated and cross-channel information is mixed through 1×1 convolution;

[0071] The sigmoid function is used to adjust the attention weights in the two-dimensional binomial distribution to generate an attention map.

[0072] Specifically, the EMA module uses a parallel subnetwork approach to effectively capture multi-scale spatial information and cross-channel dependencies. It has two parallel branches: a 1x1 branch with two routes and a 3x3 branch with one route. In the 1x1 branch, each route uses a one-dimensional global average pooling to encode channel information along the horizontal and vertical spatial directions. These operations produce two encoded feature vectors representing global information, which are then concatenated along the height direction. A 1x1 convolutional layer is subsequently applied to the concatenated output to maintain channel integrity and capture cross-channel interactions by mixing information between different channels. The output is split into two vectors, and a nonlinear Sigmoid function is used to adjust the attention weights in a two-dimensional binomial distribution. The channel-wise attention maps are then combined by multiplication within each group to enhance cross-channel interaction capabilities.

[0073] The method of the present invention enables the model to obtain feature information of different directions and scales through a multi-directional feature extraction method, provides richer feature representation for subsequent target detection, and helps to improve the detection capability of targets of different directions and shapes.

[0074] In one embodiment, the width ratio between the front and rear frames and the 2D detection frame is calculated using the first formula and the second formula:

[0075] scale 1 =(sub_x 1 -x 1 ) / (x 2 -x 1 );

[0076] scale 2 =(sub_x 2 -x 2 ) / (x 2 -x 1 );

[0077] Among them, scale 1 Corresponding relationship between the upper left corner coordinates; scale 2 is the coordinate correspondence of the lower right corner; x 1 is the horizontal coordinate of the upper left corner of the 2D detection box; x 2 is the horizontal coordinate of the lower right corner of the 2D detection box; sub_x 1 is the horizontal coordinate of the upper left corner of the front and rear frame; sub_x 2 It is the horizontal coordinate of the lower right corner of the front and rear frames;

[0078] The sigmoid activation function is calculated by the third formula, which includes:

[0079]

[0080] Among them, pred_scale is the ratio of the front and rear of the vehicle to the whole vehicle output by the Head network, and x is the output of the convolutional layer in the Head network.

[0081] In actual use, according to the first formula, the second formula and the positional relationship of the two boxes, another constraint can be obtained: scale 1 、scale 2 are all between 0 and 1. When the network is input to the head-detect layer, due to the constraint of scale, it needs to be constrained between 0 and 1. Therefore, the sigmoid activation function designed in the network layer can limit the output.

[0082] The method of the present invention calculates the corresponding relationship between the upper left corner and the lower right corner coordinates through the first formula and the second formula respectively, and can accurately obtain the width ratio relationship between the front and rear frames and the 2D detection frame. This precise ratio relationship helps the model to more accurately locate the positions of the front and rear of the vehicle during the detection process, thereby improving the accuracy of vehicle width detection. In addition, the ratio relationship pred_scale between the front and rear of the vehicle and the whole vehicle output by the Head network is standardized through the third formula, so that the output result has a clear range, which is conducive to the subsequent processing and analysis of the detection results.

[0083] Step 103: Use the training data set to train the optimized Yolov8 model and output the detection model.

[0084] In one embodiment, the optimized Yolov8 model is trained using a training data set, and the output detection model includes:

[0085] Input the training data set into the optimized Yolov8 model;

[0086] The proportional loss and category loss are calculated separately to update the model parameters. When the loss value no longer decreases or the number of training rounds reaches the set number of training rounds, the training is stopped and the detection model is output.

[0087] In this embodiment, since the true value and the predicted value of the front-to-rear ratio are both between 0 and 1, the mean square error (MSE) is selected as the loss function of the front-to-rear ratio relationship, and the binary cross entropy loss is selected as the category loss function.

[0088] The method of the present invention can optimize the performance of the model on different tasks in a targeted manner by calculating the two losses separately and updating the model parameters, thereby ensuring that the model can accurately detect the vehicle width and identify the vehicle category in practical applications.

[0089] In one embodiment, the proportional loss is calculated by a fourth formula, which includes:

[0090]

[0091] Among them, loss_scale is the proportional loss; N is the number of predicted detection boxes; scale is the ratio of the real front and rear of the vehicle to the whole vehicle; pred_scale is the ratio of the front and rear of the vehicle output by the Head network to the whole vehicle;

[0092] The class loss is calculated by the fifth formula. The fourth formula includes:

[0093]

[0094] Where sub_cls_loss is the category loss; N is the number of predicted detection boxes; i is an integer from 1 to N; y i is the label of the front and rear of the vehicle; σ is the sigmoid activation function; p i The labels predicted by the network model.

[0095] The method of the present invention provides clear calculation methods for proportional loss and category loss through the fourth formula and the fifth formula, respectively, so that the model has a clear optimization goal during the training process. This helps to guide the update direction of the model parameters, accelerate the convergence speed of the model, and improve the training efficiency.

[0096] Step 104: Input the target vehicle image into the detection model to obtain the width information of the target vehicle.

[0097] The vehicle width detection method based on Yolov8 provided by an embodiment of the present invention first collects vehicle image data to construct a training data set, and annotates the image data in the training data set, wherein the annotated content includes 2D detection frame coordinates and front and rear frame coordinates; the Yolov8 model is optimized, including: adding the output of the p2 layer in the Neck network, replacing the C2f module with the C2f-Att module, and introducing the front and rear detection head head-detect in the Head network; using the training data set to train the optimized Yolov8 model and output a detection model; inputting a target vehicle image into the detection model to obtain the width information of the target vehicle; wherein the front and rear detection head head-detect is introduced into the head network structure of the Yolov8 model, including: the width ratio relationship between the front and rear frames and the 2D detection frame and a sigmoid activation function with an output value between 0 and 1. The present invention learns the front and rear frames based on the 2D detection frame, so that the Yolov8 model can simultaneously output the 2D detection frame and the front and rear frames. With the help of the front and rear frames, the width information of the target vehicle can be accurately obtained, thereby significantly improving the accuracy of distance measurement and further improving the safety and reliability of the autonomous driving system.

[0098] Based on the above Figure 1 The vehicle width detection method based on Yolov8 described in the corresponding embodiment is as follows: an embodiment of the device of the present invention, which can be used to execute the embodiment of the method of the present invention.

[0099] The embodiment of the present invention provides a vehicle width detection device based on Yolov8, such as Figure 6 As shown, the device comprises:

[0100] The construction module 201 is used to collect vehicle image data to construct a training data set, and annotate the image data in the training data set, wherein the annotated content includes 2D detection frame coordinates and front and rear frame coordinates of the vehicle;

[0101] The optimization module 202 is used to optimize the Yolov8 model, specifically to add the output of the p2 layer in the Neck network, replace the C2f module with the C2f-Att module, and introduce the head-detect in the Head network, including the width ratio relationship between the head-detect box and the 2D detection box and the sigmoid activation function with an output value between 0 and 1;

[0102] A training module 203 is used to train the optimized Yolov8 model using a training data set and output a detection model;

[0103] The detection module 204 is used to input the target vehicle image into the detection model to obtain the width information of the target vehicle.

[0104] The vehicle width detection device based on Yolov8 provided by the embodiment of the present invention includes a construction module 201, an optimization module 202, a training module 203 and a detection module 204; the construction module 201 collects vehicle image data to construct a training data set, and annotates the image data in the training data set, and the annotation content includes 2D detection frame coordinates and front and rear frame coordinates; the optimization module 202 optimizes the Yolov8 model, including: adding the output of the p2 layer in the Neck network, replacing the C2f module with the C2f-Att module, and introducing the front and rear detection head head-detect in the Head network, including the width ratio relationship between the front and rear frames and the 2D detection frame and the sigmoid activation function with an output value between 0 and 1; the training module 203 uses the training data set to train the optimized Yolov8 model and outputs a detection model; the detection module 204 inputs the target vehicle picture into the detection model to obtain the width information of the target vehicle. The present invention learns the front and rear frames based on the 2D detection frame, so that the Yolov8 model can simultaneously output the 2D detection frame and the front and rear frames. With the help of the front and rear frames, the width information of the target vehicle can be accurately obtained, thereby significantly improving the accuracy of distance measurement and further improving the safety and reliability of the autonomous driving system.

[0105] Based on the above Figure 1 In the corresponding embodiment, the vehicle width detection method based on Yolov8 is described. Another embodiment of the present invention further provides a vehicle width detection device based on Yolov8. The vehicle width detection device based on Yolov8 includes a processor and a memory. The memory stores at least one computer instruction, which is loaded and executed by the processor to implement the above Figure 1 The vehicle width detection method based on Yolov8 described in the corresponding embodiment.

[0106] Based on the above Figure 1 In the corresponding embodiment of the vehicle width detection method based on Yolov8, the embodiment of the present invention further provides a computer-readable storage medium. For example, the non-temporary computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. The storage medium stores at least one computer instruction for executing the above Figure 1 The vehicle width detection method based on Yolov8 described in the corresponding embodiment will not be repeated here.

[0107] Those skilled in the art will readily appreciate other embodiments of the present invention after considering the specification and practicing the disclosure disclosed herein. This application is intended to cover any variations, uses or adaptations of the present invention that follow the general principles of the present invention and include common knowledge or customary techniques in the art that are not disclosed by the present invention. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present invention are indicated by the claims.

[0108] It should be understood that the present invention is not limited to the exact construction that has been described above and shown in the drawings and that various modifications and changes may be made without departing from the scope thereof. The scope of the present invention is limited only by the appended claims.

Claims

1. A vehicle width detection method based on Yolov8, characterized in that: The method comprises: Collect vehicle image data to construct a training data set, and annotate the image data in the training data set, wherein the annotated content includes 2D detection frame coordinates and front and rear frame coordinates of the vehicle; The Yolov8 model was optimized, including: adding the output of the p2 layer in the Neck network, replacing the C2f module with the C2f-Att module, and introducing the head-detect function in the Head network; The optimized Yolov8 model is trained using the training data set to output a detection model; Inputting the target vehicle image into the detection model to obtain the width information of the target vehicle; The head-detect is introduced into the head network structure of the Yolov8 model, including: the width ratio relationship between the head-detect box and the 2D detection box and a sigmoid activation function with an output value between 0 and 1.

2. The vehicle width detection method based on Yolov8 according to claim 1, characterized in that: The C2f-Att module uses the EMA module to redistribute feature weights, specifically: The image features output by the backbone are divided into two groups. One group enters the two paths of the 1×1 branch to extract horizontal and vertical features, and the other group extracts global information through the 3×3 branch. The outputs of the two branches are concatenated and cross-channel information is mixed through 1×1 convolution; The sigmoid function is used to adjust the attention weights in the two-dimensional binomial distribution to generate an attention map.

3. The vehicle width detection method based on Yolov8 according to claim 1, characterized in that: The width ratio between the front and rear frames and the 2D detection frame is calculated using the first and second formulas: scale1=(sub_x1-x1) / (x2-x1); scale2=(sub_x2-x2) / (x2-x1); Among them, scale1 is the coordinate correspondence of the upper left corner; scale2 is the coordinate correspondence of the lower right corner; x1 is the horizontal coordinate of the upper left corner of the 2D detection frame; x2 is the horizontal coordinate of the lower right corner of the 2D detection frame; sub_x1 is the horizontal coordinate of the upper left corner of the front and rear frames; sub_x2 is the horizontal coordinate of the lower right corner of the front and rear frames; The sigmoid activation function is calculated by a third formula, which includes: Among them, pred_scale is the ratio of the front and rear of the vehicle to the whole vehicle output by the Head network, and x is the output of the convolutional layer in the Head network.

4. The vehicle width detection method based on Yolov8 according to claim 1, characterized in that: The optimized Yolov8 model is trained using the training data set, and the output detection model includes: Inputting the training data set into the optimized Yolov8 model; The proportional loss and the category loss are calculated respectively to update the model parameters. When the loss value no longer decreases or the number of training rounds reaches the set number of training rounds, the training is stopped and the detection model is output.

5. The vehicle width detection method based on Yolov8 according to claim 4 is characterized in that: The proportional loss is calculated by a fourth formula, which includes: Among them, loss_scale is the proportional loss; N is the number of predicted detection boxes; scale is the ratio of the real front and rear of the vehicle to the whole vehicle; pred_scale is the ratio of the front and rear of the vehicle output by the Head network to the whole vehicle; The class loss is calculated by the fifth formula, and the fourth formula includes: Where sub_cls_loss is the category loss; N is the number of predicted detection boxes; i is an integer from 1 to N; y i is the label of the front and rear of the vehicle; σ is the sigmoid activation function; p i The labels predicted by the network model.

6. The vehicle width detection method based on Yolov8 according to claim 1, characterized in that: The labeling of the image data in the training data set includes: Use the labelimg labeling tool to label the image data in the training data set.

7. A vehicle width detection device based on Yolov8, characterized in that: include: A construction module is used to collect vehicle image data to construct a training data set, and to annotate the image data in the training data set, wherein the annotated content includes 2D detection frame coordinates and front and rear frame coordinates of the vehicle; The optimization module is used to optimize the Yolov8 model. Specifically, it is used to add the output of the p2 layer in the Neck network, replace the C2f module with the C2f-Att module, and introduce the head-detect in the Head network; A training module, used to train the optimized Yolov8 model using the training data set and output a detection model; The detection module is used to input the target vehicle image into the detection model to obtain the width information of the target vehicle.

8. A vehicle width detection device based on Yolov8, characterized in that: The Yolov8-based vehicle width detection device includes a processor and a memory, wherein the memory stores at least one computer instruction, and the instruction is loaded and executed by the processor to implement the steps performed in the Yolov8-based vehicle width detection method described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The storage medium stores at least one computer instruction, which is loaded and executed by the processor to implement the steps performed in the vehicle width detection method based on Yolov8 as described in any one of claims 1 to 6.

Citation Information

Cited By

  • Small target identification method and system based on cascade hierarchy detection and self-comparison

    CN120356125A

  • A small target recognition method and system based on cascade hierarchical detection and self-comparison

    CN120356125B