Target detection method based on deep learning
By improving the YOLOv3 model, jump connection, multi-scale detection and optimization loss functions are introduced, which solves the problem of low accuracy and recall of YOLOv1 and YOLOv2 in small object detection, and achieves efficient small object detection in complex environments.
Patent Information
- Application Number
- CN202510356248.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-25
AI Technical Summary
When YOLOv1 and YOLOv2 detect small targets, there are problems with low accuracy and recall, especially in tasks with high discrimination accuracy requirements, which cannot effectively detect small targets.
The improved YOLOv3 model is adopted, and feature fusion is enhanced by introducing jump connections, a multi-scale detection mechanism is adopted, and the positive and negative sample imbalance problem is solved by using Focal Loss, and the coordinate loss is weighted, the network structure and loss function are optimized, and the object detection and recognition is combined with a non-maximum suppression algorithm.
The accuracy and accuracy of small target detection are improved, especially the detection ability of small targets in complex environments, achieving efficient and accurate target detection and recognition.
Smart Images

Figure CN120375041A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of object detection, and specifically to an object detection method based on deep learning. Background Art
[0002] In recent years, the YOLO (You Only Look Once) series of algorithms have received more attention from researchers due to their faster speed and superior object detection effect. The YOLO algorithm transforms object detection into a regression problem and can obtain the detection result through one forward propagation, so its speed is very fast. At the same time, this method is end-to-end training, and the measurement of the entire process is relatively small, achieving the detection purpose.
[0003] When detecting small objects, since small objects are objects with relatively small length and width, and the pixels of small objects are very few in the image, they are prone to being confused with the background, and the position accuracy requirement is also very high. Therefore, the comparison accuracy and recall rate of small object detection in YOLOv1 and YOLOv2 are lower than those of large objects. Especially for tasks with higher discriminant accuracy requirements and higher detection accuracy requirements, they are even more unable to handle. Summary of the Invention
[0004] In view of this, the present invention provides an object detection method based on deep learning, which can improve the detection accuracy of small objects.
[0005] An object detection method based on deep learning includes the following steps:
[0006] Step 1, Image Acquisition: Obtain a real-time or pre-stored image of the scene to be detected through an image acquisition device;
[0007] Step 2, Image Input: Input the image obtained in Step 1 into an improved YOLOv3 model;
[0008] Step 3, Object Classification: Use the improved YOLOv3 model to classify the objects to be detected in the image into preset categories, including the first type of fixed object and the second type of moving object;
[0009] Step 4, Safety Distance Judgment: Based on the classification result, judge whether the distance between the first type of fixed object and the second type of moving object in the image meets the preset safety distance threshold;
[0010] Step 5, Information Extraction: For the second type of moving object that meets the preset safety distance threshold, extract the position information and shape information of the second type of moving object in the image;
[0011] Step 6, Object Detection and Recognition: Based on the position information and shape information of the second type of moving object in the image, identify the identity information of the object.
[0012] Further, in step 2, the image is preprocessed before input, and the preprocessing includes normalization, data augmentation, and resizing the image. Among them, the normalization process scales the image pixel values to a preset range; data augmentation randomly rotates, flips, crops, and scales the image; resizing the image adjusts the image size to be consistent with the input requirements of the YOLOv3 model.
[0013] Further, the position information includes the bounding box coordinates, and the shape information includes the size and aspect ratio.
[0014] Further, the improvements of the improved YOLOv3 model on the YOLOv3 model specifically include:
[0015] Introduce skip connections in the Darknet-53 network to fuse shallow feature maps with deep feature maps and enhance the small object detection ability;
[0016] Adopt a multi-scale detection mechanism to predict anchor boxes on three feature maps of 13×13, 26×26, and 52×52 respectively. Each feature map corresponds to anchor boxes of different sizes, and the sizes of the anchor boxes are set according to the target statistical information of the training dataset;
[0017] Loss function optimization: Introduce Focal Loss to solve the problem of unbalanced positive and negative samples, and weight the coordinate loss to improve the small object detection accuracy.
[0018] Further, the judgment method for the safety distance threshold in step 4 is as follows:
[0019] According to the bounding box coordinates of the first type of fixed object and the second type of moving object in the image, calculate the Euclidean distance between them;
[0020] If the distance is less than the preset threshold, it is determined to be an unsafe state, otherwise it is determined to be a safe state.
[0021] Further, the training method of the improved YOLOv3 model includes:
[0022] Use the buoy image dataset for training. The dataset contains buoy images under different weather and lighting conditions and is divided into training set, test set, and validation set;
[0023] Adopt the PyTorch framework and set the unified learning rate, batch size, and number of iterations;
[0024] In the pre-training stage, load the pre-trained weights of Darknet-53 and monitor the loss function and validation set performance metrics during the training process.
[0025] Further, the annotation format of the buoy image dataset is a.txt file in the YOLO format, which contains the upper-left and lower-right coordinates of the buoy bounding box and the class information.
[0026] Further, the evaluation metrics of the improved YOLOv3 model include:
[0027] Precision, Recall, F1-score, and mean average precision (mAP);
[0028] The mean average precision (mAP) is calculated from the data in the test set that was not involved in training, and is used to verify the model's ability to detect small targets.
[0029] Further, the identification of the second type of moving object includes:
[0030] Based on the bounding box coordinates and shape information, determine the specific type through the class confidence level output by the improved YOLOv3 model;
[0031] Combine the non-maximum suppression (NMS) algorithm to remove redundant detection boxes and retain the detection result with the highest confidence level.
[0032] The present invention proposes an object detection method based on deep learning. This algorithm takes improving the detection accuracy of small targets as the main research object, and realizes the enhancement of the multi-scale detection mechanism, the optimization of the network structure, the improvement of the loss function and other technical means to improve the ability to detect small targets in complex environments on the basis of high detection accuracy. The multi-scale detection mechanism, network structure optimization and loss function in the YOLOv3 algorithm are mainly reflected in the steps of feature extraction, multi-scale detection and object classification and recognition after the image input. These steps together constitute the core framework of the YOLOv3 algorithm, realizing efficient and accurate object detection and recognition, and providing the algorithm of the corresponding model. Through computer simulation, it is found that this algorithm still has a high detection accuracy for small targets in complex environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 It is a multi-scale detection mechanism diagram;
[0034] Figure 2 It is the optimized network structure diagram of the embodiment of the present invention;
[0035] Figure 3 It is the overall design diagram of an object detection method based on deep learning according to an embodiment of the present invention;
[0036] Figure 4 It is the operation architecture schematic diagram of the embodiment of the present invention;
[0037] Figure 5Training diagram of the dataset in the embodiment of the present invention;
[0038] Figure 6 Prediction process diagram of the embodiment of the present invention;
[0039] Figure 7 Detection result diagram of small targets under hydrology in complex sea conditions by the trained model in the embodiment of the present invention;
[0040] Figure 8 Detection result diagram of small targets under illumination in complex sea conditions by the trained model in the embodiment of the present invention;
[0041] Figure 9 Detection result diagram of small targets under the sky background in complex sea conditions by the trained model in the embodiment of the present invention;
[0042] Figure 10 Confidence calculation diagram of the embodiment of the present invention, indicating how many of the detected positive classes are truly positive classes;
[0043] Figure 11 Harmonic mean calculation diagram of precision and recall in the embodiment of the present invention. The higher the F1 value in the figure, the better the model performs in both the precision and recall dimensions;
[0044] Figure 12 Mean average precision calculation diagram of the embodiment of the present invention, centrally representing the accuracy and recall metrics of detection.
[0045] Figure 13 Flowchart of an object detection method based on deep learning in the embodiment of the present invention. Detailed implementation manners
[0046] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0047] Please refer to Figure 13 , the embodiment of the present invention provides an object detection method based on deep learning, including the following steps:
[0048] Step 1, Image acquisition: First, an image of the object to be detected in the current scene is acquired through a camera or other image acquisition devices. This step is the basis of the object detection process, ensuring that the algorithm can process real-time or pre-stored image data.
[0049] Step 2, Image Input and Preprocessing: Input the acquired image into the improved YOLOv3 model. Before input, certain preprocessing can be performed on the image, such as normalization, data augmentation, resizing the image, etc., to meet the requirements of the model input.
[0050] Step 3, Object Classification: Use the improved YOLOv3 model to classify the objects to be detected in the image, and classify the objects to be detected in the image into preset categories. The present invention adds feature fusion, which is mainly reflected in the basic network feature extraction. The basic network feature extraction is mainly reflected in the feature extraction network step. Specifically, YOLOv3 uses a deep convolutional neural network (such as Darknet-53) as the backbone network to extract features from the input image. These features contain hierarchical information of the image, which helps subsequent object detection and recognition. Darknet-53 consists of multiple convolutional layers, pooling layers, and skip connections. By continuously using operations such as multiple convolutional and pooling operations, feature maps with richer semantic meanings can be obtained, and these feature maps will be used for subsequent multi-scale prediction and object detection. Efficient and accurate object detection and recognition are achieved. YOLOv3 is implemented to combine shallow features with deep features using skip connections. This feature fusion method enhances the model's ability to detect small objects. Shallow feature maps have higher resolution and rich detail information, suitable for detecting small objects; while deep feature maps have stronger semantic information, suitable for detecting large objects. Through feature fusion, YOLOv3 can utilize the advantages of both shallow and deep features simultaneously to improve the detection accuracy of small targets.
[0051] The present invention adds the output of multi-scale feature maps. The output of multi-scale feature maps in the algorithm is reflected in: after using the Darknet-53 network to extract features, through convolutional and upsampling operations at different levels, 3 different-scale feature maps (13*13, 26*26, 52*52) are generated, and each grid on each feature map predicts 3 anchor boxes for detecting targets of different sizes. YOLOv3 is implemented to use Darknet-53 as the backbone network, and this network outputs multiple-scale feature maps at different levels. Specifically, YOLOv3 outputs detection results from three different scales (13x13, 26x26, 52x52), corresponding to different levels of the network respectively. This design of multi-scale output enables YOLOv3 to detect objects of different sizes at different scales, especially for the detection of small targets.
[0052] Step 4, Safety distance judgment: Based on the classification results, further determine whether the distance between the first type of fixed object and the second type of moving object in the image meets the preset safety distance threshold. This safety distance is preset by engineers according to the actual application scenario and is used to evaluate whether the relative position relationship between objects is safe. The algorithm for safety distance judgment aims to compare the actual distance (such as Euclidean distance) between targets with the preset safety distance threshold to determine whether there is a violation of the safety distance.
[0053] Step 5, Information extraction: For the second type of moving object that meets the safety distance condition, extract the position information (such as bounding box coordinates) and shape information (such as size, aspect ratio, etc.) of the second type of moving object in the image. The position information and shape information are crucial for subsequent object detection tasks. The present invention adds fine-grained bounding box prediction, which is implemented in YOLOv3 by using a larger number of anchor boxes and combining a more refined bounding box prediction method to more accurately locate the position and shape of small targets. YOLOv3 adopts a more refined loss calculation method, calculating the position loss (including center point coordinates and width and height), confidence loss, and class loss respectively. Secondly, to address the problem of difficult detection of small objects, YOLOv3 weights the regression loss of small objects in the position loss, improving the detection accuracy of small objects.
[0054] Step 6, Object detection and recognition: Based on the position information and shape information of the second type of moving object, identify the identity information of the object. This step not only confirms the existence of the object but also further identifies the identity information of the object, that is, which specific type of moving object it is.
[0055] In summary, the application of the present invention in object detection realizes the rapid and accurate detection and recognition of different types of objects in the scene through a series of core steps such as image acquisition, classification, safety distance judgment, information extraction, and object detection and recognition.
[0056] In addition, in the YOLOv3 algorithm, there is no clear concept of the first type of fixed object and the second type of moving object. The types and quantities of objects that the algorithm can detect depend on the quality and diversity of the training dataset.
[0057] Verify the YOLOv3 object detection algorithm proposed by the present invention. In the simulation experiment, select the PyTorch framework experimental platform as a server with an NVIDIA graphics card. For the YOLOv3 algorithm, based on ensuring experimental fairness, train the algorithm fairly with the same parameters and training strategies (such as learning rate, batch size, number of iterations, etc.).
[0058] Simulation: Object detection model training and detection
[0059] To verify the detection effect of the present invention on sea surface buoys, the improved YOLOv3 model is compared respectively. The required buoy image dataset includes some buoy images obtained under different weather conditions and the corresponding buoy annotations, which are the position and type information of the buoys appearing in the images, so as to divide the buoy dataset into a training set, a test set and a validation set.
[0060] 1. Image acquisition:
[0061] Collect 500 buoy images from a specific upstream data source, screen and identify buoy images from multiple angles and different sizes in terms of illumination and weather, and select the number of screened buoy images to be 500, considering the collation of image numbers before and after filtering.
[0062] Use an annotation tool to annotate all the buoys in each frame of the picture, and generate a corresponding buoy annotation file (such as a.txt file (in yolo format) recording the position coordinate information of the buoy in the picture (the upper left and lower right coordinates of the buoy) and the buoy type (which can also be directly selected in the annotation tool)).
[0063] 2. Image input and preprocessing:
[0064] Preprocess the dataset. Before pre-training the dataset to assist the model in learning more effectively, preprocessing will be carried out first.
[0065] Normalization: Scale or compress the pixel values so that their range is within an acceptable range, reducing the influence of image similarity on illumination or contrast differences.
[0066] Data augmentation: Adopt a series of image augmentation techniques, including operations such as random rotation, flipping, cropping and scaling, so as to expand the training set, increase the diversity of the dataset and make it adapt to more types of scenarios.
[0067] Resize the image: Resize the size of the image to be consistent to adapt to the input of the YOLOv3 model.
[0068] 3. Model configuration:
[0069] The object detection model uses YOLOv3 and corresponding configurations are made for the task. The configuration items include but are not limited to the following:
[0070] (1) Implementation of the multi-scale detection mechanism
[0071] Based on Darknet-53, YOLOv3 adds a multi-scale detection mechanism, and by making predictions on feature maps of different scales, the detection ability for objects of different sizes is improved.
[0072] YOLOv3 optimizes small object detection. It uses the Darknet-53 convolutional network to extract features and makes full use of the different sizes of features at different levels. Therefore, different output feature maps can be obtained. Object detection is performed on the feature maps at different layers by using multiple output layers of different sizes at multiple scales.
[0073] In YOLOv3, multi-scale prediction is efficiently calculated based on three groups of feature maps obtained by downsampling the input image by 32 times, 16 times, and 8 times. On the feature maps at each scale, the network sets a group of anchor boxes, and these anchor boxes are used for localization and class prediction. The number, size, and other information of the anchor boxes at each level are directly set using the statistical information of the target sizes in the training set.
[0074] YOLOv3 realizes the detection of targets of different scales by performing object detection on multi-scale feature maps. Moreover, it also supports detecting multiple targets at once, improving the detection speed while adapting to detection scenarios of different scales.
[0075] Three aspects of the implementation of the YOLOv3 multi-scale detection mechanism:
[0076] Scale transformation: Similar to YOLOv1 and YOLOv2, YOLOv3 does not scale the input image during training. Instead, it uses the computational graph structure and simultaneously uses feature maps at three scales for prediction during inference. These three scales are ml-search[Y1], Y2, and Y3. These scales represent the outputs of three scales of YOLOv3 and are used for detecting objects of different sizes. By simultaneously predicting targets at three different scales, YOLOv3 significantly enhances the model's ability to detect targets of different scales
[18] .
[0077] More abundant anchor boxes: Each scale range of feature maps in YOLOv3 has different numbers of anchor boxes with different sizes and aspect ratios. YOLOv3 designs the anchor boxes according to the target sizes in the training set, so as to more accurately suppress and predict targets of different sizes. The anchor boxes in YOLOv1 and v2 are relatively fixed, and the same anchor boxes are used in multiple scale ranges. The algorithm uses the K-means clustering method to set multiple prior boxes (anchor boxes) for each downsampling scale. These anchor boxes have different sizes and ratios and can cover a wider range of target sizes. When predicting on feature maps at multiple scales, each grid cell on each feature map will predict multiple anchor boxes, thereby improving the detection accuracy and recall rate for targets of different sizes.
[0078] Specifically, YOLOv3 generates a total of 9 types of anchor boxes on feature maps of 3 different scales. These anchor boxes are used to match and predict object bounding boxes during the training and inference phases, achieving precise object localization in the multi-scale detection mechanism. The above optimizations are reflected in the multi-scale detection mechanism as shown in the following figure:
[0079] Figure 1 It is a diagram of the multi-scale detection mechanism;
[0080] YOLOv3 selects the Darknet-53 network as the feature extractor. The Darknet-53 network architecture is complex, has many parameters, and is deep in layers. Therefore, Darknet-53 is more conducive to extracting image features, can extract rich information from images, and thus can perform more precise feature extraction on the images behind the input, which is beneficial to the subsequent detection process.
[0081] Detection speed: Compared with YOLOv2, although the model of YOLOv3 is larger and the computational complexity of the algorithm increases, due to the efficient feature extraction ability of Darknet-53, which is comparable to SSD512, and the object detection part uses an improved multi-scale detection network, it is faster in real-time detection.
[0082] (2) Optimization of the network structure
[0083] YOLOv3 uses Darknet-53 as the feature extraction network, which contains multiple residual blocks and can effectively extract deep features in images. Network structure optimization: For the network structure of YOLOv3, we have further optimized it, including introducing attention mechanisms, using more efficient convolution operations, etc., to improve the detection accuracy and speed of the model.
[0084] Basic composition of the YOLOv3 network structure:
[0085] Feature extraction network (backbone network): YOLOv3 selects Darknet-53 as the feature extraction network. Darknet-53 is composed of a convolutional layer (conV), a batch normalization layer (batch-normalization, BN), and a Leaky ReLU activation layer (activation) stacked and combined. Darknet-53 can continuously perform convolution and pooling operations on the input image to obtain deep and rich features.
[0086] After migrating the anchor into the Darknet-53 network, YOLOv3 concatenates different numbers of detection heads on the feature maps at different positions. The detection heads are responsible for predicting object bounding boxes, confidence levels, and object categories on feature maps of different scales. Inside, multiple convolutional layers further process the feature maps.
[0087] Advantages of Darknet-53 as the backbone:
[0088] High performance: Darknet-53 has demonstrated excellent performance in large-scale image classification tasks such as ImageNet, which fully proves its strong ability in feature extraction and enables YOLOv3 to more accurately identify various types of targets in images
[19] .
[0089] Lightweight: Compared with other deep convolutional neural networks such as ResNet and VGGNet, Darknet-53 has fewer parameters and lower computational complexity. This feature enables YOLOv3 to have a faster detection speed while maintaining high performance and is very suitable for use in resource-constrained environments.
[0090] Easy to train: Darknet-53 effectively improves the training efficiency of the network by adopting techniques such as batch normalization and LeakyReLU. In addition, the multi-scale prediction strategy of YOLOv3 further enhances the training stability and convergence speed of the network, making network training simpler and more efficient. The above optimizations are reflected in the network structure diagram as follows:
[0091] Figure 2 Is the optimized network structure diagram;
[0092] (3) Improvement of the loss function
[0093] The loss function of YOLOv3 consists of three parts: coordinate loss, confidence loss, and classification loss.
[0094] Regarding the improvement of the loss function, the present invention introduces Focal Loss to solve the problem of imbalance between positive and negative samples, enabling the model to pay more attention to difficult-to-separate samples during the training process.
[0095] In addition, the present invention also optimizes the coordinate loss and adopts a more reasonable IoU (Intersection over Union) calculation method to make the matching between the predicted box and the ground truth box more accurate.
[0096] Composition of the YOLOv3 loss function:
[0097] Classification Loss: This is an index used to evaluate the accuracy of the model when predicting the category to which the target belongs. In YOLOv3, the cross-entropy loss function is adopted to calculate this loss to ensure the accuracy of the model in the classification task.
[0098] Coordinate Loss: This loss is used to measure the accuracy of the model in predicting the position of the target bounding box. To ensure that the model can accurately locate the target, the mean squared error loss (MSE) is used to calculate the coordinate loss.
[0099] Confidence Loss: This is a loss used to evaluate the confidence of the model in judging whether the target actually exists in the image, and the cross-entropy loss function is used to calculate it.
[0100] Total Loss: This is calculated by weighted summation of the above losses. During the training process, the goal of optimization is to minimize this total loss, so as to optimize the performance of the model in the object detection task.
[0101] Improvement of the YOLOv3 Loss Function:
[0102] Increase the weight of the xy loss: Increase the weight of the coordinate loss in the YOLOv3 total loss function. The coordinate loss is directly related to the position of the target bounding box, especially for small targets. If the weight of the coordinate loss is too small, the network will optimize the position of the target bounding box relatively less during the training process. Therefore, increasing the weight of this item can make the network pay more attention to the position of the target bounding box and can detect small targets more accurately.
[0103] Adopt the idea of solving the imbalance between positive and negative samples: YOLOv3 adopts Focal Loss to deal with the problem of imbalance between positive and negative samples. Focal Loss reduces the weight of easy-to-classify samples, making the network more sensitive to difficult-to-classify samples, and enhancing the detection ability of the model in complex scenarios to a certain extent.
[0104] Optimization of the confidence loss: To enhance the model's judgment on the existence of the target, the confidence loss can be optimized, which can better distinguish whether there is a target in the image, thus avoiding misjudgments of missed detections and false detections, and further improving the detection accuracy.
[0105] Anchor boxes: Different-sized anchor boxes are selected according to the different sizes and shapes of the buoys in the image, and they are adjusted to the sizes required for the buoy detection task.
[0106] Set the loss function to the default loss function of YOLOv3, including coordinate loss, confidence loss, and classification loss. The operation of the above functions is reflected in the overall design diagram of the algorithm (as Figure 3 shown) and the operation architecture description diagram (as Figure 4 shown).
[0107] 4. Pre-training process:
[0108] After the dataset preprocessing and model setting preparations are completed, pre-training of the NLR and PT embedding models is carried out.
[0109] As Figure 5 shown, the specific training process is as follows;
[0110] 4.1 Load the dataset: Put the preprocessed images and labeled files into the training framework.
[0111] 4.2 Initialize the model: Before starting the training of the YOLOv3 model, select pre-trained weights for initialization.
[0112] Hyperparameters: When training a neural network, in addition to the parameters learned by the model, some hyperparameters are set manually, such as the learning rate, batch size, and number of iterations, etc.
[0113] 4.3 Start training: In the PyTorch training framework, when training reaches 25,000 steps, it can basically reach the minimum loss function at 25,000 times. The overall process is shown in Figure 6 the prediction process diagram shown.
[0114] 4.4 Monitor the training process: Use the tools provided by the training framework to monitor the training process, including checking the changes in the loss function, performance metrics on the validation set, etc.
[0115] 4.5 Save the model: After training is completed, save the model weights and configuration files in a timely manner for subsequent operation.
[0116] 4.6 Run tests, detect small targets under different variables, observe the detection results, and verify the detection ability
[0117] Among them, Figure 7 is the detection result diagram of the small targets under hydrology in complex sea surface conditions by the trained model; Figure 8 is the detection result diagram of the small targets under illumination in complex sea surface conditions by the trained model; Figure 9 is the detection result diagram of the small targets under the sky background in complex sea surface conditions by the trained model.
[0118] 5. Evaluation metrics: Calculate the small target detection accuracy
[0119] When evaluating the performance of object detection, common evaluation criteria are adopted, including Precision, Recall, F1-Score, and mean Average Precision (mAP). These metrics can comprehensively reflect the performance of buoy detection. In particular, mAP can effectively consider the imbalance of input samples in object detection. To verify the detection effects of different pre-trained models, a certain amount of data that has not participated in training is randomly selected from the entire dataset for verification. The detection ability of the detection model is verified through the classification accuracy of small targets, mAP, F1 score, recall rate, etc., so as to correct and compare different pre-trained models and judge the accuracy and feasibility of the final model, where Figure 10 is a confidence calculation graph, indicating how many of the detected positive classes are truly positive classes;
[0120] Figure 11 is a calculation graph of the harmonic mean of precision and recall. The higher the F1 value in the graph, the better the model performs in both the precision and recall dimensions;
[0121] Figure 12 is a calculation graph of mean average precision, centrally representing the detection accuracy and recall rate metrics;
[0122] 6. Instructions for use:
[0123] 6.1. Detecting objects in images:
[0124] Use the command-line tool or Python script of Darknet to load the pre-trained weights and configuration file (such as yolov3.cfg).
[0125] Specify the image file to be detected and run the detection command.
[0126] The program will output the detected object classes, confidence levels, and bounding box information, and generate a detection result image with bounding boxes.
[0127] 6.2. Detecting objects in videos:
[0128] Similar to detecting images, but a video file needs to be specified as the input.
[0129] The program will process the video frame by frame and output the detection results in real time.
[0130] 6.3. Custom dataset training:
[0131] Prepare your own dataset and annotate it in the YOLOv3 format (usually using txt files to record the class and bounding box information of each object).
[0132] Modify the configuration file to adapt to the custom dataset (e.g., modify the number of categories, input image size, etc.).
[0133] Use the training tool of Darknet to train and generate a custom weight file.
[0134] 6.4. Adjust the detection parameters:
[0135] By modifying the configuration file or command-line parameters, some parameters in the detection process can be adjusted, such as thresholds, non-maximum suppression (NMS), etc., to optimize the detection results.
[0136] 6.5. Technical effects: YOLOv3 achieves multi-scale detection by making predictions on feature maps of different scales. Specifically, YOLOv3 connects three detection heads (Heads) after the last three residual blocks of the Darknet-53 network, and each detection head corresponds to a feature map of a scale. Smaller feature maps have a larger receptive field and are suitable for detecting larger objects; while larger feature maps have a smaller receptive field and are suitable for detecting smaller objects. In this way, YOLOv3 can capture the features of target objects at different scales, thereby improving the detection ability for objects of different sizes. As Figure 12 shown, the average detection accuracy mAP value for small targets can reach 96.71%.
[0137] The present invention introduces a multi-scale detection mechanism in YOLOv3. By setting detection heads with multiple resolutions and distributing them at different levels of the network, the effect of capturing and detecting feature maps of various scales is achieved. Specifically, YOLOv3 adopts the idea of multi-scale prediction and divides the network into three branches: Y1, Y2, and Y3, which are respectively responsible for detecting targets of different scales. Finally, comprehensive coverage of multi-scale object detection is realized, thereby improving the ability to detect small targets in complex environments.
[0138] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A target detection method based on deep learning, characterized in that It includes the following steps: Step 1, Image acquisition: Obtain real-time or pre-stored images of the scene to be detected through an image acquisition device; Step 2, Image input: Input the image obtained in Step 1 into the improved YOLOv3 model; Step 3, Target classification: Use the improved YOLOv3 model to classify the objects to be detected in the image into preset categories, including the first type of fixed object and the second type of moving object; Step 4, Safety distance judgment: Based on the classification results, judge whether the distance between the first type of fixed object and the second type of moving object in the image meets the preset safety distance threshold; Step 5, Information extraction: For the second type of moving object that meets the preset safety distance threshold, extract the position information and shape information of the second type of moving object in the image; Step 6, Target detection and recognition: Based on the position information and shape information of the second type of moving object in the image, identify the identity information of the object.
2. The object detection method based on deep learning according to claim 1, characterized in that: In Step 2, the image is preprocessed before input. The preprocessing includes normalization, data augmentation, and resizing the image. Among them, the normalization process scales the image pixel values to a preset range; data augmentation performs random rotation, flipping, cropping, and scaling on the image; Resizing the image is to adjust the image size to be consistent with the input requirements of the YOLOv3 model.
3. The object detection method based on deep learning according to claim 1, characterized in that: The position information includes the bounding box coordinates, and the shape information includes the size and aspect ratio.
4. The object detection method based on deep learning according to claim 1, characterized in that: The improvements of the improved YOLOv3 model on the YOLOv3 model specifically include: Introduce skip connections in the Darknet-53 network to fuse shallow feature maps and deep feature maps, enhancing the small target detection ability; Adopt a multi-scale detection mechanism to predict anchor boxes on three feature maps of 13×13, 26×26, and 52×52 respectively. Each feature map corresponds to anchor boxes of different sizes, and the sizes of the anchor boxes are set according to the target statistical information of the training dataset; Loss function optimization: Introduce FocalLoss to solve the problem of unbalanced positive and negative samples, and weight the coordinate loss to improve the small target detection accuracy.
5. The object detection method based on deep learning according to claim 1, characterized in that: The judgment method of the safety distance threshold in Step 4 is: According to the bounding box coordinates of the first type of fixed object and the second type of moving object in the image, calculate the Euclidean distance between the two; If the distance is less than the preset threshold, it is determined to be an unsafe state, otherwise it is determined to be a safe state.
6. The object detection method based on deep learning according to claim 1, wherein: The training method of the improved YOLOv3 model includes: Use the buoy image dataset for training. The dataset contains buoy images under different weather and lighting conditions and is divided into training set, test set, and validation set; Adopt the PyTorch framework, and set the unified learning rate, batch size, and number of iterations; In the pre-training stage, load the pre-trained weights of Darknet-53, and monitor the loss function and validation set performance metrics during the training process.
7. The object detection method based on deep learning according to claim 6, characterized in that: The annotation format of the buoy image dataset is a.txt file in YOLO format, which contains the upper left and lower right coordinates and category information of the buoy bounding box.
8. The object detection method based on deep learning according to claim 1, characterized in that: The evaluation metrics of the improved YOLOv3 model include: Precision, Recall, F1-score, and mean average precision (mAP); The mean average precision (mAP) is calculated from the data in the test set that was not involved in training, and is used to verify the model's detection ability for small targets.
9. The object detection method based on deep learning according to claim 1, characterized in that: The identity recognition of the second type of moving object includes: Based on the bounding box coordinates and shape information, the specific type is determined by the class confidence output by the improved YOLOv3 model; Combined with the non-maximum suppression (NMS) algorithm to remove redundant detection boxes and retain the detection result with the highest confidence.