A large-scale multi-target detection method for open-pit mine dumps

By improving Yolov8n network architecture and data enhancement technology, a multi-objective detection model is built, which solves the problems of low recognition accuracy and high model complexity of small and medium-sized targets in open-pit mine excretion sites, and realizes lightweight real-time detection on edge devices to adapt to severe weather conditions and ensure safety monitoring.

CN120411913BActive Publication Date: 2025-09-02TAIYUAN UNIVERSITY OF TECHNOLOGY +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510898687.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-09-02
Estimated Expiration
2045-07-01

AI Technical Summary

Technical Problem

The prior art cannot accurately identify small targets in open-pit mine soil discharge sites. For example, staff members, especially in severe weather conditions, target detection accuracy is greatly reduced, and the existing network model has large parameters and complex calculations, making it difficult to deploy on edge devices in real time.

Method used

Using the improved Yolov8n network architecture, a multi-objective network model is built by replacing modules and adding small object detection layers, combined with data enhancement technology, and deploying it to edge devices for real-time detection.

Benefits of technology

The target detection accuracy is improved, the number of parameters and model size is reduced, so that the model is lightly deployed on edge devices, real-time detection is realized, adapting to severe weather conditions, and ensuring real-time monitoring and protection of security personnel.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120411913B_ABST
    Figure CN120411913B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of target detection technology, and aims to solve the current problem of inaccurate detection of small targets in a large environment. A large-scale multi-target detection method for an open-pit mine spoil dump is provided, comprising the following steps: constructing a multi-target network model based on Yolov8n, taking the Yolov8n network as the basis, replacing all C2f modules in the Yolov8n network with C3Ghost modules, replacing the Conv module with the GSConv module in the backbone network of the Yolov8n network, replacing the Conv module with the GhostConv module in the neck network, introducing the C3STR module at the same time, and adding a small target detection layer; then using pre-processed historical multi-target video images to train and optimize the multi-target network model to obtain an optimized multi-target network model; inputting the multi-target video images to be tested into the optimized multi-target network model for target recognition. The present invention not only improves the target detection accuracy in a large environment, but also reduces the number of parameters and model size.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of target detection, and in particular relates to a large-scale multi-target detection method for an open-pit mine spoil dump. Background Art

[0002] YOLOv8 is the latest iteration of the YOLO family of real-time object detectors, delivering cutting-edge performance in both accuracy and speed. Building on previous YOLO versions, YOLOv8 introduces new features and optimizations, making it ideal for a wide range of object detection tasks.

[0003] In open-pit mine dumps, bulldozers need to work 24 hours a day, and their operating areas are very dangerous. Therefore, accurate identification of the surrounding environment of the bulldozers is required. However, current target recognition in large environments such as open-pit mines cannot accurately identify small targets such as workers, especially under the influence of severe weather, and the target detection accuracy is greatly reduced. Summary of the Invention

[0004] In order to solve at least one of the above-mentioned technical problems existing in the prior art, the present invention provides a large-scale multi-target detection method for an open-pit mine spoil dump.

[0005] The present invention is implemented by the following technical solution: a large-scale multi-target detection method for an open-pit mine dump includes the following steps:

[0006] Acquire historical multi-target video images, and perform data annotation on target objects contained in the historical multi-target video images;

[0007] Performing data enhancement on the historical multi-target video images after data annotation, and dividing the data-enhanced historical multi-target video image dataset into a training set and a test set according to a preset division ratio;

[0008] A multi-target network model is constructed based on Yolov8n. The multi-target network model is based on the Yolov8n network, and all C2f modules in the Yolov8n network are replaced with C3Ghost modules that are a fusion of C3 modules and Ghost modules. The GSConv module is used to replace the Conv module in the backbone network of the Yolov8n network, and the GhostConv module is used to replace the Conv module in the neck network. At the same time, the C3STR module that is a fusion of the C3 module and the Swin-transformer module is introduced into the 8th and 12th layers of the Yolov8n network, and a small target detection layer with a feature map size of 160×160 is added.

[0009] The multi-objective network model is trained and optimized using the training set, and tested using the test set to obtain an optimized multi-objective network model;

[0010] The multi-target video image to be tested is input into the optimized multi-target network model, and the category of each target object in the multi-target video image to be tested is obtained using the optimized multi-target network model.

[0011] Preferably, the data processing process of the C3Ghost module, which is the fusion of the C3 module and the Ghost module, is as follows:

[0012] The multi-target image to be tested is used as the first input feature map of the C3Ghost module. First, the first input feature map passes through a GhostConv layer, then passes through a C3 module, and then passes through a GhostConv layer to obtain a first output feature map. The first input feature map and the first output feature map are added through a residual connection, and the number of channels is adjusted through a 1×1 convolution layer. Finally, the C3Ghost module feature map is output through the SiLU activation function; the C3 module includes at least one Bottleneck module.

[0013] Preferably, the data processing process of the GhostConv layer is:

[0014] The first input feature map is passed through the standard convolution layer on the main branch of the GhostConv layer to extract features to obtain a real feature map, and then the ghost feature map is generated by the depthwise separable convolution layer of the auxiliary branch of the GhostConv layer;

[0015] The real feature map and the ghost feature map are spliced ​​in the channel dimension, and finally the splicing result is channel rearranged to obtain the GhostConv layer feature map.

[0016] Preferably, the data processing process of the GSConv module is:

[0017] The second input feature map of the GSConv module is first obtained by standard convolution to obtain the GSConv first feature map, and then the feature-mixed GSConv second feature map is obtained by depth-wise separable convolution. The GSConv first feature map and the GSConv second feature map are concatenated in the channel dimension to obtain the GSConv third feature map, and finally the channels are rearranged to obtain the GSConv feature map.

[0018] Preferably, the data processing process of the C3STR module is:

[0019] The third input feature map of the C3STR module first passes through a 3×3 convolution layer, then passes through the Swin-transformer module and a 1×1 convolution layer to obtain the first feature map of the C3STR module, and then the third input feature map passes through a 1×1 convolution layer to adjust the number of channels, and then performs a residual connection with the first feature map of the C3STR module, and finally outputs the C3STR module feature map through the SiLU activation function.

[0020] Preferably, the data processing process of the small target detection layer is:

[0021] The fifth layer in the backbone network is stacked with the upsampling layer of the neck network, and then after passing through the C3Ghost module and upsampling, a small target feature layer containing small target feature information is obtained;

[0022] The small target feature layer and the third layer in the backbone network are stacked and then output through the decoupling head layer.

[0023] Preferably, data labeling is performed on the target objects contained in the historical multi-target video images, including:

[0024] Classifying the target objects contained in the historical multi-target video images into at least people, cars, trucks, buses, and bulldozers, and performing data annotation using an image annotation tool;

[0025] Convert the file format corresponding to the historical multi-target video image with data annotation into YOLO format.

[0026] Preferably, the data enhancement of the historical multi-target video images after data annotation includes:

[0027] Noise is added to simulate bad weather, and data enhancement is performed using an automatic color equalization algorithm, a dark channel prior dehazing algorithm, and salt and pepper noise and Gaussian noise.

[0028] Preferably, it also includes:

[0029] The optimized multi-objective network model is deployed on the edge device Jetson agx orin, and tensorRT is used for accelerated reasoning.

[0030] Compared with the prior art, the present invention has the following beneficial effects:

[0031] The present invention discloses a large-scale multi-target detection method for open-pit mine dumping sites. Improvements are made to the basic Yolov8n network. The C3Ghost module is used to replace the C2f module. The GSConv module is used to replace the Conv module in the backbone network. The GhostConv module is used to replace the Conv module in the neck network. The C3STR module is used to replace the C3Ghost module in the 8th and 12th layers. At the same time, a small target detection layer with a feature map of 160×160 is added. The improved multi-target network model significantly reduces weight while improving target detection accuracy. Furthermore, the improved multi-target network model is deployed to the edge device Jetson agx after accelerated inference through tensorRT. Real-time detection is achieved on orin, which not only makes full use of the image data in the large scene of the open-pit mine spoil dump, but also solves the problems of dust, strong light and other factors that have a great impact on the image data through data enhancement. At the same time, noise is added to simulate various severe weather conditions. The improved multi-target network model in this application improves the target detection accuracy in large scenes while reducing the number of parameters and model size, which reduces the computing power requirements of the multi-target network model, facilitates the deployment of the multi-target network model on edge devices for real-time detection, and also facilitates security personnel to complete observation and monitoring tasks in real time, so as to take corresponding protective measures in time. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0033] Figure 1 1 is a flow chart of a method for detecting multiple targets in a large-scale scene of an open-pit mine dump according to an embodiment of the present invention;

[0034] Figure 2 Schematic diagram of the layout of industrial cameras installed in different directions on a bulldozer according to an embodiment of the present invention;

[0035] Figure 3 2 is a schematic diagram of comparing historical multi-target video images processed by data enhancement according to an embodiment of the present invention;

[0036] Figure 4 This is a diagram of the Yolov8n-GhostGC network structure of an embodiment of the present invention;

[0037] Figure 5 is a structural diagram of a C3Ghost module according to an embodiment of the present invention;

[0038] Figure 61 is a network structure diagram of an unimproved Yolov8n model according to an embodiment of the present invention;

[0039] Figure 7 is a structural diagram of the GhostConv module according to an embodiment of the present invention;

[0040] Figure 8 : is a structural diagram of the GSConv module of an embodiment of the present invention;

[0041] Figure 9 This is a structural diagram of a C3STR module according to an embodiment of the present invention;

[0042] Figure 10 1 is a schematic diagram of the detection results of the multi-target network model according to an embodiment of the present invention;

[0043] Figure 11 is a bounding box loss verification curve graph of the Yolov8n-GhostGC model according to an embodiment of the present invention;

[0044] Figure 12 3 is a classification loss verification curve diagram of the Yolov8n-GhostGC model according to an embodiment of the present invention;

[0045] Figure 13 3. It is a distributed focus loss verification curve diagram of the Yolov8n-GhostGC model according to an embodiment of the present invention;

[0046] Figure 14 3. This is a mean precision verification curve diagram of the Yolov8n-GhostGC model of an embodiment of the present invention with an IOU threshold of 0.5;

[0047] Figure 15 3. This is a graph showing the mean precision verification curve of the Yolov8n-GhostGC model with an IOU threshold of 0.5 to 0.95 according to an embodiment of the present invention;

[0048] Figure 16 is the confusion matrix of the Yolov8n-GhostGC model validation set according to an embodiment of the present invention. DETAILED DESCRIPTION

[0049] The technical solutions in the embodiments of the present invention are clearly and completely described in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other implementations derived by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts are within the scope of protection of the present invention.

[0050] It should be noted that the structures, proportions, sizes, etc. illustrated in the drawings of this specification are only used to match the contents disclosed in the specification for people familiar with this technology to understand and read, and are not used to limit the conditions under which the present invention can be implemented. Therefore, they have no substantive technical significance. Any structural modification, change in proportional relationship or adjustment of size should fall within the scope of the technical content disclosed in the present invention without affecting the efficacy and purpose that can be achieved by the present invention. It should be noted that in this specification, relational terms such as first and second are only used to distinguish one entity from several other entities, and do not necessarily require or imply any actual relationship or order between these entities.

[0051] The purpose of the present invention is to provide a large-scale multi-target detection method for open-pit spoil dumps. Through an improved Yolov8n network architecture, a multi-target network model is constructed to realize the application of Yolov8n in the large-scale environment of open-pit mines. While improving the accuracy of target detection, the number of parameters and model size are reduced, so that the computing power requirements of the multi-target network model are reduced, the difficulty of deployment on open-pit mine bulldozers is reduced, and it is more conducive to deployment on edge devices.

[0052] In order to more clearly introduce the above-mentioned objects, features and advantages of the present invention, further detailed description is given below with reference to the accompanying drawings and specific embodiments.

[0053] like Figure 1 As shown, an embodiment of the present invention provides a large-scale multi-target detection method for an open-pit mine dump, comprising the following steps:

[0054] S1: Acquire historical multi-target video images, and perform data annotation on target objects contained in the historical multi-target video images.

[0055] Optionally, data annotation is performed on the target objects contained in the historical multi-target video image, including: dividing the target objects contained in the historical multi-target video image into at least people, cars, trucks, buses and bulldozers, and using an image annotation tool to perform data annotation; converting the file format corresponding to the historical multi-target video image with the data annotation into YOLO format.

[0056] In this embodiment, industrial cameras installed in different directions of the bulldozer are used to collect multi-target video images of large scenes in the open-pit mine dumping yard and store them to obtain historical multi-target video images. The detection range is the front and rear sides of the bulldozer, the left and right sides of the discharge arm, and the left and right sides of the front end of the discharge arm. The collected image data are annotated using an image annotation tool. There are five annotation categories, namely people, cars, trucks, buses and bulldozers. The corresponding label names are 0: person, 1: car, 2: truck, 3: bus, 4: bulldozer, and finally the annotated file format is converted into YOLO format.

[0057] In this embodiment, the layout diagram of the industrial camera is as follows Figure 2 As shown, the industrial camera used is a Hikvision industrial camera with 8 million pixels and a variety of focal lengths and field angles to choose from. The present invention is not limited to this, and it is sufficient to capture as many targets as possible.

[0058] S2: Perform data enhancement on the historical multi-target video images after data annotation, and divide the data-enhanced historical multi-target video image dataset into a training set and a test set according to a preset division ratio.

[0059] Optionally, data enhancement is performed on the historical multi-target video images after data annotation, including: adding noise to simulate bad weather, and using an automatic color balancing algorithm, a dark channel prior dehazing algorithm, and salt and pepper noise and Gaussian noise for data enhancement.

[0060] In this embodiment, a variety of data enhancement processing is performed on the historical multi-target video images after data annotation, including but not limited to the Automatic Color Enhancement (ACE) algorithm, the Dark Channel prior defogging algorithm, salt and pepper noise, and Gaussian noise. The above data enhancement methods are used to solve the problem that dust and strong light have a great influence on detection, and noise is added to simulate various bad weather conditions. The historical multi-target video images after data enhancement processing are as follows: Figure 3 As shown in the figure, from left to right are the comparisons of the original historical multi-target video image, the image processed by the DarkChannel dark channel prior dehazing algorithm, and the image processed by the automatic color balancing algorithm.

[0061] In this example, by simulating varying degrees of inclement weather, such as rain and fog, the data-augmented historical multi-target video images were divided into a training set and a validation set at a ratio of 9:1. In this example, a total of 2,000 data-augmented historical multi-target video images were used, 1,800 of which were selected as the training set and 200 as the validation set. Each video image in the training set can include one or more types of objects.

[0062] In this embodiment, the automatic color balancing algorithm includes two steps: adjusting the color and spatial domain of the image and dynamically expanding the corrected image. The color and spatial domain adjustments are performed by imitating the lateral inhibition and regional adaptability of the visual system. The formula is as follows:

[0063]

[0064] Among them, Subset is the image pixel set; For the channel In the pixel Pixel value after region adaptive filtering; for 、 The grayscale difference between two pixels expresses the biological lateral inhibition; is the distance metric function; As a brightness representation function (odd function), this embodiment selects the classic Saturation function.

[0065] The selection of different brightness functions and parameters controls the degree of contrast enhancement. The larger the slope of the classic saturation function before saturation, the more obvious the contrast enhancement. Finally, the saturation function was selected as the relative brightness representation function. Its formula is as follows:

[0066]

[0067] in, is the grayscale difference between two pixels on the same color channel; is a parameter that controls the degree of nonlinearity of the function.

[0068] The corrected image is globally adjusted for dynamic range, ensuring that it satisfies the grayscale world theory and white spot hypothesis. The automatic color balancing algorithm, which is applied to a single-channel image, is then extended to a three-channel image in the RGB color space, processing each channel separately and then integrating them.

[0069] In this embodiment, the DarkChannel dark channel prior dehazing algorithm is based on the dark channel prior theory, which states that at least one color channel in a natural scene has very low intensity values ​​in a local area. By estimating the relationship between atmospheric light intensity and the distance between pixels, the haze information in the image is inferred and used to remove the haze in the image. The formula for blurring the image is as follows:

[0070]

[0071] in, is the position of the pixel; is the observed image intensity; is the scene radiance of the image (fog-free image); A is the global atmospheric light; is the transmittance.

[0072] The dark channel prior theory formula is as follows:

[0073]

[0074] in, is the dark channel value of the corresponding image; is the pixel y in the color channel The intensity value on ; The position of the pixel A local window centered on For color channels exist The values ​​are taken from three channels, which correspond to the red, green, and blue channels respectively.

[0075] The transmittance is estimated by dark channel prior and atmospheric light, and the formula is as follows:

[0076]

[0077] in, is the observed pixel In the color channel The pixel value on ; For atmospheric light in the color channel Strength value on ; is the estimated transmittance.

[0078] Finally, the transmittance and atmospheric light are used to reconstruct the fog-free image. The formula is as follows:

[0079]

[0080] in, Generally it is 0.1.

[0081] In this embodiment, salt and pepper noise is generated by randomly scattering a large number of black dots (pepper) and white dots (salt) on the image by specifying the signal-to-noise ratio.

[0082] Gaussian noise is noise whose probability density function follows Gaussian distribution (normal distribution). The probability density function of Gaussian noise is as follows:

[0083]

[0084] in, is the grayscale value of the video image pixel; is the expected value of the pixel; is the standard deviation of the pixels; is the variance of the pixel, is the base of natural logarithms.

[0085] S3: A multi-target network model is constructed based on Yolov8n. The multi-target network model is based on the Yolov8n network, and all the C2f modules in the Yolov8n network are replaced by C3Ghost modules that are a fusion of C3 modules and Ghost modules. The GSConv module is used to replace the Conv module in the backbone network of the Yolov8n network, and the GhostConv module is used to replace the Conv module in the neck network. At the same time, the C3STR module that is a fusion of C3 modules and Swin-transformer modules is introduced into the 8th and 12th layers of the Yolov8n network, and a small target detection layer with a feature map size of 160×160 is added.

[0086] In this embodiment, Figure 4 The figure shows the Yolov8n-GhostGC network structure based on Yolov8n to build a multi-target network model. Figure 6 The Yolov8n shown is used as a skeleton model, using an improved Ghost module, replacing the C2f module in the network with the C3Ghost module that fuses the C3 module and the Ghost module, and replacing the Conv module with the GhostConv layer in the neck network; replacing the Conv module with the GSConv module in the backbone network of the skeleton model; the GSConv module is a mixed convolution using standard convolution SC, depthwise separable convolution DWC, and shuffle; replacing the C3Ghost module with the C3STR module that fuses the C3 module and the Swin-transformer module in the eighth and twelfth layers of the skeleton model; adding a small target detection layer with a feature map of 160×160 to the skeleton model, which includes an additional fusion feature and an additional detection head.

[0087] Optionally, the data processing process of the C3Ghost module that integrates the C3 module and the Ghost module is as follows: the multi-target image to be tested is used as the first input feature map of the C3Ghost module, first the first input feature map passes through a GhostConv layer, then passes through a C3 module, and then passes through a GhostConv layer to obtain a first output feature map, the first input feature map and the first output feature map are added through a residual connection, and the number of channels is adjusted through a 1×1 convolution layer, and finally the C3Ghost module feature map is output through a SiLU activation function; the C3 module includes at least one Bottleneck module.

[0088] In this embodiment, Figure 5 The following is the structure diagram of C3Ghost module. Figure 6 The Yolov8n model network structure shown in the figure is improved. The C3Ghost module, which combines the C3 and Ghost modules, replaces the C2f module in the network. The GhostConv module replaces the Conv module in the neck network. The Ghost module in the C3Ghost module generates more feature maps through low-cost operations. It first uses a portion of the original feature map and then applies a series of simple linear transformations (cheap operations) to generate more feature maps, called Ghost feature maps. These Ghost feature maps fully reveal the information in the original features. The GhostConv module increases the width of the network by generating more virtual feature maps without significantly increasing the number of model parameters and computational complexity, allowing the model to more effectively exchange and fuse information between feature maps of different scales. The neck network (Neck) is responsible for multi-scale feature fusion. Compared with the GSConv module, the GhostConv module can provide better feature fusion, higher accuracy, and better computational cost-effectiveness when replacing the Conv module in the neck network. Therefore, the GhostConv module is used to replace the Conv module in the neck network.

[0089] Optionally, the data processing process of the GhostConv layer is as follows: extract features from the first input feature map through a standard convolution layer on the main branch of the GhostConv layer to obtain a real feature map, and then generate a ghost feature map through a depth-separable convolution layer of an auxiliary branch of the GhostConv layer; splice the real feature map and the ghost feature map in the channel dimension, and finally rearrange the splicing result in channels to obtain a GhostConv layer feature map.

[0090] Optionally, the data processing process of the GSConv module is as follows: the second input feature map of the GSConv module is first obtained by standard convolution to obtain the GSConv first feature map, and then the feature-mixed GSConv second feature map is obtained by depth-separable convolution, the GSConv first feature map and the GSConv second feature map are spliced ​​in the channel dimension to obtain the GSConv third feature map, and finally the channels are rearranged to obtain the GSConv feature map.

[0091] In this embodiment, Figure 7 The structure diagram of the GhostConv layer is shown in Figure 1, which uses the GSConv module to replace the Conv module in the backbone network of the skeleton model; the structure diagram of the GSConv module is shown in Figure 1. Figure 8As shown in the figure, it uses a mixed convolution of standard convolution SC, depth-wise separable convolution DWC and shuffle; the main task of the backbone network Backbone is to extract features from the first input feature map. Compared with the unimproved GhostConv layer, the GSConv module used in this embodiment can achieve performance improvement while being lightweight. It is more effective than the unimproved GhostConv layer in feature extraction, especially in scenarios where it is necessary to maintain lightweight while improving model performance. Therefore, the GSConv module is used to replace the Conv module in the Backbone network.

[0092] Optionally, the data processing process of the C3STR module is as follows: the third input feature map of the C3STR module first passes through a 3×3 convolution layer, and then passes through a Swin-transformer module and a 1×1 convolution layer to obtain the first feature map of the C3STR module, and then the third input feature map passes through a 1×1 convolution layer to adjust the number of channels, and then performs a residual connection with the first feature map of the C3STR module, and finally outputs the C3STR module feature map through the SiLU activation function.

[0093] In this embodiment, Figure 9 The following figure shows the structure of the C3STR module. The C3STR module combines the features of the C3 module and the Swin-transformer module, demonstrating enhanced feature fusion capabilities. The Swin-transformer module captures global information through a self-attention mechanism, while the C3 module, through its structure, enhances the extraction and fusion of local features. Placing the C3STR module in the eighth and twelfth layers of the Yolov8n network can better leverage these advantages, particularly when the feature maps of the video image being tested are large in scale and require more contextual information to improve detection accuracy.

[0094] Optionally, the data processing process of the small target detection layer is: stacking the 5th layer in the backbone network with the upsampling layer of the neck network, and then passing through the C3Ghost module and upsampling to obtain a small target feature layer containing small target feature information; stacking the small target feature layer and the 3rd layer in the backbone network, and then outputting it through the decoupling head layer.

[0095] In this embodiment, Figure 4The small object detection layer shown in the figure includes an additional fusion feature and an additional detection head. Specifically, the fifth layer of the backbone network is first stacked with the upsampling layer of the neck network. After passing through the C3Ghost module and upsampling, a feature layer containing small object feature information is obtained. This feature layer is then stacked with the third layer of the backbone network. Finally, the feature layer is output through a new decoupling head layer. With the addition of detection heads, feature information from smaller objects is passed to the other three scale feature layers, thereby improving detection accuracy and expanding the range.

[0096] S4: Using the training set to train and optimize the multi-objective network model, and using the test set to test it, to obtain an optimized multi-objective network model.

[0097] In this embodiment, the superiority of the multi-objective network model can be determined by various performance evaluation indicators during the training process; the performance indicators specifically include accuracy , recall rate , average precision , the average precision when the IoU threshold is 0.5 , parameter quantity. The calculation formulas for each evaluation index are as follows:

[0098] 、

[0099] 、

[0100] 、

[0101]

[0102] in, is the precision rate, which is used to judge the probability of detecting a positive sample; The recall rate is used to judge the missed detection of the detection box. The larger the value, the lower the missed detection rate of the multi-target network model. is the number of positive samples correctly identified as positive samples; is the number of negative samples mistakenly identified as positive samples; is the number of positive samples mistakenly identified as negative samples; is the average precision, which is The area under the curve represents the average precision of the multi-objective network model at different recall rates; It is the average precision when the IoU threshold is 0.5, which measures the average precision of the multi-target network model on multiple detection categories.

[0103] In this embodiment, the improved multi-objective network model of this application is used to conduct actual detection of large-scale scenes in open-pit mine dumps, and the following results are obtained: Figure 10The test result diagram shown is Figure 10 It can be seen that various targets in the spoil dump can be accurately identified, such as Figure 11-Figure 15 They respectively represent the validation curves when the validation set is used for at least 400 iterations of the validation experiment; Figure 11-13 The following are the bounding box loss curve, classification loss curve, and distribution focus loss curve. It can be seen from the figures that through continuous iterative updates, the loss value between the predicted label and the true label of the multi-objective network model gradually decreases, indicating that the performance of the multi-objective network model is gradually enhanced; Figure 14 、 Figure 15 The graphs are respectively the average precision mean curve with an IoU threshold of 0.5 and the average precision mean curve with an IoU threshold of 0.5 to 0.95. It can be seen from the graph that through continuous iterative updates, the corresponding average precision mean gradually increases, indicating that the detection accuracy of the multi-target network model is gradually enhanced. Figure 16 The confusion matrix of the validation set is shown, where each value is the ratio between the predicted label and the true label.

[0104] S5: Inputting the multi-target video image to be tested into the optimized multi-target network model, and using the optimized multi-target network model to obtain the category of each target object in the multi-target video image to be tested.

[0105] In this embodiment, before the multi-target video image to be tested is input into the optimized multi-target network model, data enhancement processing can also be performed on the multi-target video image to be tested. The specific processing method refers to the data enhancement method for historical multi-target video images and will not be described in detail here. It is sufficient to enhance the multi-target video image to be tested into a video image that is easier to detect by the multi-target network model. The present invention is not limited to this.

[0106] In this embodiment, as shown in Table 1 below, in order to verify the effectiveness of the proposed application, the original Yolov8n network model is used as the baseline, and the ablation experiment is used to verify the superiority of the multi-target network model Yolov8n-GhostGC model proposed in this application. The performance evaluation indicators include accuracy and parameter amount.

[0107] Table 1

[0108]

[0109] As shown in Table 2 below, in order to verify the effectiveness of the proposed model, a comparative analysis was conducted with Yolov5n, Yolov8n, and Yolov11n. The improved model of this application has significantly improved the detection accuracy of four types of targets: person, car, truck, and bus, with only a slight decrease in bulldozer. The number of parameters has also been significantly reduced. Compared with the original model, the detection accuracy is improved by 2.7%, the number of parameters is reduced by 44%, and it is easier to deploy on the edge device for real-time detection.

[0110] Table 2

[0111]

[0112] Optionally, the method further includes: deploying the optimized multi-objective network model to the edge device Jetson agxorin, and using tensorRT for accelerated reasoning.

[0113] In this embodiment, the lightweight multi-objective network model is deployed on the edge device Jetson agx orin, and tensorRT is used to accelerate reasoning and improve real-time detection performance. Since bulldozers in large scenes of open-pit mine dumps work 24 hours a day and the working area is very dangerous, high real-time performance is very important. Current network models are generally large and have many parameters, and there are differences in the performance of equipment at the deployment end, resulting in slow reasoning speed and high latency, which cannot meet the high real-time operation requirements. TensorRT is a high-performance deep learning reasoning optimizer and runtime acceleration library that can provide low-latency, high-throughput deployment reasoning for deep learning applications, optimize the trained model, and thus improve model efficiency.

[0114] In this embodiment, the lightweight multi-target network model is converted into ONNX format, and then converted into the engine format specified by tensorRT; tensorRT is used for accelerated reasoning and deployed to Jetson agx orin for real-time detection.

[0115] This paper discloses a large-scale multi-target detection method for open-pit mine dumps, which effectively enhances the collected image data. Improvements are made to the basic model Yolov8n, replacing the C2f module with the C3Ghost module, replacing the Conv module with the GSConv module in the backbone network, replacing the Conv module with the GhostConv module in the neck network, replacing the C3Ghost module with the C3STR module in the eighth and twelfth layers, and adding a small target detection layer with a feature map of 160×160. The improved multi-target network model significantly reduces weight while improving target detection accuracy. Finally, the improved multi-target network model is deployed on the edge device Jetson AgXorin after accelerated inference through TensorRT to achieve real-time detection. This application not only makes full use of the video image data in the large scene of the open-pit mine spoil dump, but also solves the problems of dust, strong light, etc. that have a great impact on the image data through data enhancement, and adds noise to simulate various severe weather conditions. Moreover, the improved Yolov8n-GhostGC model can not only improve the accuracy of target detection, but also reduce the number of parameters and model size. The computing resources required to run the improved multi-target network model are reduced, and the computing power requirements are lowered, which is more conducive to deployment on edge devices for real-time detection, making it easier for security personnel to complete observation and monitoring tasks in real time, thereby facilitating the timely adoption of corresponding protective measures.

[0116] The foregoing description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be readily conceived by a person skilled in the art within the technical scope disclosed herein should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A large-scale multi-target detection method for open-pit mine dumping sites, characterized in that: The steps include: Acquire historical multi-target video images, and perform data annotation on target objects contained in the historical multi-target video images; Performing data enhancement on the historical multi-target video images after data annotation, and dividing the data-enhanced historical multi-target video image dataset into a training set and a test set according to a preset division ratio; A multi-target network model is constructed based on Yolov8n. The multi-target network model is based on the Yolov8n network, and all C2f modules in the Yolov8n network are replaced with C3Ghost modules that are a fusion of C3 modules and Ghost modules. The GSConv module is used to replace the Conv module in the backbone network of the Yolov8n network, and the GhostConv module is used to replace the Conv module in the neck network. At the same time, the C3STR module that is a fusion of the C3 module and the Swin-transformer module is introduced into the 8th and 12th layers of the Yolov8n network, and a small target detection layer with a feature map size of 160×160 is added. The multi-objective network model is trained and optimized using the training set, and tested using the test set to obtain an optimized multi-objective network model; The multi-target video image to be tested is input into the optimized multi-target network model, and the category of each target object in the multi-target video image to be tested is obtained using the optimized multi-target network model.

2. The method for detecting multiple targets in a large-scale open-pit mine dump according to claim 1, characterized in that: The data processing process of the C3Ghost module, which is the fusion of the C3 module and the Ghost module, is as follows: The multi-target image to be tested is used as the first input feature map of the C3Ghost module. First, the first input feature map passes through a GhostConv layer, then passes through a C3 module, and then passes through a GhostConv layer to obtain a first output feature map. The first input feature map and the first output feature map are added through a residual connection, and the number of channels is adjusted through a 1×1 convolution layer. Finally, the C3Ghost module feature map is output through the SiLU activation function; the C3 module includes at least one Bottleneck module.

3. The method for detecting multiple targets in a large-scale open-pit mine dump according to claim 2, wherein: The data processing process of the GhostConv layer is: The first input feature map is passed through the standard convolution layer on the main branch of the GhostConv layer to extract features to obtain a real feature map, and then the ghost feature map is generated by the depthwise separable convolution layer of the auxiliary branch of the GhostConv layer; The real feature map and the ghost feature map are spliced ​​in the channel dimension, and finally the splicing result is channel rearranged to obtain the GhostConv layer feature map.

4. The method for detecting multiple targets in a large-scale open-pit mine dump according to claim 1, wherein: The data processing process of the GSConv module is as follows: The second input feature map of the GSConv module is first obtained by standard convolution to obtain the GSConv first feature map, and then the feature-mixed GSConv second feature map is obtained by depth-wise separable convolution. The GSConv first feature map and the GSConv second feature map are concatenated in the channel dimension to obtain the GSConv third feature map, and finally the channels are rearranged to obtain the GSConv feature map.

5. The method for detecting multiple targets in a large-scale open-pit mine dump according to claim 1, wherein: The data processing process of the C3STR module is as follows: The third input feature map of the C3STR module first passes through a 3×3 convolution layer, then passes through the Swin-transformer module and a 1×1 convolution layer to obtain the first feature map of the C3STR module, and then the third input feature map passes through a 1×1 convolution layer to adjust the number of channels, and then performs a residual connection with the first feature map of the C3STR module, and finally outputs the C3STR module feature map through the SiLU activation function.

6. The method for detecting multiple targets in a large-scale open-pit mine dump according to claim 1, wherein: The data processing process of the small target detection layer is as follows: The fifth layer in the backbone network is stacked with the upsampling layer of the neck network, and then after passing through the C3Ghost module and upsampling, a small target feature layer containing small target feature information is obtained; The small target feature layer and the third layer in the backbone network are stacked and then output through the decoupling head layer.

7. The method for detecting multiple targets in a large-scale open-pit mine dump according to claim 1, characterized in that: Data annotation is performed on the target objects contained in the historical multi-target video image, including: Classifying the target objects contained in the historical multi-target video images into at least people, cars, trucks, buses, and bulldozers, and performing data annotation using an image annotation tool; Convert the file format corresponding to the historical multi-target video image with data annotation into YOLO format.

8. The method for detecting multiple targets in a large-scale open-pit mine dump according to claim 1, wherein: The data enhancement of the historical multi-target video images after data annotation includes: Noise is added to simulate bad weather, and data enhancement is performed using an automatic color equalization algorithm, a dark channel prior dehazing algorithm, and salt and pepper noise and Gaussian noise.

9. The method for detecting multiple targets in a large-scale open-pit mine dump according to claim 1, wherein: Also includes: The optimized multi-objective network model is deployed on the edge device Jetson agx orin, and tensorRT is used for accelerated reasoning.

Citation Information

Patent Citations

  • Deep learning-based chicken and egg target detection method for chicken farm

    CN118537802A

  • Road crack detection method, medium and product

    US20250174019A1