Ground target statistical method and system based on YOLOV10 model

By performing deep optimization and channel pruning in the YOLOV10 model, combined with an improved multi-objective tracking model, the problem of insufficient tracking accuracy of the drone on the ground in complex environments is solved, and efficient and accurate target tracking is achieved.

CN120198656AInactive Publication Date: 2025-06-24HUNAN GREAT WALL GALAXY TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510687145.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-06-24
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In complex environments, the tracking accuracy of the drone to small ground targets is limited, and the prior art is difficult to effectively improve detection accuracy without increasing time cost.

Method used

The ground target statistics method based on the YOLOV10 model is adopted, and the fine-grained feature extraction capability of small targets is improved through deep optimization feature detection model and channel pruning technology. The improved ByteTrack multi-objective tracking model and Kalman filtering are combined to achieve continuous and stable tracking of small targets.

Benefits of technology

It significantly improves the spatial resolution adaptability of the drone at low altitude viewing angle, achieves continuous and stable tracking of small ground targets, improves detection accuracy and reduces computing load.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198656A_ABST
    Figure CN120198656A_ABST
Patent Text Reader

Abstract

The invention relates to a ground target statistical method and system based on a YOLOV10 model. The method comprises the following steps: acquiring a video stream data set of a ground target through an unmanned aerial vehicle; and a YOLOV10 feature detection model is constructed. And decoding the video stream data set to obtain a video frame image, marking a small target of the video frame image to train a modified model, pruning the trained model, inputting the video frame image into the pruned model to extract features, and outputting small target detection frame information. And creating a small target initial track according to the small target detection frame information. And predicting the position of the next frame of small target according to the current frame of small target initial trajectory by adopting Kalman filtering to obtain a target prediction frame, matching the target prediction frame with the small target detection information of the current moving trajectory by adopting a Hungary algorithm, and counting the small target according to a regional technical strategy to obtain a tracking result of the ground target. By adopting the method, the tracking precision of the unmanned aerial vehicle on the small ground target in a complex environment can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multi-target detection and statistics, and particularly to a method and system for ground target statistics based on the YOLOV10 model. Background Art

[0002] In recent years, with the further development of artificial intelligence and unmanned aerial vehicle (UAV) technology, the fields involved in artificial intelligence have become more and more extensive, closely related to our lives, and also contributing to various industries and fields. Target detection, tracking, and statistics have been widely applied in various fields, but there are still many challenges and problems. Ground target statistics technology refers to accurately identifying and determining the ground and sea areas through specific technical means and counting the quantity, so as to provide accurate guidance and guarantee for decision-making.

[0003] With the continuous development of UAV technology, ground target detection technology has also been more widely applied. In the wide application, important information can be determined by identifying and counting ground targets through UAVs. The application scenarios include the statistics of the number of people on the ground, vehicles, ships in ports, and livestock in pastures. For ground targets under the vision of UAVs, the environment is complex and the targets are small. For small targets, the biggest problem is that their sizes are usually small, the perception range of target detection algorithms for small targets is limited, and the number of pixels of small targets is small, resulting in serious loss of detailed information and reduced inspection accuracy. Currently, the common methods to improve accuracy are to increase the image size and select a more complex network model, which will both lead to an increase in time cost. Summary of the Invention

[0004] Based on this, in view of the above technical problems, it is necessary to provide a method and system for ground target statistics based on the YOLOV10 model that can improve the tracking accuracy of UAVs for small ground targets in complex environments.

[0005] A method for ground target statistics based on the YOLOV10 model, the method comprising: Obtaining a video stream data set of ground targets through a UAV.

[0006] Constructing a YOLOV10 feature detection model.

[0007] Decoding the video stream data set to obtain video frame images, and after labeling the small targets in the video frame images, inputting them into the YOLOV10 feature detection model for small target detection training to obtain a trained YOLOV10 feature detection model.

[0008] After pruning the trained YOLOV10 feature detection model, inputting the video frame images into the pruned YOLOV10 feature detection model for feature extraction, and outputting small target detection box information.

[0009] Create the initial trajectory of small targets using the multi-object tracking ByteTrack model based on the small target detection box information.

[0010] Use Kalman filtering to predict the positions of small targets in the next frame based on the initial trajectories of small targets in the current frame, obtain the target prediction boxes, use the Hungarian algorithm to match the target prediction boxes with the small target detection information of the current running trajectories, and count the small targets according to the regional technology strategy to obtain the tracking results of ground targets.

[0011] A ground target statistics system based on the YOLOV10 model, the system includes: A data acquisition module, used to acquire the video stream dataset of ground targets through a drone.

[0012] A feature detection model construction module, used to construct a YOLOV10 feature detection model.

[0013] A model training module, used to decode the video stream dataset to obtain video frame images, after labeling the small targets in the video frame images, input them into the YOLOV10 feature detection model for small target detection training to obtain a trained YOLOV10 feature detection model.

[0014] A detection box information acquisition module, used to prune the trained YOLOV10 feature detection model, input the video frame images into the pruned YOLOV10 feature detection model for feature extraction, and output the small target detection box information.

[0015] An initial trajectory creation module, used to create the initial trajectory of small targets using the multi-object tracking ByteTrack model based on the small target detection box information.

[0016] A statistics module, used to use Kalman filtering to predict the positions of small targets in the next frame based on the initial trajectories of small targets in the current frame, obtain the target prediction boxes, use the Hungarian algorithm to match the target prediction boxes with the small target detection information of the current running trajectories, and count the small targets according to the regional technology strategy to obtain the tracking results of ground targets.

[0017] The above-mentioned ground target statistics method and system based on the YOLOV10 model adopt a deeply optimized YOLOV10 feature detection model in the detection layer to enhance the fine-grained feature extraction ability for small targets. Combining channel pruning technology to compress redundant parameters, while retaining the key feature extraction ability, it reduces the model's computational load, enabling the drone platform to process high-resolution video streams in real time and significantly improving the model's adaptability to spatial resolution from a low-altitude perspective. In the tracking layer, a multi-object association framework based on improved ByteTrack is constructed, and a spatio-temporal domain joint criterion is innovatively introduced - using Kalman filtering to perform Gaussian process modeling on the target motion state and combining the improved Hungarian algorithm to establish an IoU-DIoU double-threshold dynamic matching mechanism to effectively solve the problem of frequent ID switching in dense scenes. For complex meteorological interference, a regional perception enhancement strategy is designed: by establishing an adaptive ROI (region of interest) hierarchical strategy, local super-resolution reconstruction is performed on suspicious target areas; at the same time, a cross-frame optical flow constraint mechanism is integrated to filter out abnormal detection frames using the motion vector consistency between adjacent frames. At the system level, an end-to-end joint training paradigm is adopted, incorporating detection confidence and tracking continuity into the joint loss function, enabling the model to have the ability to adapt to dynamic environments. It has successfully achieved continuous and stable tracking of ground small targets by drones in complex weather environments, providing reliable technical support for accurate situation awareness. Brief Description of the Drawings

[0018] Figure 1 It is a schematic diagram of the system framework of the ground target statistics method based on the YOLOV10 model in an embodiment; Figure 2 It is a schematic flowchart of the ground target statistics method based on the YOLOV10 model in an embodiment; Figure 3 It is a partial network structure diagram of the YOLOV10 model after pruning in an embodiment; Figure 4 It is a block diagram of the structure of the ground target statistics system based on the YOLOV10 model in an embodiment. Detailed Description of the Invention

[0019] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0020] The ground target statistics method based on the YOLOV10 model provided by the present invention can be applied to, for example Figure 1In the system framework shown, it includes four main modules: input video stream, target detection, target tracking and counting, and statistical result display. Among them, target detection includes: data marking, designing a target detection model, optimizing the target detection model, model training, model pruning, and outputting detection information. Target tracking and counting includes: inputting target detection information, distinguishing high and low detection frames, Kalman filter prediction, matching, trajectory update, and target counting.

[0021] In one embodiment, as Figure 2 shown, a ground target statistical method based on the YOLOV10 model is provided. Taking the application of this method to Figure 1 the system framework in Step 202, obtain a video stream dataset of ground targets through a drone.

[0022] Step 204, construct a YOLOV10 feature detection model.

[0023] Step 206, decode the video stream dataset to obtain video frame images. After annotating the small targets in the video frame images, input them into the YOLOV10 feature detection model for small target detection training to obtain a trained YOLOV10 feature detection model.

[0024] Step 208, after pruning the trained YOLOV10 feature detection model, input the video frame images into the pruned YOLOV10 feature detection model for feature extraction, and output small target detection box information.

[0025] Step 210, create an initial trajectory of small targets according to the small target detection box information using the multi-target tracking ByteTrack model.

[0026] Step 212, use the Kalman filter to predict the position of small targets in the next frame according to the initial trajectory of small targets in the current frame to obtain target prediction boxes. Use the Hungarian algorithm to match the target prediction boxes with the small target detection information of the current running trajectory, and count small targets according to the regional technology strategy to obtain the tracking result of ground targets.

[0027] In the above-mentioned ground target statistics method based on the YOLOV10 model, in the detection layer, a deeply optimized YOLOV10 feature detection model is adopted to enhance the fine-grained feature extraction ability for small targets; combined with channel pruning technology to compress redundant parameters, while retaining the key feature extraction ability, the computational load of the model is reduced, enabling the UAV platform to process high-resolution video streams in real time, and significantly improving the spatial resolution adaptability of the model from a low-altitude perspective. In the tracking layer, a multi-target association framework based on improved ByteTrack is constructed, and a spatio-temporal domain joint criterion is innovatively introduced - using Kalman filtering to perform Gaussian process modeling on the target motion state, and combining the improved Hungarian algorithm to establish an IoU-DIoU double-threshold dynamic matching mechanism, effectively solving the problem of frequent ID switching in dense scenarios. Aiming at complex meteorological interference, a regional perception enhancement strategy is designed: by establishing an adaptive ROI (region of interest) hierarchical strategy, local super-resolution reconstruction is implemented for suspicious target areas; at the same time, a cross-frame optical flow constraint mechanism is integrated, and abnormal detection boxes are filtered using the motion vector consistency between adjacent frames. At the system level, an end-to-end joint training paradigm is adopted, incorporating detection confidence and tracking continuity into the joint loss function, enabling the model to have the ability to adapt to dynamic environments. It has successfully achieved continuous and stable tracking of small ground targets by UAVs in complex weather conditions, providing reliable technical support for accurate situation awareness.

[0028] In one embodiment, an Xsmall detection head is introduced into the head layer of the YOLOV1010 model, and the Xsmall detection head is connected to the convolutional layer of the backbone network to obtain the YOLOV10 feature detection model.

[0029] In one embodiment, the video stream dataset is decoded to obtain the current video frame images of ground targets under different weather conditions, and the small targets in the current video frame images are labeled to obtain the labeled video frame images. The labeled video frame images are divided into a training set, a validation set, and a test set according to a preset ratio and input into the YOLOV10 feature detection model for training to obtain the trained YOLOV10 feature detection model. The data format of each labeled video frame image in the labeled video frame image dataset is the two-dimensional coordinate information of the small target, the rectangle width and height information, and the class ID.

[0030] In one embodiment, an L1 regularization term is added to the original loss function of the trained YOLOV10 feature detection model to obtain a sparse loss function: ; where, is the sparse loss function, is the original loss function, It is the L1 regularization loss function. Obtain the weights of the trained YOLOV10 feature detection model for training on the training set, sort them according to the absolute values of the weights in the convolutional layer, and prune the weights corresponding to the smallest absolute values.

[0031] It should be noted that in order to improve the execution efficiency of the network, sparse pruning is selected to optimize the execution efficiency. By adding an L1 regularization term to the loss function, the training objective of the network can be changed from simply minimizing the loss function to minimizing the loss while keeping the weights sparse. During the network training process, with the effect of sparse regularization, some weights will gradually become very small. When the absolute value of the weight is less than a set threshold, these weights are pruned accordingly.

[0032] In one embodiment, the labeled video frame images are sequentially input into the pruned YOLOV10 feature detection model for feature extraction, and the position information and confidence results of small targets are output.

[0033] In one embodiment, according to the high and low confidence results in the small target detection box information, the detector of the multi-target tracking ByteTrack model is used to classify the detection boxes, obtaining high-score detection boxes and low-score detection boxes, and creating initial trajectories of small targets based on the high-score detection boxes and low-score detection boxes.

[0034] In one embodiment, the Kalman filter is used to predict the position of the small target detection box in the next frame of the labeled video frame image according to the current initial trajectory of the small target, obtaining the target prediction box. The Hungarian algorithm is used to track and match the target prediction box with the existing drone operation trajectories. For low-score detection boxes, calculate the IOU between the low-score detection box and the target prediction box that was not matched in the previous time step, obtaining the matched operation trajectories and low-score detection boxes, the unmatched operation trajectories, and the unmatched low-score detection boxes. Update the tracking box in the small target tracking trajectory to the detection box according to the matched operation trajectories and low-score detection boxes. For the high-score detection boxes that fail in tracking matching, match the unmatched high-score detection boxes with the to-be-activated operation trajectories, obtaining the matched to-be-activated operation trajectories and high-score detection boxes, the unmatched to-be-activated operation trajectories, and the high-score detection boxes of the unmatched to-be-activated operation trajectories. After deleting the unmatched to-be-activated operation trajectories, if the confidence of the high-score detection box of the unmatched to-be-activated operation trajectory is greater than the preset threshold, a new small target tracking trajectory is created. If the confidence of the high-score detection box of the unmatched to-be-activated operation trajectory is less than the preset threshold, discard the high-score detection box of the unmatched to-be-activated operation trajectory, and update the operation trajectory according to the matched to-be-activated operation trajectories and high-score detection boxes. Set the video frame image corresponding to the middle area of the video stream as the counting area, and count the number of times each continuously tracked small target passes through the technical area to obtain the tracking result of the ground target.

[0035] In one embodiment, as Figure 3 shown, a partial network structure diagram after pruning the YOLOV10 model is provided, introducing an XSmall detection head with a resolution of 160 pixels × 160 pixels. This head significantly reduces the downsampling to only two stages, enabling it to retain more detailed and richer features of small targets. It is connected to the feature fusion layer of the same scale in the backbone, and the rest is consistent with the network structure of YOLOV10.

[0036] It should be noted that adding the XSmall head to the head layer can improve feature fusion by integrating finer-grained and high-resolution features into the detection process.

[0037] In one embodiment, first, the video stream of the drone dataset in different weather conditions is decoded to obtain the current video frame image, and the small targets in the current video frame image are labeled; the labeled data is trained based on the improved YOLOV10; the trained YOLOV10 model is pruned.

[0038] Furthermore, the test video frame pictures in the drone dataset are input into the pruned YOLOV10 model, and the target objects in the image can be detected, and the position and size of the detection boxes of the target objects are obtained; the detection results are connected to the improved ByteTrack method to achieve target tracking based on the Kalman filter method, continuously updating the statistical result list, and the region counting strategy is used to count the targets.

[0039] Furthermore, to enhance the robustness of the model in the actual environment, weather factors are crucial. In this application, the drone dataset includes drone videos in three weather conditions: cloudy, rainy, and sunny. Then, they are decoded into single pictures, and the small targets in the images are labeled. The labeling format is: x, y, width, height, id. Where x and y are positions, width and height are the width and height of the rectangle, and id is the small target category.

[0040] Furthermore, the improved YOLOV10 means that the original YOLOV10 downsamples the feature map from stage P1 to stage P5. For our input image size of 640, when they reach the detection head, the resolutions of the generated feature maps are 80 (P3) pixels, 40 (P4) pixels, and 20 (P5) pixels respectively. To improve the detection ability for small targets, this method introduces an XSmall detection head with a resolution of 160 pixels × 160 pixels. This head significantly reduces the downsampling to only two stages, enabling it to retain more detailed and richer features of small targets. It is connected to the features of the same scale in the backbone, as Figure 3 shown, Figure 3It is a partial network structure diagram, and the rest is the same as the network structure of YOLOV10. Adding the XSmall head to the head layer can improve feature fusion by integrating finer-grained and high-resolution features into the detection process.

[0041] Furthermore, the already labeled dataset is divided into a training set, a validation set, and a test set in the ratio of 8:1:1 for training until the model converges. The improved network can further improve the accuracy of identifying small targets, and its map50 has increased by 20%.

[0042] Furthermore, the process of pruning the model is as follows: 1. Sparse training: Use L1 regularization to help the model learn more sparse weights; gradually increase the sparsity intensity to make some weights gradually become zero during the training process.

[0043] 2. Pruning: The weights of the convolutional layer can be sorted according to their absolute values, and the smallest weights are cropped; use the parameters (i.e., the scaling parameters γ and β) in the Batch Normalization (BN) layer to determine which convolutional kernels need to be pruned, so that the convolutional kernels corresponding to the BN parameters close to zero are pruned. In the pruning process, if the γ and β of an entire layer in the network are relatively small, it will be completely deleted during the proportional pruning process, resulting in network tomography and unable to transmit information normally. To solve this problem, this application ensures that at least 8 channels are retained in each layer to prevent the model from being truncated.

[0044] 3. Fine-tuning: After pruning, the accuracy of the model usually decreases, so fine-tuning is required to restore and optimize the performance. This application fine-tunes the pruned model with a lower learning rate. The parameters of the pruned model are reduced by 20%, and map50 is reduced by 5%.

[0045] Furthermore, decode the test video stream into single images and input them into the pruned YOLOV10 network in sequence to obtain the position information and confidence results of the objects in the image. Connect the results to the improved ByteTrack. The steps of the improved ByteTrack are as follows: 1. The detector obtains the detection boxes of the objects, which are divided into high-confidence detection boxes and low-confidence detection boxes according to the confidence level, and then creates initial trajectories.

[0046] 2. Use Kalman filtering to predict the positions of the fruit detection boxes in the next frame of the image to obtain the target prediction boxes.

[0047] 3. Use the Hungarian algorithm to perform tracking and matching on the object detection boxes and the existing trajectories in each frame of the video sequence, and assign a unique ID number to each object.

[0048] 4. Then, for the low-score bounding boxes, calculate the Intersection over Union (IOU) between the low-score bounding boxes and the prediction bounding boxes that were not matched in the previous step. Use the Hungarian algorithm to match the IOU and obtain three results: the matched trajectories and low-score bounding boxes, the trajectories that were not successfully matched, and the low-score bounding boxes that were not successfully matched. After successful matching, update the bounding boxes in the tracking trajectories to the detection bounding boxes.

[0049] 5. Finally, for the high-score detection bounding boxes that were not matched, match them with the trajectories whose status is not activated and obtain three results: match, unmatched trajectories, and unmatched detection bounding boxes. For the matched ones, update the status. For the unmatched trajectories, mark them for deletion. For the unmatched detection bounding boxes, if the confidence is greater than the high threshold + 0.1, create a new tracking trajectory; if it is less, discard it. Create, delete, and return the tracking trajectories. 6. Use the region counting strategy to count the objects.

[0050] Furthermore, the counting region is set in the middle region of the video, that is, 60% and 95% of the horizontal and vertical widths. Most of the objects within the field of view of this drone are in a relatively stable state, and the phenomenon of frequent switching of object IDs is not common; at the same time, considering the efficient utilization of the video, while reducing duplicate counting, a large counting region is ensured as much as possible. When each continuously tracked object passes through the set counting region, the statistical quantity is incremented by 1, and the counting result is statistically counted in real time.

[0051] It should be noted that first, videos collected under different weather conditions are used to train the model, which can enhance the robustness of the model; improve the network structure of YOLOV10, and adding an XSmall head to the head layer can improve feature fusion by integrating finer-grained and high-resolution features into the detection process, thereby improving the detection accuracy of the model for small objects; select sparse pruning to optimize the execution efficiency. By adding an L1 regularization term to the loss function, the execution efficiency of the network is improved, and the network parameters are reduced by 20%; connect the detection results to the improved ByteTrack method to implement object tracking based on the Kalman filter method, continuously update the result list, and use the region counting strategy to count the objects. In the prior art, the training dataset does not consider the data differences brought by weather, and there is no specific optimization for small objects, resulting in the need to further improve the inspection accuracy and the detection efficiency of the model also needs to be improved.

[0052] It should be understood that although Figure 1 the steps in the flowchart Figure 1At least some of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages do not necessarily need to be executed and completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages does not necessarily need to be sequential, but can be executed alternately or in turn with at least some of the sub-steps or stages of other steps or other steps.

[0053] In one embodiment, as Figure 4 shown, a ground target statistics system based on the YOLOV10 model is provided, including: a data acquisition module 402, a feature detection model construction module 404, a model training module 406, a detection box information acquisition module 408, an initial trajectory creation module 410, and a statistics module 412, where: The data acquisition module 402 is used to obtain a video stream data set of ground targets through a drone.

[0054] The feature detection model construction module 404 is used to construct a YOLOV10 feature detection model.

[0055] The model training module 406 is used to decode the video stream data set to obtain video frame images. After labeling the small targets in the video frame images, the labeled images are input into the YOLOV10 feature detection model for small target detection training to obtain a trained YOLOV10 feature detection model.

[0056] The detection box information acquisition module 408 is used to prune the trained YOLOV10 feature detection model, and then input the video frame images into the pruned YOLOV10 feature detection model for feature extraction, and output small target detection box information.

[0057] The initial trajectory creation module 410 is used to create an initial trajectory of small targets using the multi-object tracking ByteTrack model according to the small target detection box information.

[0058] The statistics module 412 is used to predict the position of small targets in the next frame according to the initial trajectory of small targets in the current frame using Kalman filtering to obtain target prediction boxes, match the target prediction boxes with the small target detection information of the current running trajectory using the Hungarian algorithm, and count the small targets according to the regional technology strategy to obtain the tracking results of ground targets.

[0059] For the specific limitations of the ground target statistics system based on the YOLOV10 model, reference can be made to the limitations of the ground target statistics method based on the YOLOV10 model in the foregoing text, which will not be elaborated herein. Each module in the above-mentioned ground target statistics system based on the YOLOV10 model can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to the above-mentioned modules.

[0060] Those skilled in the art can understand that Figure 4 the structure shown in is only a block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0061] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above embodiments of the methods. Among them, any reference to the memory, storage, database, or other media used in the various embodiments provided by the present invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0062] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0063] The above-described embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the appended claims.

Claims

1. A method for counting ground targets based on the YOLOV10 model, characterized in that, The method includes: Obtaining a video stream data set of ground targets through a drone; Constructing a YOLOV10 feature detection model; Decoding the video stream data set to obtain video frame images, annotating small targets in the video frame images, and then inputting them into the YOLOV10 feature detection model for small target detection training to obtain a trained YOLOV10 feature detection model; After pruning the trained YOLOV10 feature detection model, inputting the video frame images into the pruned YOLOV10 feature detection model for feature extraction, and outputting small target detection box information; Creating an initial small target trajectory according to the small target detection box information using the multi-object tracking ByteTrack model; Using Kalman filtering to predict the small target position in the next frame according to the initial small target trajectory in the current frame to obtain a target prediction box, using the Hungarian algorithm to match the target prediction box with the small target detection information of the current running trajectory, and counting small targets according to the regional technology strategy to obtain the tracking result of the ground target.

2. The method according to claim 1, characterized in that, Constructing a YOLOV10 feature detection model includes: Introducing an Xsmall detection head into the head layer of the YOLOV10 model, and connecting the Xsmall detection head to the convolutional layer of the backbone network to obtain a YOLOV10 feature detection model.

3. The method according to claim 2, wherein Decoding the video stream data set to obtain video frame images, annotating small targets in the video frame images, and then inputting them into the YOLOV10 feature detection model for small target detection training to obtain a trained YOLOV10 feature detection model, including: Decoding the video stream data set to obtain current video frame images of ground targets under different weather conditions, and annotating small targets in the current video frame images to obtain annotated video frame images; Dividing the annotated video frame images into a training set, a validation set, and a test set according to a preset ratio, and inputting them into the YOLOV10 feature detection model for training to obtain a trained YOLOV10 feature detection model; The data format of each annotated video frame image in the annotated video frame image data set is small target two-dimensional coordinate information, rectangular width and height information, and class ID.

4. The method according to claim 3, characterized in that, Pruning the trained YOLOV10 feature detection model includes: Adding an L1 regularization term to the original loss function of the trained YOLOV10 feature detection model to obtain a sparse loss function: ; Among them, is the sparse loss function, is the original loss function, is the L1 regularization loss function; Obtaining the weights of the trained YOLOV10 feature detection model for training on the training set, sorting the weights according to the absolute values of the weights in the convolutional layer, and pruning the weights corresponding to the smallest absolute value.

5. The method according to claim 4, characterized in that, Inputting the video frame images into the pruned YOLOV10 feature detection model for feature extraction, and outputting small target detection box information, including: Inputting the annotated video frame images into the pruned YOLOV10 feature detection model for feature extraction one by one as single pictures, and outputting the position information and confidence results of small targets.

6. The method according to claim 5, characterized in that, Creating an initial small target trajectory according to the small target detection box information using the multi-object tracking ByteTrack model, including: Classify the detection boxes using the detector of the multi-object tracking ByteTrack model according to the confidence results in the small target detection box information, obtaining high-confidence detection boxes and low-confidence detection boxes, and create initial small target trajectories based on the high-confidence detection boxes and the low-confidence detection boxes.

7. The method according to claim 6, wherein Use Kalman filtering to predict the positions of small targets in the next frame based on the initial small target trajectories in the current frame, obtaining target prediction boxes. Use the Hungarian algorithm to match the target prediction boxes with the small target detection information of the current running trajectories, and statistically calculate the matching results of the small targets according to the regional technology strategy to obtain the tracking results of the ground targets, including: Use Kalman filtering to predict the positions of the small target detection boxes in the next frame of the annotated video frame image based on the current initial small target trajectories, obtaining target prediction boxes. Use the Hungarian algorithm to perform tracking matching on the target prediction boxes and the running trajectories of the currently existing drones. For low-confidence detection boxes, calculate the IOU between the low-confidence detection boxes and the target prediction boxes that were not matched in the previous time step, obtaining the matched running trajectories and the low-confidence detection boxes, the unmatched running trajectories, and the unmatched low-confidence detection boxes. Update the tracking boxes in the small target tracking trajectories to detection boxes according to the matched running trajectories and the low-confidence detection boxes. For the high-confidence detection boxes that fail in tracking matching, match the unmatched high-confidence detection boxes with the running trajectories to be activated, obtaining the matched running trajectories to be activated and the high-confidence detection boxes, the unmatched running trajectories to be activated, and the high-confidence detection boxes of the unmatched running trajectories to be activated. After deleting the unmatched running trajectories to be activated, if the confidence of the high-confidence detection boxes of the unmatched running trajectories to be activated is greater than the preset threshold, create new small target tracking trajectories; if the confidence of the high-confidence detection boxes of the unmatched running trajectories to be activated is less than the preset threshold, discard the high-confidence detection boxes of the unmatched running trajectories to be activated, and update the running trajectories according to the matched running trajectories to be activated and the high-confidence detection boxes. Set the video frame image corresponding to the middle region of the video stream as the counting region, and count the number of times each continuously tracked small target passes through the technology region to obtain the tracking results of the ground targets.

8. A ground target statistics system based on the YOLOV10 model, characterized in that, The system includes: A data acquisition module for obtaining a video stream dataset of ground targets through a drone. A feature detection model construction module for constructing a YOLOV10 feature detection model. A model training module for decoding the video stream dataset to obtain video frame images, annotating the small targets in the video frame images, and then inputting them into the YOLOV10 feature detection model for small target detection training to obtain a trained YOLOV10 feature detection model. A detection box information acquisition module for pruning the trained YOLOV10 feature detection model, inputting the video frame images into the pruned YOLOV10 feature detection model for feature extraction, and outputting small target detection box information. An initial trajectory creation module, which is used to create an initial trajectory of the small target by using the multi-object tracking ByteTrack model according to the small target detection box information; A statistics module, which is used to predict the position of the small target in the next frame according to the initial trajectory of the small target in the current frame by using Kalman filtering to obtain a target prediction box, match the target prediction box with the small target detection information of the current running trajectory by using the Hungarian algorithm, and count the small targets according to the regional technology strategy to obtain the tracking result of the ground target.

Citation Information

Patent Citations

  • Improved YOLO and SIFT combined multi-small-target detection and tracking method for unmanned aerial vehicle

    CN111666871A

  • Field fruit counting method and system based on video target tracking

    CN117036238A

  • Traffic tracking detection system for view angle of unmanned aerial vehicle

    CN118918148A

  • Water column detection tracking algorithm based on ByteTrack

    CN119992401A