Beef cattle multi-target tracking method driven by double detection and pseudo-depth motion model
Through the dual detection and pseudo-depth motion model-driven method, combining high-precision and efficient object detection networks and pseudo-deep Kalman filters, the detection accuracy and calculation burden problems of multi-objective tracking in beef cattle breeding are solved, and efficient and stable tracking is achieved in complex environments.
Patent Information
- Application Number
- CN202510297451.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-07-04
AI Technical Summary
The existing multi-objective tracking methods are difficult to take into account the detection accuracy and calculation burden in beef cattle breeding, especially in complex environments when target occlusion and lighting changes are not good.
A dual detection and pseudo-depth motion model-driven method is adopted, combining high-precision context-aware object detection teacher network and efficient context-aware object detection student network, a pseudo-depth enhancement Kalman filter is introduced to improve the target tracking stability through pseudo-depth information, and a Hungarian algorithm is used to perform data correlation and linear interpolation processing missing detection.
It realizes multi-target tracking with high accuracy and low computational burden in beef cattle breeding scenarios, improves the stability and continuity of the target in complex environments, adapts to the needs of multiple breeding scenarios, and provides technical support for intelligent animal husbandry management.
Smart Images

Figure CN120259362A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of video image understanding, relates to multi-target tracking of beef cattle, and particularly relates to a multi-target tracking method for beef cattle driven by a dual detection and pseudo-depth motion model. Background Art
[0002] With the development of the beef cattle industry and the expansion of the breeding scale, the traditional manual observation mode, which is time-consuming, labor-intensive and has high labor costs, has been difficult to meet the needs of large-scale beef cattle breeding. What needs to be solved urgently at present is how to monitor the beef cattle breeding process through scientific means, timely discover and handle abnormal situations in breeding, and improve animal productivity, health and welfare. Based on the multi-target tracking method, the videos collected by the monitoring equipment in the farm are analyzed to realize the individual target tracking of beef cattle, which provides support for behavior recognition, physical sign monitoring and the establishment of connections for individual beef cattle.
[0003] In the prior art, various attempts have been made for the application of multi-target tracking technology in the field of livestock breeding, which can generally be divided into a back-end tracking optimization algorithm based on Hungary (KM matching), a multi-target tracking algorithm based on a single target tracker, and an end-to-end tracking method. The back-end tracking optimization algorithm based on Hungary (KM matching). Representative algorithms include SORT (Simple Online and Realtime Tracking) and DeepSORT (Simple Online and Realtime Tracking with a Deep Association Metric). Such algorithms can meet the requirements of real-time tracking, but their tracking effects largely depend on the accuracy of the target detector and the degree of feature differentiation. The multi-target tracking algorithm based on a single target tracker. The classic algorithm in this type of method is the kernel correlation filtering algorithm (KCF). This algorithm assigns a single target tracker to each target, and the tracking effect is very good, but it consumes a lot of computing resources in the case of a large number of targets, and has high requirements for memory and CPU. The end-to-end multi-target tracking algorithm. This type of algorithm integrates the detection and tracking processes in a unified framework, and optimizes the entire tracking process through end-to-end training, with relatively high computing efficiency. However, the training process of the end-to-end model is complex, requires a large amount of data and computing resources. And due to the high coupling of the detection and tracking processes, the adaptability of the model in different scenarios is poor. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the purpose of the present invention is to provide a multi-target tracking method for beef cattle driven by a dual detection and pseudo-depth motion model, so as to solve the technical problem that it is difficult for the multi-target tracking method in the prior art to balance the detection accuracy and the computing burden.
[0005] To solve the above technical problems, the present invention is implemented by the following technical solutions:
[0006] A multi-object tracking method for beef cattle driven by dual detection and pseudo-depth motion model, the method comprising the following steps:
[0007] Step S1, constructing a beef cattle detection and multi-object tracking data set.
[0008] Step S2, implementing a high-precision context-aware object detection teacher network HP-CAOD-TN:
[0009] Step S3, training the high-precision context-aware object detection teacher network HP-CAOD-TN:
[0010] Based on the beef cattle detection and multi-object tracking data set obtained in step S1, the high-precision context-aware object detection teacher network HP-CAOD-TN obtained in step S2 is trained.
[0011] Step S4, implementing an efficient context-aware object detection student network EC-CAOD-SN:
[0012] In the high-precision context-aware object detection teacher network HP-CAOD-TN trained in step S3, the context-aware boundary fusion pyramid module CABFP in the context-aware attention-enhanced feature fusion module CAAFM is replaced with a student-context-aware boundary fusion pyramid S-CABFP (Student-Context-Aware Boundary Fusion Pyramid) module to form a student-context-aware attention-enhanced feature fusion module S-CAAFM (Student-Context-Aware Attention-Enhanced Feature Fusion Module), thereby implementing an efficient context-aware object detection student network EC-CAOD-SN.
[0013] Step S5, model distillation:
[0014] Model distillation is achieved by calculating the similarity between the student feature map of the efficient context-aware object detection student network EC-CAOD-SN obtained in step S4 and the teacher feature map of the high-precision context-aware object detection teacher network HP-CAOD-TN obtained in step S3, and finally obtaining the target dual detection network.
[0015] Step S6, implementing a pseudo-depth enhanced Kalman filter P-DEKF:
[0016] Integrate the pseudo-depth information into the state vector of the Kalman filter, and expand the dimension of the state vector from 8 dimensions (x, y, w, h, vx, vy, vw, vh) to 10 dimensions (x, y, w, h, d, vx, vy, vw, vh, vd).
[0017] Where:
[0018] x and y represent the central position coordinates of the target bounding box;
[0019] w and h represent the width and height of the target bounding box;
[0020] d represents the pseudo-depth value, which indicates the distance between the target and the camera;
[0021] vx and vy represent the speeds in the x and y directions of the position;
[0022] vw and vh represent the changing speeds of the width and height;
[0023] vd represents the changing speed of the pseudo-depth.
[0024] Step S7, a multi-target tracking method for beef cattle driven by double detection and pseudo-depth motion model:
[0025] Based on the target double detection network obtained in S5, use the pseudo-depth enhanced Kalman filter P-DEKF in step S6 to simulate the state of the target for pseudo-depth motion model driving, and use the Hungarian algorithm for frame-by-frame data association to jointly achieve multi-target tracking.
[0026] Step S8, post-processing:
[0027] Adopt the linear interpolation method to fill the missing detection boxes at the missing detection positions in each trajectory.
[0028] Compared with the prior art, the present invention has the following technical effects:
[0029] (Ⅰ) The present invention is specially designed for the multi-target tracking task of beef cattle. By introducing the high-precision context-aware object detection teacher network HP-CAOD-TN and the efficient context-aware object detection student network EC-CAOD-SN, the present invention not only achieves high-precision object detection in multi-target tracking, but also takes into account the low computational burden. The combination of the two enables the system to flexibly and efficiently adapt to the complex task requirements of various breeding scenarios.
[0030] (II) The present invention incorporates pseudo-depth information into the state space of the Kalman filter, allowing the filter to indirectly perceive the depth changes of the target based on plane detection. Without relying on real depth sensors or accurate depth measurements, the system can infer the distance relationship of the target based on its position and size in the image. The pseudo-depth enhanced Kalman filter P-DEKF significantly improves the tracking stability and continuity of the target under dense occlusion and complex backgrounds, making up for the shortcomings of traditional two-dimensional tracking methods in depth judgment.
[0031] (III) The beef cattle multi-target tracking dataset constructed by the present invention covers a variety of real breeding scenarios, and the method specifically solves practical problems such as dense occlusion, complex background, and illumination changes. Based on this, the proposed beef cattle multi-target tracking method driven by dual detection and pseudo-depth motion model can meet the needs of beef cattle positioning and tracking in various production environments, providing technical support for intelligent and automated modern animal husbandry management.
[0032] (IV) The student model of the present invention inherits the key features of the teacher model while significantly reducing the computational cost. The introduction of pseudo-depth information not only does not significantly increase the computational complexity, but also improves the accuracy and robustness of target tracking, thus providing a solution that combines performance and efficiency.
[0033] (V) To address the problem of missing detections in multi-target tracking, the present invention adopts a linear interpolation method to effectively fill in the missing detection frames in the trajectory and fully utilizes the spatiotemporal continuity of the trajectory, thereby maintaining the integrity of the tracking.
[0034] (VI) The present invention combines the appearance-free model AFLink and predicts the connectivity between trajectories only through spatiotemporal information, which further enhances the stability and accuracy of the method in complex environments such as occlusion and irregular target motion. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 It is a framework diagram of the beef cattle multi-target tracking method driven by dual detection and pseudo-depth motion model of the present invention.
[0036] Figure 2 It is a structural diagram of the high-precision context-aware object detection teacher network HP-CAOD-TN of the present invention.
[0037] Figure 3 It is a structural diagram of the parallel context-aware feature network of the present invention.
[0038] Figure 4 It is a structural diagram of the context-aware attention enhancement feature fusion module of the present invention.
[0039] Figure 5 It is a structural diagram of the dual-pool channel attention module of the present invention.
[0040] Figure 6 It is the structural diagram of the triple attention fusion module of the present invention.
[0041] Figure 7 It is the structural diagram of the dual gating mechanism for controlling the feature fusion module of the present invention.
[0042] Figure 8 It is the structural diagram of the spectrum enhancement fusion module of the present invention.
[0043] Figure 9 It is the structural diagram of the OmniKernel module of the present invention.
[0044] Figure 10 It is the structural diagram of the student-context aware attention enhanced feature fusion module S-CAAFM module of the present invention.
[0045] Figure 11 It is the result comparison diagram of the method of the present invention and the comparative method.
[0046] The following further elaborates on the specific content of the present invention in conjunction with embodiments. Specific Embodiments
[0047] It should be noted that all models, networks, modules, and methods in the present invention, unless otherwise specified, all adopt the models, networks, modules, and methods known in the prior art.
[0048] In order to meet the requirements of beef cattle target tracking in the actual breeding process and cope with challenges in complex environments, such as complex backgrounds and dense occlusions, the present invention first collected a real beef cattle multi-object tracking dataset covering a variety of scenarios. Based on this dataset, the present invention proposed a beef cattle multi-object tracking method driven by dual detection and pseudo-depth motion model (Dual Detection and Pseudo-Depth Motion Model Driven Beef Cattle Multi-Object Tracking Method, DDPDM). This method is specifically designed for beef cattle target tracking and has the following innovative components:
[0049] The High-Precision Context-Aware Object Detection Teacher Network (HP-CAOD-TN) of the present invention consists of multiple key modules and is specifically designed to solve object detection problems in complex environments. First, the Parallelized Context-Aware Feature (PCAF) is proposed and designed, and the Parallelized Context-Aware Feature Network (PCAFNet) is constructed by combining it with a basic convolutional layer, thereby extracting richer feature information and enhancing the understanding of objects. Next, the Context-Aware Attention-Enhanced Feature Fusion Module (CAAFM) is proposed, which combines the Attention-based Intra-scale Feature Interaction (AIFI) and the Context-Aware Boundary Fusion Pyramid (CABFP), significantly improving the ability to capture object boundaries and details. To further improve the object detection accuracy, the teacher network also introduces an IoU-aware query selection mechanism and a decoder with an auxiliary prediction head to ensure more accurate and efficient object detection.
[0050] To improve the computational efficiency and reduce the computational burden, the present invention also designs the Efficient Context-Aware Object Detection Student Network (EC-CAOD-SN). The network structure is similar to that of the teacher network, but it proposes the Student-Context-Aware Boundary Fusion Pyramid module (S-CABFP) to replace the CABFP module, significantly reducing the computational overhead. By applying the model distillation technology, the student network can achieve high-efficiency and accurate object detection while maintaining a low computational burden and approaching or even exceeding the detection accuracy of the teacher network.
[0051] Another important component of the present invention is the Pseudo-Depth Enhanced Kalman Filter (P-DEKF). By introducing pseudo-depth information and incorporating it as part of the Kalman filter state space, this method significantly improves the stability and continuity of object tracking without relying on real depth sensors or precise depth measurements. This innovative design greatly enhances the robustness of multi-object tracking in complex environments, especially in cases of object occlusion and dense distribution, effectively maintaining the accuracy and coherence of object tracking.
[0052] Integrating the above technologies, the present invention proposes a multi-object tracking method for beef cattle driven by dual detection and pseudo-depth motion models, which takes into account both accuracy and efficiency and provides an efficient and stable real-time multi-object tracking solution. To address the problem of detection loss in multi-object tracking (e.g., failure to detect objects in some frames), the present invention adopts a linear interpolation method to fill in the missing detection boxes and make full use of the temporal and spatial continuity of the trajectories. At the same time, combined with the Appearance-Free Linkage (AFLink), the connectivity between trajectories is predicted using spatio-temporal information, further improving the accuracy and robustness of tracking.
[0053] Specific embodiments of the present invention are given below. It should be noted that the present invention is not limited to the following specific embodiments, and any equivalent transformation based on the technical solutions of this application falls within the protection scope of the present invention.
[0054] Embodiment:
[0055] This embodiment provides a multi-object tracking method for beef cattle driven by dual detection and pseudo-depth motion models. As Figure 1 shown, this method includes the following steps:
[0056] Step S1, constructing a beef cattle detection and multi-object tracking dataset:
[0057] Collect videos from five different breeding scenarios from real monitoring perspectives, annotate the videos using the DarkLabel tool to obtain annotation data in standard data format, and convert the annotation data into a beef cattle detection and multi-object tracking dataset.
[0058] In this step, these videos cover various interferences and challenges, such as occlusion, shadow, similar appearance, etc., and are closer to the actual cattle farm environment. Cattle herds are usually raised collectively and often do not lie at the center of the image.
[0059] Specifically, in this embodiment, videos of 5 different breeding scenarios from real monitoring perspectives were collected. The average duration of the videos was 35 minutes, in MP4 format, with a resolution of 1920×1080 pixels and a frame rate of 24 frames per second. Particularly, in different scenarios of beef cattle breeding, there are various interference factors affecting the detection performance, including: the scale change of beef cattle in the video due to the camera being installed in the corner of the shed roof; the change in light intensity caused by the large time span of video sampling; the mutual occlusion between beef cattle; the occlusion caused by external objects such as the shed roof and fences; the dust interference caused by behaviors such as beef cattle running; and the background interference generated when the appearance characteristics of beef cattle are not as obvious as those of pedestrians, especially when the body color of beef cattle is similar to the color of the muddy ground. Subsequently, we used the DarkLabel tool to annotate the beef cattle videos to obtain the standard data format, and converted the annotated data into a customized dataset suitable for beef cattle detection and multi-object tracking, namely the beef cattle detection and multi-object tracking dataset. The beef cattle target detection part in this dataset is named Beef-Cattle, including 11,936 training set pictures, 1,492 validation set pictures, and 1,492 test set pictures, and the ratio of the training set, validation set, and test set is 8:1:1. There are a total of 5 videos for testing beef cattle multi-object tracking.
[0060] Step S2, implement the High-Precision Context-Aware Object Detection Teacher Network HP-CAOD-TN:
[0061] Specifically, in this step, as Figure 2 shown, the key components of the High-Precision Context-Aware Object Detection Teacher Network HP-CAOD-TN (High-Precision Context-Aware Object Detection Teacher Network, HP-CAOD-TN) include: the Parallel Context-Aware Feature Network PCAFNet (Parallelized Context-Aware Feature Network), the Context-Aware Attention-Enhanced Feature Fusion Module CAAFM (Context-Aware Attention-Enhanced Feature Fusion Module), the IoU-Aware Query Selection Mechanism, and the decoder with an auxiliary prediction head.
[0062] The Parallel Context-Aware Feature Network PCAFNet consists of a basic convolution module and a PCAF module (Parallelized Context-Aware Feature).
[0063] The Context-Aware Attention Augmented Feature Fusion Module CAAFM combines the Attention-based Intra-scale Feature Interaction module AIFI and the Context-Aware Boundary Fusion Pyramid CABFP.
[0064] The Context-Aware Boundary Fusion Pyramid CABFP includes the Dual-Pooling Channel Attention Module DPCA, the Tri-Attention Fusion Module TAFM, the Dual-Gate Attention Fusion Module DGAFM to control the feature fusion module, and the Spectrum-Enhanced Fusion Module SEFM.
[0065] Furthermore, in step S2, the construction method of the High-Precision Context-Aware Object Detection Teacher Network HP-CAOD-TN is as follows:
[0066] Step S201, construct the Parallel Context-Aware Feature Network PCAFNet:
[0067] Send the input image data into the Parallel Context-Aware Feature Network PCAFNet to generate four levels of feature maps {P2, P3, P4, P5}.
[0068] In this embodiment, the specific structure of the Parallel Context-Aware Feature Network is as Figure 3 shown. Here, the PCAF module, as an enhanced variant of the C2f architecture, aims to optimize the feature extraction performance through parallel processing and advanced attention mechanisms. While maintaining a compact architecture, this module can efficiently handle computational tasks. By introducing the Parallelized Patch-Aware Attention (PPA) mechanism in the bottleneck layer, the PCAF module significantly improves the model's ability to capture complex feature representations while avoiding a substantial increase in computational overhead. The specific operations are as follows:
[0069] The module first receives an input feature map with c1 channels. By using a 1x1 convolutional layer cv1, the channel dimension is expanded from c1 to 2c, where c is the number of hidden channels determined by the expansion factor e. Then, the output of cv1 is split along the channel dimension to obtain two feature maps y1 and y2, each with c channels. The branch y1 is directly passed to the concatenation stage without additional processing, retaining the original feature information.
[0070] Different from the traditional C2f module, PCAF integrates n parallel PPA modules, and each PPA module independently processes the y2 branch to achieve diverse feature extraction. The PPA module focuses on the significant regions of the feature map through the attention mechanism, thereby enhancing the discriminative ability of the model in the spatial dimension. Subsequently, the outputs of the PPA modules are concatenated with the y1 branch along the channel dimension to obtain a combined feature map with (2 + n)c channels.
[0071] Finally, a 1x1 convolutional layer cv2 is used to reduce the number of channels of the concatenated feature map from (2 + n)c back to the desired output channel number c2. At this time, an activation function (such as FReLU) can also be introduced in this layer as needed to further increase non-linearity and enhance the feature representation ability. Finally, the output is a feature map with c2 channels, which contains the refined information processed and aggregated by the PCAF module.
[0072] Step S202, construct a context-aware attention enhanced feature fusion module CAAFM:
[0073] Input the P5 feature map obtained in step S201 into the internal interaction of the scale-internal feature interaction module AIFI of the attention mechanism, and output the interacted F5 feature map.
[0074] Input the P2 feature map, P3 feature map, P4 feature map, and F5 feature map obtained in step S201 into the context-aware boundary fusion pyramid CABFP, perform cross-scale feature fusion and concatenation on the P2 feature map, P3 feature map, P4 feature map, and F5 feature map, and output the Memory feature map.
[0075] Specifically in this embodiment, the structure of the context-aware attention enhanced feature fusion module CAAFM is as Figure 4 shown. The P5 feature map is first input into the scale-internal feature interaction module AIFI of the attention mechanism, and the interacted F5 feature map is output after internal processing. Input the P2, P3, P4, and F5 feature maps into the context-aware boundary fusion pyramid CABFP. The pyramid performs cross-scale feature fusion and concatenation on the input feature maps through four sub-modules: DPCA, TAFM, DGAFM, and SEFM, and finally generates the Memory feature map.
[0076] The construction process in step S202 further includes the following steps:
[0077] Step S20201, construct the dual-pool channel attention module DPCA:
[0078] The dual-pool channel attention module (DPCA) is as Figure 5 shown, aiming to dynamically adjust the importance of each channel in the input feature map by combining adaptive average pooling and adaptive max pooling operations, thereby enhancing the expression ability of key features. The specific process is as follows:
[0079] First, the input feature Figure X has C channels. The module processes the input feature map through two independent pooling operations: one is adaptive average pooling (AdaptiveAvgPool2d(1)), which compresses the spatial dimension of each channel into a single value to capture global context information; the other is adaptive max pooling (AdaptiveMaxPool2d(1)), which also compresses the spatial dimension of each channel into a single value, but focuses on capturing the significant features in the feature map. Next, the two feature maps after the pooling operations pass through a shared 1x1 convolutional layer (conv1) respectively, reducing the number of channels from C to where r is the channel compression ratio. Subsequently, after passing through the ReLU activation function (ReLU) to increase the non-linear ability, and then through the second 1x1 convolutional layer (conv2) to restore the number of channels to the original C. Then, the two feature maps processed by the average pooling path and the max pooling path are added together to obtain a comprehensive attention feature map. This comprehensive feature map further passes through the Sigmoid activation function (Sigmoid) to generate channel attention weights between 0 and 1. Finally, the output result is determined according to the flag bit flag during the initialization of the module. If flag = True, then the generated attention weights are multiplied element-wise with the original input feature Figure X to achieve weighted calibration of the features and enhance the feature expression of important channels; if flag = False, then only the attention weights are returned without weighting the input feature map. Through the above process, the dual-pool channel attention module (DPCA) can effectively identify and strengthen the important channel features in the input feature map, while suppressing the unimportant channel features, thereby improving the performance of the overall model.
[0080] Step S20202, construct the triple attention fusion module TAFM:
[0081] The triple attention fusion module (TAFM) is as Figure 6As shown, it aims to achieve multi-level weighted fusion of the input feature map by integrating channel attention, spatial attention, and pixel attention mechanisms, thereby enhancing the feature representation ability. The specific process is as follows:
[0082] First, the module receives a pair of input feature maps x and y, which usually come from different network branches or different scale levels. The module generates an initial fused feature map initial = x + y by adding these two feature maps element-wise. This operation effectively integrates the feature information from different sources and provides a fused basic feature for the subsequent attention mechanisms. Next, the module applies the channel attention and spatial attention mechanisms to the initial fused feature map respectively. Specifically, the channel attention module (ChannelAttention_CGA) calculates the importance weight cattn for each channel, emphasizing the channel features that contribute significantly to the task; while the spatial attention module (SpatialAttention_CGA) calculates the importance weight sattn for each spatial position, highlighting the significant spatial regions in the feature map. The two attention weights cattn and sattn are added together to obtain the comprehensive attention weight pattn1 = sattn + cattn, further integrating the attention information at the channel and spatial levels. Subsequently, the pixel attention module (PixelAttention_CGA) uses the comprehensive attention weight pattn1 to perform pixel-level attention calculation on the initial fused feature map, generating the final pixel attention weight pattn2 = σ(PixelAttention_CGA(initial, pattn1)), where σ represents the Sigmoid activation function, restricting the weight between 0 and 1 to ensure the smoothness and stability of the weighted fusion. In the feature recalibration stage, the module
[0083] performs weighted fusion on the input feature maps x and y according to the pixel attention weight pattn2, obtaining the fused feature map: result = initial + pattn2 × x + (1 - pattn2) × y. This operation dynamically adjusts the contribution degrees of different feature maps, not only strengthening the expression of important features but also suppressing unimportant features, thus enhancing the richness of feature representation and the discriminative ability of the model. Finally, the module further processes the fused feature map through a 1x1 convolutional layer (Conv2d(1x1)) to integrate and compress the feature information, generating the final output feature map. This convolutional operation not only helps with feature integration but also improves the feature expression ability, providing high-quality feature inputs for the subsequent network layers. In summary, the triple attention fusion module (TAFM) achieves multi-level weighted fusion of the input feature map by integrating channel attention, spatial attention, and pixel attention mechanisms, significantly enhancing the feature expression ability and model performance.
[0084] Step S20203, construct a dual gating mechanism to control the feature fusion module DGAFM:
[0085] The dual gating mechanism to control the feature fusion module (DGAFM) is as Figure 7 shown. Through the dual gating mechanism, dynamically control the fusion ratio of high-level features and low-level features to achieve more flexible and refined feature integration. The specific process is as follows:
[0086] First, the module receives a pair of input feature maps, H_feature (high-level feature) and L_feature (low-level feature), which usually come from different network branches or different scale levels. To start feature fusion, the module performs 1x1 convolution on L_feature and H_feature respectively through fully connected layers fc1 and fc2, reducing the number of channels from the original value to half (input_dim / / 2) to achieve feature transformation and channel compression. Next, the module passes the transformed L_feature and H_feature through the Sigmoid activation function respectively to generate gating weights g_L_feature and g_H_feature. These gating weights are used to dynamically adjust the fusion ratio of features, ensuring that important features are strengthened while unimportant features are suppressed. Subsequently, the module passes the L_feature and H_feature processed by the gating weights through 1x1 convolutional layers d_in1 and d_in2 respectively, keeping the number of channels unchanged. This step further extracts and transforms features to prepare for subsequent fusion operations. In the feature fusion stage, the module performs weighted fusion on L_feature and H_feature, and this operation dynamically adjusts the fusion ratio of high-level and low-level features through gating weights, ensuring that important features are strengthened while suppressing unimportant features. To ensure the spatial dimension consistency between the high-level feature H_feature and the low-level feature L_feature, the module upsamples H_feature to the same size as L_feature. Then, the fused H_feature and L_feature are concatenated in the channel dimension to form a comprehensive feature map with input_dim channels. Finally, the module further processes the concatenated feature map through a 3x3 convolutional layer conv to integrate and compress feature information and generate the final output feature map out. This process not only helps with feature integration but also improves the feature expression ability, providing high-quality feature inputs for subsequent network layers. Through the above process, the Dual Gating Attention Fusion Module (DGAFM) can effectively combine high-level and low-level features, enhance the expression of key features by dynamically adjusting the fusion ratio, and suppress unimportant features, thereby improving the performance and expression ability of the overall model.
[0087] Step S20204, construct the Spectral Enhancement Fusion Module SEFM:
[0088] The Spectral Enhancement Fusion Module (SEFM) is as Figure 8 shown. Its core goal is to process the feature map through spatial convolution, frequency domain transformation, and adaptive enhancement mechanism to improve the feature expression ability. The specific process is as follows:
[0089] First, the SEFM module receives an input feature map with a shape of [batch_size, dim, H, W] and performs preliminary processing through a 1x1 convolution (Conv1). The role of this convolutional layer is to adjust the dimensions of the input features and generate a streamlined feature map. After being processed by the convolutional layer, the feature map is divided into two branches:
[0090] Ok_Branch: This branch represents a part of the input feature map (channels divided according to the ratio e), and this branch will be input into the OmniKernel module for further processing.
[0091] Identity: This branch is the remaining part of the input feature map, keeping the original features unchanged and used for subsequent fusion with the processed features.
[0092] The OmniKernel module, as Figure 9 shown, is responsible for performing complex enhancement operations on the features in both the spatial domain and the frequency domain. Its processing flow is as follows:
[0093] Preliminary convolution and activation: The input feature map first passes through a 1x1 convolution (Conv2d(1x1)), and this operation is used to adjust the dimensions of the feature map to ensure that subsequent operations can be carried out on a unified scale. Immediately afterwards, the convolution output undergoes a non-linear transformation through the GELU activation function to enhance the expressive ability of the features.
[0094] Frequency domain enhancement (FCA): In the frequency domain enhancement part, after the feature map undergoes adaptive pooling, the channel attention (x_att) is obtained. Subsequently, the feature map undergoes a Fourier transform (FFT) to convert the features from the spatial domain to the frequency domain. In the frequency domain, the channel attention is modulated with the frequency domain features through element-wise multiplication, thereby enhancing the expressive ability of the frequency domain features. After an inverse Fourier transform (IFFT), it is restored to the spatial domain and the absolute value is taken to obtain the feature map enhanced in the frequency domain.
[0095] Spatial domain enhancement (SCA): The spatial domain enhancement part further enhances the spatial information of the feature map through adaptive pooling and convolution operations. In particular, after spatial domain enhancement, the feature map is weighted and fused with the feature map enhanced in the frequency domain, enabling the model to simultaneously focus on spatial information and frequency domain information and capture more fine-grained features.
[0096] Depthwise separable convolution (DwConv): The OmniKernel module also contains multiple depthwise separable convolution operations (dw_13, dw_31, dw_33, dw_11). These convolution operations help to further extract local spatial features and perform multi-scale fusion through different convolution kernel sizes to enhance the richness of spatial features.
[0097] Adaptive Feature Re-calibration (FGM): After completing spatial and frequency domain enhancement, the feature map is further adaptively re-calibrated through the FGM module (Frequency and Spatial Modulation). The FGM module further adjusts the enhanced feature map through the adaptive fusion of the frequency domain and the spatial domain. Specifically, the FGM module uses the enhanced information in the frequency domain and the spatial domain, combined with the weighting coefficients (alpha and beta), to re-calibrate the output features, thus ensuring the maximization of the expressive ability of the final features.
[0098] Activation function and output: The processed feature map undergoes a non-linear transformation through the ReLU activation function to enhance the expressive ability of the features. Finally, a final output feature map is generated through an output convolutional layer (Conv2d(1x1)), completing the processing flow of the Ok_Branch branch of the SEFM module.
[0099] The processed Ok_Branch (output from OmniKernel) and the original Identity branch are concatenated together to form a new feature map. The concatenation operation is completed through torch.cat(). Subsequently, the concatenated feature map is further processed through another 1x1 convolution (Conv2) to generate the final output feature map.
[0100] Finally, the feature map processed by the convolutional layer is returned as the output. This output feature map integrates spatial, frequency domain information, and the original features, and has a stronger expressive ability.
[0101] The SEFM module enhances the feature expressive ability by combining spatial convolution and frequency domain transformation. Its core steps include: preliminary processing of the input features, branch operations, adaptive enhancement in the frequency domain and the spatial domain, and feature fusion. Through the OmniKernel module, SEFM can optimize the features in multiple dimensions in the spatial and frequency domains, thereby enhancing the expressiveness and robustness of the model. Finally, the fused features processed by convolution will be used as the output to provide a richer feature representation.
[0102] Step S203, construct an IoU-aware query selection mechanism:
[0103] Calculate the regression and classification losses between the predicted bounding boxes and the ground truth bounding boxes corresponding to each feature in the Memory feature map obtained in step S202, and add the IOU to the classification loss to ensure the consistency of the location confidence and the class confidence of the network output. Select the top-K features according to the losses.
[0104] Step S204, construct a decoder with an auxiliary prediction head:
[0105] The Decoder receives the top-K features selected in step S203, performs initial target queries and the Memory feature sequence output by the Hybrid Encoder, iteratively optimizes the target queries through multi-layer self-attention and cross-attention, and finally outputs the category and coordinates of the detection box; the Head maps the category and coordinates of the detection box output by the Decoder onto the detection box.
[0106] Step S3, training the High-Precision Context-Aware Object Detection Teacher Network HP-CAOD-TN:
[0107] Based on the beef cattle detection and multi-object tracking dataset obtained in step S1, train the High-Precision Context-Aware Object Detection Teacher Network HP-CAOD-TN obtained in step S2. The loss function L used for training includes the target bounding box regression loss L box (Bounding Box Regeression Loss) and the classification loss L cls (Classificition Loss), that is, L = L box +L cls ; Backpropagate the loss function L and repeat the iteration until the iteration number reaches the preset initial value to complete the training.
[0108] In this step, the target bounding box regression loss L box is defined as the weighted sum of generalized IoU (GIoU) and L1 loss, that is
[0109]
[0110] In the formula:
[0111] represents the predicted bounding box;
[0112] b represents the ground truth bounding box;
[0113] L GIoU represents the GIoU loss;
[0114] λ represents the weight parameter used to adjust the proportion of the L SIoU loss in the L box loss;
[0115] represents the weight parameter used to adjust the proportion of the L1 loss in the L box loss.
[0116] In this step, the classification loss L cls adopts the variable focal loss (VFL, varifocal loss), that is:
[0117]
[0118] In the formula:
[0119] VFL represents variable focal loss;
[0120] p represents the predicted instance-aware classification score, which is the classification score of each instance predicted by the model;
[0121] q represents the target score;
[0122] Both α and κ represent hyperparameters.
[0123] Step S4, implement the efficient context-aware object detection student network EC-CAOD-SN:
[0124] In the high-precision context-aware object detection teacher network HP-CAOD-TN trained in step S3, replace the context-aware boundary fusion pyramid module CABFP in the context-aware attention-enhanced feature fusion module CAAFM with the student-context-aware boundary fusion pyramid S-CABFP (Student-Context-Aware Boundary Fusion Pyramid) module to form the student-context-aware attention-enhanced feature fusion module S-CAAFM (Student-Context-Aware Attention-Enhanced Feature Fusion Module), thereby implementing the efficient context-aware object detection student network EC-CAOD-SN.
[0125] In this embodiment, the student-context-aware attention-enhanced feature fusion module S-CAAFM is as Figure 10 shown.
[0126] Step S5, model distillation:
[0127] Implement model distillation by calculating the similarity between the student feature map of the efficient context-aware object detection student network EC-CAOD-SN obtained in step S4 and the teacher feature map of the high-precision context-aware object detection teacher network HP-CAOD-TN obtained in step S3, and finally obtain the target dual detection network.
[0128] In step S5, the specific method of model distillation is as follows:
[0129] Step S501, feature map alignment operation:
[0130] For the feature maps generated by the student model and the teacher model, we first need to perform an alignment operation. The feature maps of the student model and the teacher model may differ in terms of the number of channels, size, etc. Therefore, we need to align them through a convolutional layer so that their feature maps match in terms of the number of channels.
[0131] aligned_feat s = Conv2d(f s ), aligned_feat t = Conv2d(f t );
[0132] Where:
[0133] f s and f t respectively represent the feature maps of the student and teacher models;
[0134] aligned_feat s and aligned_feat t respectively represent the aligned feature maps of the student and teacher models;
[0135] Conv2d represents a 1x1 convolutional operation that maps the feature maps of the student and teacher to the same number of channels.
[0136] In this way, we ensure the consistency of the student and teacher feature maps in terms of channels.
[0137] Step S502, Normalize the feature maps:
[0138] Normalize the aligned feature maps to ensure that the features in each channel have zero mean and unit variance.
[0139]
[0140] Where:
[0141] x represents the input feature map or feature, usually the output of the student or teacher model;
[0142] μ represents the mean of the input x, which is calculated batch-wise and represents the average value of each channel.
[0143] σ represents the standard deviation of the input x, representing the distribution range of each channel.
[0144] After normalization, the features have zero mean and unit standard deviation. The purpose of normalization is to make the features more stable during training and avoid difficulties in training caused by too large or too small numerical ranges.
[0145] Step S503, calculate the similarity of the feature maps:
[0146] Calculate the cross-channel correlation between the student and teacher models, and quantify their similarity by calculating the correlation matrix of the student and teacher feature maps. To obtain the similarity, we first flatten the feature maps into two-dimensional matrices and then calculate the correlation between each pair of channels.
[0147]
[0148] In the formula:
[0149] f s and f t represent the feature maps of the student and teacher models respectively;
[0150] bsz represents the batch size;
[0151] ch represents the number of channels;
[0152] H and w represent the height and width of the feature map respectively.
[0153] By flattening each feature map into a two-dimensional matrix, we can obtain a feature map with a shape of bsz×ch×(H×w). emd s and emd t represent the correlation matrices of the student and teacher feature maps respectively, and represent the transposes of the student and teacher feature maps.
[0154] Step S504: Normalize the correlation matrix:
[0155] Normalize the calculated correlation matrix with the L2 norm to ensure that the calculated similarity has a consistent scale.
[0156]
[0157] In the formula:
[0158] ‖ ‖2 represents the L2 norm, which is used to normalize the correlation matrix to make it have a unified scale;
[0159] emd s represents the correlation matrix of the student feature map, which calculates the similarity between channels. By normalizing, the norm of each matrix is 1, avoiding instability in training caused by overly large values.
[0160] emd t represents the correlation matrix of the teacher feature map, and the normalization operation is the same as that of the student feature map.
[0161] Step S505, calculate the loss:
[0162] For each pair of student and teacher feature maps, calculate the difference between their correlation matrices.
[0163]
[0164] Where:
[0165] emd s and emd t represent the correlation matrices of the student and the teacher;
[0166] i and j represent the indices in the batch;
[0167] bsz represents the batch size;
[0168] ch represents the number of channels;
[0169] loss represents the loss, which is used to measure the difference in cross-channel correlation between the student and the teacher.
[0170] Step S506, weighted loss:
[0171] The loss of each layer will be weighted according to its importance and aggregated to obtain the final total loss.
[0172]
[0173] Where:
[0174] L represents the total number of feature layers;
[0175] w i represents the weight of each layer;
[0176] loss represents the loss calculated for each layer.
[0177] total_loss represents the weighted sum of the losses, which is used to represent the overall difference between the student model and the teacher model in terms of feature maps.
[0178] Step S507, backpropagation and optimization:
[0179] After calculating the total loss, use the backpropagation algorithm to optimize the parameters of the student model so that its performance in the feature space is as close as possible to that of the teacher model. The backpropagation process updates the parameters of the student model through the total loss.
[0180]
[0181] Where:
[0182] represents the gradient of the parameters θ of the student model s and the parameters are updated through backpropagation.
[0183] Step S6, implement the Pseudo-Depth Enhanced Kalman Filter (P-DEKF):
[0184] In this step, the pseudo-depth is a method for indirectly estimating the depth of an object (i.e., the distance from the camera) based on the position and size of the object in the image. It is assumed that the larger the y coordinate in the image (i.e., the closer the object is to the bottom of the image), the smaller the pseudo-depth, indicating that the object is farther from the camera.
[0185] Integrate the pseudo-depth information into the state vector of the Kalman filter, expanding the dimension of the state vector from 8 dimensions (x, y, w, h, vx, vy, vw, vh) to 10 dimensions (x, y, w, h, d, vx, vy, vw, vh, vd).
[0186] Where:
[0187] x and y represent the center position coordinates of the object bounding box;
[0188] w and h represent the width and height of the object bounding box;
[0189] d represents the pseudo-depth value, indicating the distance of the object from the camera;
[0190] vx and vy represent the velocities in the x and y directions of the position;
[0191] vw and vh represent the change velocities of the width and height;
[0192] vd represents the change velocity of the pseudo-depth.
[0193] In step S6, the expansion method includes the following steps:
[0194] Step S60, define the state vector:
[0195] The expanded state vector is defined as:
[0196]
[0197] Step S602, establish the motion model:
[0198] P-DEKF assumes that the motion of the object follows a constant velocity model, i.e., the position and pseudo-depth change linearly according to their respective velocities. The state transition matrix F is defined as:
[0199]
[0200] Where:
[0201] Δt is the time step.
[0202] This matrix ensures that the position and pseudo-depth change linearly with the time step, while the velocities remain constant.
[0203] Step S603, establish an observation model:
[0204] The observation vector z includes the center position, width, height, and pseudo-depth of the target:
[0205]
[0206] The observation matrix H maps the state vector to the observation space:
[0207]
[0208] This matrix indicates that the observed values directly correspond to the position, size, and pseudo-depth parts in the state vector.
[0209] Step S604, initialize the pseudo-depth enhanced Kalman filter P-DEKF:
[0210] In the initialization step of the pseudo-depth enhanced Kalman filter P-DEKF, the state vector and covariance matrix are set based on the initial measurements (including pseudo-depth). The initial state vector x0 contains the observed position, size, and pseudo-depth, and the velocity part is initialized to zero:
[0211]
[0212] Where:
[0213] x0, y0, w0, h0, d0 represent the initial observed values, and the velocity part is initialized to zero.
[0214] The covariance matrix P0 is set as a diagonal matrix according to the uncertainties of the position, size, and pseudo-depth:
[0215]
[0216] Where:
[0217] σ represents the uncertainty of each state variable.
[0218] Step S605, prediction:
[0219] In the prediction step, the state and covariance at the next time step are predicted based on the motion model:
[0220] x k|k-1 = F · x k-1|k-1 ;
[0221] P k|k-1 = F · P k-1|k-1 · F T + Q;
[0222] Where:
[0223] F represents the state transition matrix;
[0224] F T represents the transpose of the state transition matrix;
[0225] Q represents the process noise covariance matrix, reflecting the uncertainty of model prediction.
[0226] x k|k-1 represents the predicted state vector, predicting the state at the current time k based on the state and motion model at the previous time k-1;
[0227] P k|k-1 represents the predicted covariance matrix, predicting the state covariance at the current time k based on the covariance and motion model at the previous time k-1.
[0228] Step S606, update:
[0229] In the update step, fuse the new measurement value into the predicted state.
[0230] y k = z k - H·x k|k-1 ;
[0231] S k = H·P k|k-1 ·H T + R;
[0232]
[0233] x k|k = x k|k-1 + K k ·y k ;
[0234] P k|k = (I - K k ·H)·P k|k- 1;
[0235] Where:
[0236] y k represents the innovation vector, the measurement residual, representing the difference between the observed value and the predicted observed value;
[0237] S k represents the innovation covariance, representing the uncertainty of the innovation vector, combining the predicted covariance and the measurement noise;
[0238] K k represents the Kalman gain, used to balance the weights of the prediction and the measurement value, determining how to update the state vector;
[0239] R represents the measurement noise covariance matrix, which represents the noise in the measurement process and reflects the uncertainty of the measurement value;
[0240] x k|k represents the updated state vector, which combines the predicted value and the measurement value, and is the state estimate at the current moment after update;
[0241] P k|k represents the updated covariance matrix, which is the updated state covariance and reflects the uncertainty of the state estimate;
[0242] I represents the identity matrix, which is used to ensure dimension consistency in matrix operations;
[0243] Step S607, calculate the gating distance:
[0244] The gating distance is used to determine whether the measurement value matches the predicted value. After introducing the pseudo-depth, the calculation of the gating distance takes into account the additional depth information, thereby improving the accuracy of the matching. Commonly used distance metrics include the Mahalanobis Distance and the Euclidean Distance:
[0245]
[0246] By setting the gating threshold based on the chi-square distribution, the system can effectively filter out the measurement values that do not conform to the prediction, ensuring the accuracy and robustness of the tracking.
[0247] Step S608, calculate the pseudo-depth:
[0248] The calculation of the pseudo-depth is based on the position and size of the target in the image, and the specific formula is as follows:
[0249] lendth = 2000 - (y top-left + h);
[0250] where: y top-left represents the y coordinate of the upper left corner of the target bounding box; h represents the height of the target bounding box.
[0251] In this embodiment, the pseudo-depth enhanced Kalman filter (P-DEKF) significantly improves the performance of the multi-object tracking system in complex aquaculture environments by integrating pseudo-depth information into the state vector. Its specific process includes the expansion of the state vector, the establishment of the motion model and the observation model, the initialization of the state, the implementation of the prediction and update steps, and the calculation of the gating distance and pseudo-depth. Pseudo-depth not only enhances the ability to distinguish targets and prediction accuracy but also improves the stability and robustness of tracking. Moreover, it does not rely on additional depth sensors, making the system more efficient and practical in actual applications. Through this innovative method, the multi-object tracking requirements of beef cattle in the farm are effectively met, providing an efficient, stable, and accurate real-time tracking solution.
[0252] Step S7, a multi-object tracking method for beef cattle driven by double detection and pseudo-depth motion model:
[0253] Based on the target double detection network obtained in S5, use the pseudo-depth enhanced Kalman filter (P-DEKF) in step S6 to simulate the state of the target for pseudo-depth motion model driving, and use the Hungarian algorithm for frame-by-frame data association to jointly achieve multi-object tracking.
[0254] Step S8, post-processing:
[0255] Adopt the linear interpolation method to fill in the missing detection boxes at the missing detection positions in each trajectory.
[0256] In this embodiment, in multi-object tracking, missing detections (i.e., the target is not detected in some frames) will lead to a decrease in tracking accuracy. To solve this problem, the linear interpolation method is adopted to fill in the missing detection boxes at the missing detection positions in each trajectory. This operation makes full use of the continuity of the trajectory in time and space. At the same time, combined with the appearance-free model AFLink, solely relying on spatio-temporal information to predict the connectivity between two trajectories, thereby further improving the accuracy and robustness of tracking.
[0257] In this embodiment, the evaluation method of the target double detection network obtained in step S5 is as follows: Given beef cattle image data, input it into the target double detection network obtained in step S5, and output the category and coordinates of the beef cattle detection box. The evaluation metrics for measuring detection accuracy are Precision, Recall, mAP0.5, and mAP0.5:0.95. In addition to measuring detection accuracy, metrics such as the number of model parameters (Params), computational complexity (Gflops), inference time (Inference Time), and FPS (Frame Per Second) are also very important for measuring the overall performance of the model.
[0258] In this embodiment, the evaluation of the beef cattle multi-object tracking method in step S7 is as follows: HOTA (Higher Order Tracking Accuracy), MOTA (Multiple Object Tracking Accuracy), IDF1 (ID F1 Score), and ID Switch (Identity Switch) are used for evaluation.
[0259] Comparative Example 1:
[0260] In this comparative example, the proposed dual detection network is compared with 20 classic and advanced object detection methods. The specific methods for comparison are Faster R-CNN (NeurIPS'2015), RetinaNet (ICCV'2017), Cascade R-CNN (CVPR'2018), Libra R-CNN (CVPR'2019), ATSS (CVPR'2020), Double-Head R-CNN (CVPR'2020), GFL (Advances in Neural Information Processing Systems'2020), EfficientDet (CVPR'2020), Dynamic R-CNN (ECCV'2020), DDOD (ACM MM'2021), VarifocalNet (CVPR'2021), TOOD (ICCV'2021), Deformable DETR (ICLR'2021), DAB-DETR (ICLR'2022), DDQ (CVPR'2023), DINO (ICLR'2023), H-DINO (CVF'2023), Align-DETR (arXiv'2023), DiffusionDet (ArXiv'2023), CO-DETR (ICCV'2023). There are 5 evaluation metrics, namely mAP50, mAP50-95, Params, Gflops, and FPS.
[0261] Table 1 compares the embodiments of the present invention with other state-of-the-art object detection methods
[0262]
[0263]
[0264]
[0265] Experimental verification: The results of this comparative example are shown in Table 1. We comprehensively compared the dual-detection network of the embodiment of the present invention with 20 classic object detection methods. The evaluation metrics covered mAP50, mAP50-95, Params, Gflops, and FPS. The experimental results show that the student network and the teacher network of object detection in this embodiment perform excellently in multiple metrics. Especially in terms of precision (mAP50), the student network (90.1%) and the teacher network (90.0%) outperform all other methods, especially showing excellent performance in comparison with VarifocalNet (88.7%) and DiffusionDet (89.7%). In terms of mAP50-95, the teacher network (62.8%) and the student network (61.5%) also exceed most classic methods such as Deformable DETR (55.8%) and Cascade R-CNN (56.5%), demonstrating their advantages in multi-scale object detection. In terms of model complexity, although the number of parameters in this embodiment (44.8M for the teacher network and 36.6M for the student network) and the computational cost (108.1 Gflops for the teacher network and 61.4 Gflops for the student network) are relatively high compared to lightweight models such as EfficientDet (11.9M) and GFL (19.24M), a relatively optimized performance balance is still maintained. Finally, in terms of inference speed (FPS), the student network (56.4 FPS) is significantly superior to Double-Head R-CNN (15.9 FPS) and H-DINO (19.2 FPS), indicating its significant advantage in terms of efficiency. Overall, this embodiment achieves a good balance among precision, efficiency, and computational complexity, demonstrating its excellent performance in object detection tasks.
[0266] Comparative Example 2:
[0267] This comparative example presents a high-precision context-aware object detection teacher network, which finally forms an embodiment by gradually adding the parallel context-aware feature network PCAFNet in step S201 and the context-aware attention-enhanced feature fusion module (CAAFM) in step S202. There are 6 evaluation metrics, namely Precision, Recall, mAP50, mAP50-95, Params (M), and Gflops.
[0268] Table 2 Detection results of the teacher network after gradually adding modules
[0269]
[0270] Experimental verification: The results of this comparative example are shown in Table 2, which demonstrates the performance changes of the teacher network after gradually adding the parallel context-aware feature network PCAFNet and the context-aware attention-enhanced feature fusion module CAAFM. It can be seen from the experimental results that the baseline model performs well in terms of Precision (86.0%) and Recall (84.1%). However, after adding PCAFNet, Precision and Recall are respectively improved to 86.4% and 85.2%, indicating the effectiveness of PCAFNet in enhancing feature representation. After further adding the CAAFM module, Precision and Recall reach 89.6% and 86.7% respectively, with improvements of 3.6% and 2.6% compared to the baseline. In terms of mAP50, the baseline model is 88.4%, which is improved to 88.9% after adding PCAFNet, and when it comes to the method of this embodiment, mAP50 reaches 90.0%, demonstrating the continuous improvement of the model accuracy. In terms of mAP50-95, the baseline is 59.0%, PCAFNet is improved to 59.8%, and the final method of this embodiment reaches 62.8%, indicating that by gradually adding modules, the performance of the model on different difficulty samples has been significantly improved. Although the number of model parameters (Params) increases from 38.6M of the baseline to 44.8M, with the introduction of the PCAFNet and CAAFM modules, Gflops also increases to 108.1, showing that while improving performance, the computational complexity also increases. Generally speaking, by gradually adding the PCAFNet and CAAFM modules, the teacher network of this embodiment has achieved significant improvements in various indicators, especially in terms of precision and recall, demonstrating its superior performance in object detection tasks.
[0271] Comparative Example 3:
[0272] This comparative example presents an efficient context-aware object detection student network, which finally forms an embodiment by gradually adding the parallel context-aware feature network PCAFNet in step S201, the student-context-aware attention-enhanced feature fusion module (S-CAAFM) in step S4, and model distillation in step S5. There are 6 evaluation indicators, namely Precision, Recall, mAP50, mAP50-95, Params (M), and Gflops (computational volume).
[0273] Table 3 Detection results of the student network after gradually adding modules
[0274]
[0275]
[0276] Experimental verification: The results of this comparative example are shown in Table 3, which shows the performance changes of the student network after gradually adding the parallel context-aware feature network PCAFNet, the student-context-aware attention-enhanced feature fusion module S-CAAFM, and model distillation. The Precision of the baseline model is 86.0%, Recall is 84.1%, mAP50 is 88.4%, and mAP50-95 is 59.0%. By adding PCAFNet, Precision is improved to 86.4%, Recall is improved to 85.2%, and mAP50 is improved to 88.9%, indicating that PCAFNet effectively improves the feature expression ability. After further adding the S-CAAFM module, Precision and Recall reach 87.6% and 86.2% respectively, mAP50 is 89.0%, and mAP50-95 is improved to 60.3%. Finally, after model distillation, the student network of this embodiment reaches 89.1% Precision, 86.3% Recall, 90.1% mAP50, and 61.5% mAP50-95, demonstrating a further improvement in precision and recall. Although Params(36.6M) and Gflops
[0277] (61.4) are the same as the network after adding S-CAAFM, the overall performance is significantly improved, proving that model distillation maintains high performance while compressing the model.
[0278] Comparative Example 4:
[0279] Based on the efficient context-aware object detection student network of Comparative Example 3, this comparative example further presents the step-by-step addition of the pseudo-depth enhanced Kalman filter (P-DEKF) in step S6 and the post-processing of the beef cattle multi-object tracking method (AFLink and Interpolation) in step S8 during the construction of the efficient context-aware object detection student network of the embodiment of the present invention, finally constituting the embodiment. There are 4 evaluation indicators, namely HOTA, MOTA, IDF1, and ID Switch for evaluation.
[0280] Table 8 Ablation Experiment of Dual Detection and Beef Cattle Multi-Object Tracking Method Driven by Pseudo-Depth Motion Model
[0281] Experiment HOTA / % MOTA / % IDF1 / % IDS Baseline 63.679 75.248 75.432 99 +P-DEKF 64.993 75.314 77.256 96 +AFLink 65.015 75.32 77.737 92 +Interpolation Method of this Embodiment 65.588 75.77 77.98 81
[0282] Experimental verification: The results of this comparative example are shown in Table 8. In this comparative example, gradually adding the P-DEKF, AFLink, and Interpolation modules significantly improved the multi-object tracking performance of the model. In the baseline model, HOTA was 63.679%, MOTA was 75.248%, IDF1 was 75.432%, and ID Switch was 99. After adding P-DEKF, HOTA slightly increased to 64.993%, MOTA increased to 75.314%, IDF1 increased to 77.256%, and ID Switch decreased to 96, indicating that P-DEKF effectively improved the prediction accuracy and robustness of the model and reduced the occurrence of ID switches. Then, after adding the AFLink module, HOTA further increased to 65.015%, MOTA was 75.32%, IDF1 increased to 77.737%, and ID Switch decreased to 92, showing that AFLink optimized the correlation between targets, effectively improved the tracking accuracy, and reduced ID switches. Finally, by adding the Interpolation module, the HOTA of the model reached 65.588%, MOTA was 75.77%, IDF1 increased to 77.98%, and ID Switch further decreased to 81, showing significant improvements in stability and consistency. In summary, gradually adding the P-DEKF, AFLink, and Interpolation modules significantly improved the multi-object tracking performance, especially in terms of metrics such as accuracy, IDF1, and ID Switch, demonstrating the effectiveness of optimizing model accuracy and reducing ID switches. This series of optimization measures indicates that by combining P-DEKF with multi-object post-processing methods, the multi-object tracking ability of the model in complex scenarios can be greatly improved.
[0283] Comparative Example 5:
[0284] In this comparative example, the proposed multi-object tracking method for beef cattle driven by dual detection and pseudo-depth motion model was compared with 6 advanced multi-object tracking methods, specifically DeepOCSORT, BoTSORT, OCSORT, HybridSORT, ByteTrack, and StrongSORT. There were 4 evaluation metrics, namely HOTA, MOTA, IDF1, and ID Switch for evaluation.
[0285] Table 9 Comparative Experiment of Multi-Object Tracking Method for Beef Cattle Driven by Dual Detection and Pseudo-Depth Motion Model
[0286] Experiment HOTA / % MOTA / % IDF1 / % IDS Method of this Embodiment 65.588 75.77 77.98 81 DeepOCSORT 58.773 74.849 69.631 216 BoTSORT 61.819 75.505 74.43 162 OCSORT 59.606 75.351 70.52 205 HybridSORT 56.649 74.957 67.188 204 ByteTrack 58.674 73.368 70.87 158 StrongSORT 54.726 56.717 61.954 255
[0287] Experimental verification: The results of this comparative example are shown in Table 9. The method of this embodiment significantly leads in the HOTA index, indicating that the method can maintain a high target matching accuracy and motion trajectory consistency in the motion and occlusion scenarios of beef cattle groups. Compared with the sub-optimal methods BoTSORT (65.588%) and OCSORT (59.606%), this method has better tracking accuracy, especially in complex backgrounds. In terms of the MOTA index, the performance of the method of this embodiment (75.77%) is similar to that of BoTSORT (75.505%) and OCSORT (75.351%), with a slight lead. MOTA comprehensively considers factors such as missed detections, mis-matches, and identity switches. This method can effectively reduce missed detections and mis-matches, especially in complex environments and when the beef cattle group is relatively dense. The IDF1 index reflects the identity consistency of the target. The method of this embodiment (77.98%) performs the most prominently in this index, showing the advantage of this method in dealing with target identity consistency. In contrast, the IDF1 values of ByteTrack (70.87%) and BoTSORT (74.43%) are relatively low, indicating that they face certain challenges in dealing with the identity consistency of beef cattle targets. In terms of the IDS index, the performance of the method of this embodiment (81) is significantly better than other methods. The lower the IDS, the fewer identity switches and the more stable the tracking. Compared with the sub-optimal methods ByteTrack (158) and BoTSORT (162), this method has an obvious advantage in reducing identity switches and can effectively avoid identity switch problems in complex backgrounds.
[0288] Comparative Example 6:
[0289] In this comparative example, the proposed beef cattle multi-object tracking method driven by double detection and pseudo-depth motion model is visually compared with 6 advanced multi-object tracking methods in the same beef cattle video frames. The specific methods for comparison are DeepOCSORT, BoTSORT, OCSORT, HybridSORT, ByteTrack, and StrongSORT.
[0290] Experimental verification: The visualization of this comparative example is as Figure 11 shown. In the visualization results, the method of this embodiment demonstrates significant advantages. Compared with other methods, the method of this embodiment can provide more stable and accurate tracking results under the same conditions, showing lower mis-matches and target identity switches, and is particularly suitable for complex beef cattle tracking scenarios.
Claims
1. A multi-object tracking method for beef cattle driven by double detection and pseudo-depth motion model, the method comprising the following steps: Step S1, constructing a beef cattle detection and multi-object tracking data set; Characterized in that: Step S2, implementing a high-precision context-aware object detection teacher network HP-CAOD-TN; Step S3, training the high-precision context-aware object detection teacher network HP-CAOD-TN; Based on the beef cattle detection and multi-object tracking data set obtained in step S1, training the high-precision context-aware object detection teacher network HP-CAOD-TN obtained in step S2; Step S4, implementing an efficient context-aware object detection student network EC-CAOD-SN; In the high-precision context-aware object detection teacher network HP-CAOD-TN trained in step S3, replacing the context-aware boundary fusion pyramid module CABFP in the context-aware attention enhancement feature fusion module CAAFM with a student-context-aware boundary fusion pyramid S-CABFP module to form a student-context-aware attention enhancement feature fusion module S-CAAFM, thereby implementing an efficient context-aware object detection student network EC-CAOD-SN; Step S5, model distillation: Implementing model distillation by calculating the similarity between the student feature map of the efficient context-aware object detection student network EC-CAOD-SN obtained in step S4 and the teacher feature map of the high-precision context-aware object detection teacher network HP-CAOD-TN obtained in step S3, and finally obtaining a target double detection network; Step S6, implementing a pseudo-depth enhanced Kalman filter P-DEKF; Integrating pseudo-depth information into the state vector of the Kalman filter, and expanding the dimension of the state vector from 8 dimensions (x, y, w, h, vx, vy, vw, vh) to 10 dimensions (x, y, w, h, d, vx, vy, vw, vh, vd); Wherein: x, y represent the central position coordinates of the target bounding box; w, h represent the width and height of the target bounding box; d represents the pseudo-depth value, that is, represents the distance of the target from the camera; vx, vy represent the speeds in the x and y directions of the position; vw, vh represent the change speeds of the width and height; vd represents the change speed of the pseudo-depth; Step S7, a multi-object tracking method for beef cattle driven by double detection and pseudo-depth motion model: Based on the target double detection network obtained in S5, using the pseudo-depth enhanced Kalman filter P-DEKF in step S6 to simulate the state of the target for pseudo-depth motion model driving, and using the Hungarian algorithm for frame-by-frame data association to jointly achieve multi-object tracking.
2. The multi-object tracking method for beef cattle driven by double detection and pseudo-depth motion model as described in claim 1, wherein In step S2, the key components of the high-precision context-aware object detection teacher network HP-CAOD-TN include: a parallel context-aware feature network PCAFNet, a context-aware attention enhancement feature fusion module CAAFM, an IoU-aware query selection mechanism, and a decoder with an auxiliary prediction head; The described parallel context-aware feature network PCAFNet consists of a basic convolutional module and a PCAF module; The described context-aware attention-enhanced feature fusion module CAAFM combines an intra-scale feature interaction AIFI module based on the attention mechanism and a context-aware boundary fusion pyramid module CABFP; The described context-aware boundary fusion pyramid CABFP includes a dual-pool channel attention module DPCA, a triple attention fusion module TAFM, a dual gating mechanism to control the feature fusion module DGAFM, and a spectrum enhancement fusion module SEFM.
3. The multi-object tracking method for beef cattle driven by double detection and pseudo-depth motion model as described in claim 1, wherein In step S2, the construction method of the high-precision context-aware object detection teacher network HP-CAOD-TN is as follows: Step S201, construct a parallel context-aware feature network PCAFNet: Send the input image data into the parallel context-aware feature network PCAFNet to generate feature maps at four levels of {P2, P3, P4, P5}; Step S202, construct a context-aware attention-enhanced feature fusion module CAAFM: Input the P5 feature map obtained in step S201 into the internal interaction of the intra-scale feature interaction module AIFI of the attention mechanism, and output the interacted F5 feature map; Input the P2 feature map, P3 feature map, P4 feature map, and F5 feature map obtained in step S201 into the context-aware boundary fusion pyramid CABFP to perform cross-scale feature fusion stitching on the P2 feature map, P3 feature map, P4 feature map, and F5 feature map, and output the Memory feature map; Step S203, construct an IoU-aware query selection mechanism: Calculate the regression and classification losses of the predicted bounding box corresponding to each feature in the Memory feature map obtained in step S202 and the ground truth bounding box, and add the IOU to the classification loss to ensure the consistency of the position confidence and class confidence of the network output. Select the top-K features according to the loss; Step S204, construct a decoder with an auxiliary prediction head: The Decoder receives the top-K features screened in step S203, performs initial object queries and the Memory feature sequence output by the hybrid encoder, and iteratively optimizes the object queries through multi-layer self-attention and cross-attention, and finally outputs the category and coordinates of the detection box; The Head maps the category and coordinates of the detection box output by the Decoder onto the detection box.
4. The multi-object tracking method for beef cattle driven by double detection and pseudo-depth motion model as described in claim 1, characterized in that, In step S3, the loss function L used in training includes the target bounding box regression loss L box and the classification loss L cls in two parts, that is, L = L box + L cls ; Backpropagate the loss function L and repeat the iteration until the number of iterations reaches the preset initial value to complete the training.
5. The multi-object tracking method for beef cattle driven by double detection and pseudo-depth motion model as described in claim 1, characterized in that, In step S5, the specific process of the model distillation is as follows: Step S501, feature map alignment operation; Step S502, standardize the feature map; Step S503, calculate the similarity of the feature map; Step S504: Standardize the correlation matrix; Step S505, calculate the loss; Step S506, weighted loss; Step S507, backpropagation and optimization.
6. The multi-object tracking method for beef cattle driven by double detection and pseudo-depth motion model as described in claim 1, characterized in that, In step S6, the described extension method includes the following steps: Step S60, define the state vector; Step S602, establish a motion model; Step S603, establish an observation model; Step S604, initialize the pseudo-depth enhanced Kalman filter P-DEKF; Step S605, prediction; Step S606, update; Step S607, calculate the gating distance; Step S608, calculate the pseudo-depth.
7. The multi-object tracking method for beef cattle driven by double detection and pseudo-depth motion model as described in claim 1, characterized in that The specific process of Step S1 is as follows: Collect videos of five different breeding scenarios from a real monitoring perspective, use the DarkLabel tool to annotate the videos, obtain the annotation data in the standard data format, and convert the annotation data into a beef cattle detection and multi-object tracking dataset.
8. The multi-object tracking method for beef cattle driven by double detection and pseudo-depth motion model as described in claim 1, characterized in that This method further includes Step S8, post-processing: Adopt the linear interpolation method to fill the missing detection boxes at the missing detection positions in each trajectory.
Citation Information
Cited By
Coal mining machine roller tracking system and method
CN121010931A