Multimodal group pig individual tracking method and system based on spatial prior guidance

By combining video images and IMU signals, using momentum correction particle filtering algorithm and cross attention mechanism, the problems of missed detection, false detection and ID jump in tracking of individual pigs in group pigs are solved, and efficient tracking of individual pigs in group pigs is achieved, improving the accuracy and real-time tracking of individual pigs in group pigs are improved.

CN118968542BActive Publication Date: 2025-08-19HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410931370.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-12
Publication Date
2025-08-19
Estimated Expiration
2044-07-12

AI Technical Summary

Technical Problem

In actual breeding scenarios, existing individual tracking technology of pig farming is susceptible to factors such as similar appearance, occlusion interaction and lighting changes, resulting in frequent jumps in missed detection, false detection and IDs. Video monitoring technology relies on high frame rates to increase costs and has great real-time challenges, while IMU sensors cannot provide pixel-level spatial information and are susceptible to noise interference.

Method used

Combining video images and IMU signals, the pig position is initially estimated through a particle filtering algorithm based on momentum correction, and a detection box decoder with cross attention mechanism is established using spatial prior information to realize the complementary relationship between the two modal data to achieve continuous and accurate tracking of individuals in pig farming.

Benefits of technology

It effectively alleviates the missed detection and mis-detection problems of video monitoring technology, improves the accuracy and real-time tracking, makes full use of the position advantages of IMU signals and the boundary advantages of video images, and reduces the computing burden.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118968542B_ABST
    Figure CN118968542B_ABST
Patent Text Reader

Abstract

The present invention discloses a multimodal group pig individual tracking method and system based on spatial prior guidance. The method comprises the following steps: S1, establishing a momentum-corrected particle filter algorithm based on high-frequency IMU signals, converting the high-frequency IMU signals into spatial position coordinates to map the IMU signals and video images to the same space; S2, establishing a group pig individual tracking algorithm guided by spatial position priors, extracting the features of the spatial position obtained in step S1 and the features of the corresponding video frame image, constructing a detection frame decoder based on the cross-attention mechanism, and establishing a complementary relationship between the two modal data to achieve continuous and accurate tracking of group pig individuals. The present invention preliminarily estimates the spatial position of individual pigs through IMU signals, and uses this spatial prior information to guide the detection of individual pigs in the corresponding video frame image, thereby achieving continuous tracking of group pig individuals in the video stream.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of multi-target tracking, and in particular relates to a multi-modal group pig individual tracking method and system based on spatial priori guidance. Background Art

[0002] China is the world's leading producer of live pigs and consumer of pork, with pork consumption accounting for approximately 60% of the country's total meat consumption. [1] Shen Mingxia, Chen Jinxin, Ding Qi'an, et al. Research progress and prospects of automated pig farming equipment and technology [J]. Journal of the Chinese Society of Agricultural Machinery, 2022, 53(12): 1-19.). However, the top of my country's pig farming industry chain - the supply of high-quality breeding pigs, has long been highly dependent on imports. According to statistics, in recent years, my country's annual import volume from abroad has exceeded 20,000 heads ( [1] Zhang Haifeng, Xie Kengzheng, Liu Xiaohong, et al. Utilization and existing problems of imported breeding pigs in my country [J]. Chinese Journal of Animal Husbandry, 2023, 59(01): 330-332. DOI: 10.19556 / j.0258-7033.20220918-20.). Excessive reliance on external resources has led to a "bottleneck" dilemma facing the domestic pig breeding industry, that is, key technologies and resources are controlled by others. Breaking this constraint and achieving independent control of the breeding industry is crucial to promoting the transformation of my country's pig breeding industry from scale advantages to a balance between technology and quality. At the same time, the health status of breeding pigs directly affects the quality of domestic breeding pigs and is also related to the economic benefits and sustainable development of the entire industry.

[0003] During the growth process of pigs, their daily behaviors, such as eating, drinking and excretion, contain rich health information and are important indicators for measuring their health status and welfare level ( [1] Shen Mingxia, Chen Jinxin, Ding Qi'an, et al. Research progress and prospects of automated pig breeding equipment and technology [J]. Transactions of the Chinese Society of Agricultural Machinery, 2022, 53(12): 1-19.). Traditional pig behavior monitoring mainly relies on manual labor, which is time-consuming and labor-intensive and has the risk of cross-infection between humans and animals. The rapid development of digital technologies such as computer vision and wearable sensors, as well as artificial intelligence technologies such as machine learning and deep learning, has now successfully promoted the innovation and upgrading of automated and intelligent pig behavior recognition systems, greatly improving the efficiency of breeding pigs, while reducing labor costs and effectively suppressing the risk of pig disease transmission ( [3] He Peitong, Zhang Jianhua, Zhang Ning, Xia Xue, Chai Xiujuan. Research progress in intelligent livestock and poultry breeding management and disease diagnosis based on visual perception [J]. Journal of China Agricultural University, 2023, 28(10): 141-165.).

[0004] Multi-target tracking of group-raised pigs is the basis for accurate identification and statistical analysis of individual pig behaviors, and plays an important role in promoting the improvement of refined and digital management of pigs. The current mainstream multi-target tracking of group-raised pigs relies on video imaging technology. Among them, multi-target tracking algorithms are gradually being applied to pig video monitoring because they can effectively locate multiple targets in the video based on biometrics and track the movement path of each pig. [2] Chen, C., Zhu, W., Steibel, J., Siegford, J., Han, J., & Norton, T. (2020). Recognition of feeding behavior of pigs and determination of feeding time of each pig by a video-based deeplearning method. Computers and Electronics in Agriculture, 176, 105642.). Multi-target tracking is mainly divided into tracking by detection (TBD) and joint detection and embedding (JDE). The former performs detection and tracking in stages, resulting in the model's reasoning speed being unable to achieve real-time effects. In contrast, the JDE algorithm uses an end-to-end network to simultaneously perform target detection and tracking, which can achieve higher accuracy and faster target tracking effects ( [3] Wang, Z., Zheng, L., Liu, Y., Li, Y., & Wang, S. (2020, August). Towards real-time multi-object tracking. In European Conference on Computer Vision (pp. 107-122). Cham: Springer International Publishing.). Professor Liang Yun’s team ( [4] Tu Shuqin, Huang Lei, Liang Yun, Huang Zhengxin, Li Chengjie, & Liu Xiaolong. (2022). Multi-target tracking of group-reared pigs based on the JDE model. Transactions of the Chinese Society of Agricultural Engineering, 38(17), 186-195.) It was confirmed that the pig multi-target tracking model using the JDE algorithm has a 340% improvement in the number of frames per second compared to the TBD algorithm model. However, the detector part of the existing JDE model mostly uses the anchor frame mechanism, which is prone to the problem of a single frame covering multiple targets due to overlapping detection frames in the pig adhesion occlusion scene ( [5]Guo, Q., Sun, Y., Orsini, C., Bolhuis, JE, de Vlieg, J., Bijma, P., & de With, PH (2023). Enhanced camera-based individual pig detection and tracking for smart pig farms. Computers and Electronics in Agriculture, 211, 108009.). In response to this, the researchers further optimized the JDE algorithm and proposed the FairMOT algorithm ( [6] Zhang, Y., Wang, C., Wang, X., Zeng, W., Liu, W., 2021. Fairmot: On the fairness of detection and re-identification in multiple object tracking. Int. J. Comput. Vis. 129 (11), 3069–3087.), uses the anchor-free target detection network CenterNet as its target detection branch, and simultaneously extracts the target center ReID feature. Professor Tomas Norton's team at the University of Leuven, Belgium ( [7] Wang, M., Larsen, ML, Liu, D., Winters, JF, Rault, JL, & Norton, T. (2022). Towards reidentification for long-term tracking of group housed pigs. Biosystems Engineering, 222, 71-81.) optimized the FairMOT algorithm to achieve accurate re-identification of pigs when they reappear after being occluded. However, in actual farming scenarios, due to factors such as the similar appearance of pigs, mutual occlusion, and variable lighting conditions, existing methods are inevitably prone to problems such as missed detection, false detection, and frequent ID jumps. At the same time, the tracking process relies too much on inter-frame data correlation, which increases the difficulty of real-time tracking under conditions of fast movement of pigs and low frame rates.

[0005] Another target tracking technology is to use wearable devices, such as electronic pedometers, RFID tags and inertial measurement units (IMUs). These devices are usually fixed to different parts of the target body to obtain individual identity, physiological indicators and behavioral characteristics data, thereby realizing intelligent monitoring of individuals ( [8]He Peitong, Zhang Jianhua, Zhang Ning, Xia Xue, Chai Xiujuan. Research progress on intelligent livestock and poultry breeding management and disease diagnosis based on visual perception [J]. Journal of China Agricultural University, 2023, 28(10): 141-165.). Among them, the IMU sensor integrates accelerometers, gyroscopes and magnetometers. Because of its small size, portability, independence from external facilities, easy deployment, and unaffected by environmental factors such as occlusion and light, it has been widely used in the field of indoor pedestrian positioning. The original inertial navigation system ( [9] Weston J, Titterton D.Modern inertial navigation technology and its application[J].ElectronicsCommunication Engineering Journal,2000,12(2):49-64.)(Strapdown inertial navigation integration, SINS) is based on the integration of acceleration and angular velocity for positioning, but due to sensor noise interference, the error during independent dead reckoning will accumulate rapidly. In response to this, researchers proposed a pedestrian dead reckoning based on step length (

[10] Kang W, Han Y. SmartPDR: Smartphone-based pedestrian dead reckoning for indoor localization [J]. IEEE Sensors Journal, 2014, 15 (5): 2906-2916.) (Pedestrian Dead Reckoning, PDR) algorithm uses an accelerometer to detect the number of steps and estimate the step length, while using a gyroscope and a magnetometer to calculate the pedestrian's heading angle to obtain the pedestrian's relative position and thus achieve positioning. However, because it relies on prior knowledge of human walking patterns and needs to adjust parameters according to the user's individual gait characteristics, this hinders its widespread application in real life. Wang and Shkel et al. (

[11] Wang Y, Shkel AM. Adaptive threshold for zero-velocity detector in ZUPT-aided pedestrian inertial navigation [J]. IEEE Sensors Letters, 2019, 3 (11): 1-4.) proposed applying the Bayesian method to pedestrian dead reckoning to determine the adaptive threshold estimation step information. In addition, the magnetometer signal is easily interfered by ferromagnetic materials, and the gyroscope angle measurement has drift, which may affect the accuracy of heading estimation. To improve the accuracy of heading estimation, a common solution is to use a Kalman filter to integrate the magnetometer and gyroscope readings (

[12] Shang J, Gu F, Hu X, et al. APFiLoc: An Infrastructure-Free Indoor Localization Method Fusing Smartphone Inertial Sensors, Landmarks and Map Information [J]. Sensors, 2015, 15(10): 27251-27272.), but the classical Kalman filter can only achieve optimal estimation in an ideal linear Gaussian system. Although the extended Kalman filter and the unscented Kalman filter solve the nonlinear problem, they are still limited by the fact that the system is Gaussian noise. In order to solve the problem of nonlinear non-Gaussian system, the particle filter algorithm based on the Monte Carlo method was proposed (

[13] Liu H, Li Q, Li C, et al. Application research of an array distributed IMU optimization processing method in personal positioning in large span blind environment [J]. IEEE Access, 2020, 8: 48163-48176.) This algorithm uses a large number of particles to approximate the true posterior distribution. Due to its reliable performance, it has been widely used in indoor positioning fields such as integrated navigation. Currently, IMU sensors are mainly used for indoor positioning of humans, and their application in the animal field is still underdeveloped. Given the successful case of this technology in human positioning, its application in farms to achieve indoor positioning of animals has great potential and broad prospects.

[0006] In summary, the advantage of video imaging-based methods lies in their ability to continuously capture scene dynamics and provide high-resolution pixel-level information. However, they cannot effectively overcome the effects of pigs' similar appearance, occlusion interactions, and changing lighting conditions. In contrast, IMU-based methods do not rely on external equipment and are unaffected by factors such as occlusion and changing lighting. They can provide accurate identification and preliminary position estimates for targets, and are expected to provide accurate identification and preliminary position estimates for individual pigs. However, this method lacks the fine-grained spatial resolution of images, and data collected in complex scenes is susceptible to interference from measurement noise and hardware noise.

[0007] Video monitoring technology faces numerous challenges in actual farming scenarios. First, external conditions such as variable lighting can degrade image quality and blur target outlines, leading to missed detections during individual tracking. Second, internal factors such as pigs' similar appearance, mutual occlusion, and shifting postures can lead to false detections and frequent ID changes. Third, high-precision behavior recognition often relies on high-frame-rate video, which not only increases equipment operating costs but also significantly challenges the real-time performance of pig behavior recognition.

[0008] IMU sensors also face numerous challenges in actual farming scenarios. First, they cannot provide users with intuitive and easy-to-understand visualization results. Second, they lack the ability to capture pixel-level spatial information and cannot be used to accurately calibrate the positional boundaries of pigs in space. Third, while IMU sensors can detect subtle temporal changes in individual pig behaviors, they cannot capture global spatial changes. Furthermore, motion signals are susceptible to noise, which in turn affects the accuracy of behavior recognition.

[0009] Based on the above situation, the present invention innovatively combines video images and IMU signals, models the complementary relationship between the two, and designs a multimodal group pig individual tracking technology solution based on spatial prior guidance. Summary of the Invention

[0010] In order to solve the above-mentioned technical problems existing in the prior art, the present invention provides a multimodal group pig individual tracking method and system based on spatial prior guidance. The present invention maps two modal data, video images and motion signals, to the same space and models the complementary relationship between the two to achieve robust group pig individual tracking.

[0011] The present invention adopts the following technical solutions:

[0012] The multimodal group pig individual tracking method based on spatial prior guidance is as follows:

[0013] S1, establishes a momentum-corrected particle filter algorithm based on the high-frequency IMU signal, converts the high-frequency IMU signal into spatial position coordinates, and maps the IMU signal and the video image to the same space;

[0014] S2: Establish a group pig individual tracking algorithm guided by spatial position prior, extract the features of the spatial position obtained in step S1 and the features of the corresponding video frame image, construct a detection box decoder based on the cross-attention mechanism, and establish a complementary relationship between the two modal data to achieve continuous and accurate tracking of group pig individuals.

[0015] Preferably, step S1 is specifically as follows:

[0016] S1.1. Set the starting position coordinates of the pig target and initialize the particle swarm probability. Perform equal-interval integration operations on the motion signal collected by the IMU sensor to obtain the target's speed and direction information.

[0017] S1.2. The uncorrected speed and direction values at the current moment are uniformly recorded as Ω t , the updated value after the momentum equation is corrected is recorded as The correction value at the previous moment is expressed as Initialization setting is 0; the correction relationship is expressed as follows:

[0018]

[0019] Among them, α represents the degree of correction;

[0020] S1.3. Perform zero-speed detection on the collected motion signal data to obtain a zero-speed interval; perform threshold determination using the acceleration modulus, angular velocity modulus, and angular velocity modulus standard deviation;

[0021] S1.4. Take the velocity value solved in the zero-speed interval as the observation quantity, establish the observation equation and discretize it to obtain:

[0022] Z vk =H(X k )+V vk (5)

[0023] Z vk =[V e ,V n ,V u ] (6)

[0024] V vk =[ΔV e ,ΔV n ,ΔV u ] (7)

[0025] Where V e、V n 、V u are the velocities in the east, north and sky directions respectively; H is the system observation matrix; ΔV e , ΔV n , ΔV u are the velocity observation noises in the east, north, and sky directions, respectively;

[0026] S1.5. Use particle filter algorithm to calculate the individual position coordinates of pigs.

[0027] Preferably, in step S1.3:

[0028] Acceleration modulus conditions:

[0029]

[0030] Where, |a i | is the axis acceleration modulus; a xi ,a yi ,a zi are the acceleration measurements of the x, y, and z axes respectively; a min ,a max are the upper and lower limits of the acceleration range respectively;

[0031] Angular velocity modulus conditions:

[0032]

[0033] Where, |ω i | is the angular velocity modulus; ω xi ,ω yi ,ω zi are the angular velocity measurements of the x, y, and z axes respectively; ω max is the angular velocity interval threshold;

[0034] Acceleration variance condition:

[0035]

[0036] Where, For|a i | is the standard deviation of the acceleration modulus at the sampling midpoint, and 2n+1 is the length of the sampling interval; is the mean value of acceleration modulus; σ a is the standard deviation threshold of the acceleration modulus;

[0037] When C1, C2, and C3 conditions all meet the three conditions at the same time, that is, C1(i)C2(i)C3(i)=1, the sampling point i is determined to be a zero-speed point, and the continuous zero-speed points constitute a zero-speed interval.

[0038] Preferably, step S1.5 is as follows:

[0039] S1.5.1. At the initial time t = 0, the initial particles x0 ~ p(x0) are sampled according to the sampling distribution p(x0) to obtain the initial particle set Take the number of particles N = 100, and the initial weight of each particle is 1 / N;

[0040] S1.5.2. Constructing the importance probability density function Calculate each particle according to the following formula Corresponding weight

[0041]

[0042] Where, is the posterior probability density function; is the likelihood function of the observed quantity; is the state transition probability; is the proposed distribution function, satisfying is the weight at the k-1 moment; is the weight of particle k at time;

[0043] Normalize it and get:

[0044]

[0045] Where, is the weight of the i-th particle at the k-th moment; is the weight sum of all particles at time k; is the value after weight normalization;

[0046] S1.5.3. Based on the normalized particle set, the state value at time k is obtained from the state value at time k-1, that is, the optimal estimate. The state estimate is:

[0047]

[0048] Where, is the particle at time k;

[0049] S1.5.4. Discard particles with weights less than the set value from the particle set, re-extract particles with weights greater than the set value and copy them to establish a new particle set.

[0050] Preferably, in step S1.5.4, an adaptive threshold method is used to select new particles, and the specific steps are as follows:

[0051] 1) Select an adaptive threshold M; M satisfies equation (11);

[0052]

[0053] Where: The value range of M is [M a ,M b ]; M b is the upper limit; M a is the initial value of the threshold, t is the number of iterations; N e is the effective particle number; K e is the proportional coefficient; the threshold M is proportional to the number of iterations t and the number of effective particles N e Inversely proportional to N e Satisfy formula (12);

[0054]

[0055] Using the effective number of particles N e To measure the degree of degradation of particle weights, N is the number of particles, is the particle weight variance; Substituting (12) into (11) we get:

[0056]

[0057] When M reaches the upper limit M b When M=M a ;

[0058] 2) Retain particle judgment; if Keep otherwise

[0059] 3) Make a secondary judgment on particle degradation:

[0060]

[0061] When H1+H2=1, return to step S1.5.1 to resample the particles. The weight of the resampled particles is 1 / N, where N is the number of particles. The new particle set As the initial particle at the next moment; otherwise, return to step S1.5.2 for the next iteration until the loop ends; the position of the corresponding particle at the end of the loop is the final target position.

[0062] Preferably, step S2 is specifically as follows:

[0063] S2.1, project the position and image into the feature space of the same dimension;

[0064] S2.2. Establish the complementary relationship between IMU signal and video image; and location characteristics They are input together into the detection frame decoder based on the cross attention mechanism, and the individual detection frame features corresponding to the M-dimensional position feature vector are output. To achieve the integration of information between the two; the specific process is as follows:

[0065] P′=LN(Attn(P,P,P)+P) (19)

[0066] P″=LN(Attn(P,Z,Z)+P′) (20)

[0067] P″′=LN(MLP(P″)+P″) (21)

[0068] Z′=LN(Attn(Z,P″′,P″′)+Z) (22)

[0069] B=LN(Attn(P″′,Z′,Z″)+P″′) (23)

[0070] Where, LN(·) denotes layer normalization, and MLP(·) denotes multilayer perceptron;

[0071] The operation module in the middle of the decoder contains two position-image attention modules, one image-position attention module and a multi-layer perceptron; among them, the position-image attention module replaces the multi-head self-attention block with a cross attention block based on the Transformer layer. Realization, Q = w Q P, K = w K Z, V = w V Z; in the image-position attention module, Q = w′ Q Z, K = w′ K P, V = w′ V ·P, w Q , w K , w V , w′ Q , w′ K , and w′ V Represents the mapping matrix that can be learned;

[0072] S2.3. Obtained individual detection frame features Input to the multi-layer perceptron, and the corresponding output is the key point coordinate pair of each pig's individual bounding box, recorded as and They represent the upper left corner coordinates and lower right corner coordinates of the m-th pig's bounding box respectively.

[0073] Preferably, step S2.1 is specifically as follows:

[0074] S2.1.1, based on the position coordinates p of the individual pig at a certain moment obtained in step S1 m =(x m ,y m ), m=1,2,…,M, and input it into the learnable position encoder to generate the position features corresponding to M pigs

[0075] The position encoder is composed of a function Y based on a random Fourier feature map and a multilayer perceptron with a ReLU activation function; Function Y is expressed as:

[0076]

[0077] Where, Represents the parameters that can be learned, basis vector w i It is randomly sampled from a Gaussian distribution with 0 as mean and σ as standard deviation, that is,

[0078] S2.2.2. Select the video frame image of the corresponding moment from the video, divide it into N fixed-size image blocks, generate a vector representation of dimension N×D through flattening and mapping operations, add it to the corresponding N×D-dimensional position embedding vector and input it into the Transformer encoder to obtain the image feature vector

[0079] Preferably, the specific process of step S2.2.2 is as follows:

[0080] S2.2.2.1 Video frame image Divided into N fixed-size (P 2 ×C) image block, denoted as N=HW / P 2 , C represents the number of channels; each image block is flattened into a one-dimensional vector, and a set of image block vector representations with dimension D is generated through linear mapping Vectorize the position of the image patch to the same dimension as that of the image patch Add up to get input tokens, denoted as

[0081] S2.2.2.2. Input into the Transformer encoder with a total of L layers to obtain a new set of image block codes Recorded as image feature vector Each Transformer layer mainly consists of a multi-head self-attention block and a multi-layer perceptron block, each of which contains a layer normalization layer and a residual connection:

[0082]

[0083] The present invention also discloses a multimodal group pig individual tracking system based on spatial prior guidance, which is based on the above method and includes the following modules:

[0084] Position acquisition module: A momentum-corrected particle filter algorithm is established based on the high-frequency IMU signal to convert the high-frequency IMU signal into spatial position coordinates, so as to map the IMU signal and the video image to the same space;

[0085] Tracking module: Establish a group pig individual tracking algorithm guided by spatial position prior, extract the features of the obtained spatial position and the features of the corresponding video frame image, construct a detection box decoder based on the cross-attention mechanism, and establish a complementary relationship between the two modal data to achieve continuous and accurate tracking of group pig individuals.

[0086] This multimodal group pig tracking method and system, guided by spatial priors, uses IMU signals to initially estimate the spatial position of individual pigs. This spatial prior information guides the detection of individual pigs in corresponding video frames, enabling continuous tracking of group pigs within a video stream. This method leverages the advantages of IMU signals in initially estimating individual positions and video images in defining target boundaries, effectively alleviating the problems of missed detection, false detection, and ID jumps that can occur with video monitoring technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0087] Figure 1 This is a general flow chart of a multimodal group pig individual tracking method based on spatial prior guidance according to a preferred embodiment of the present invention.

[0088] Figure 2 It is a flow chart of the particle filter algorithm principle of the preferred embodiment.

[0089] Figure 3 Schematic diagram of a particle filter algorithm based on momentum correction in a preferred embodiment.

[0090] Figure 4 This is a schematic diagram of an algorithm for tracking individual pigs in group raised guided by spatial position priori in a preferred embodiment.

[0091] Figure 5 This is a block diagram of a multimodal group pig individual tracking system based on spatial prior guidance in a preferred embodiment of the present invention. DETAILED DESCRIPTION

[0092] The preferred embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0093] In actual farming scenarios, the video-based group pig individual tracking technology is easily affected by factors such as the similar appearance of pigs, occlusion interaction, and changes in lighting conditions. The IMU signal is not easily affected by these factors and can accurately provide individual pig position information, but it cannot directly and accurately define the target boundary of the pig in the visual space. In addition, given that motion signals are lighter than video data, in order to reduce the computational burden of multimodal data on the model, and taking into account the real-time requirements in practical applications, the present invention adopts a method of combining low-frame-rate video with high-frequency motion signals to establish a spatial prior-guided multimodal group pig individual tracking method, such as Figure 1 shown.

[0094] Firstly, a momentum-corrected particle filter algorithm is established based on the high-frequency IMU signal, and the high-frequency IMU time series signal is converted into spatial position coordinates, so that the IMU signal and the video image can be mapped to the same space; on this basis, a group-raised pig individual tracking algorithm guided by spatial position prior is established, and the features of the above-mentioned prior position and the features of the corresponding video frame image are extracted. A detection box decoder based on the cross-attention mechanism is constructed, and the complementary relationship between the two modal data is modeled to promote their deep fusion, thereby realizing continuous and accurate tracking of group-raised pig individuals.

[0095] The specific technical solutions of this embodiment are as follows:

[0096] S1. Establish a particle filter algorithm based on momentum correction

[0097] In the field of indoor positioning based on IMU signals, the widely used algorithm for solving the position of the positioning target is the particle filter algorithm based on the Monte Carlo method. This algorithm uses a particle swarm to represent probability and approximates the target motion equation by finding particles propagating in the state space, thereby obtaining the target position. Figure 2 As shown, its working principle is to estimate the state of a dynamic system from a set of observation sequences using a set of particles with different weights. The observation sequences used may contain noise, or even be incomplete, and the weights of the particles will be continuously updated iteratively. The particle filter algorithm can be applied to any state space model, and the noise distribution can also be in any form. It has great advantages in the scenario of nonlinear non-Gaussian systems. In combination with the application scenario of the present invention, considering that the data collected by the IMU sensor is easily interfered by measurement noise and hardware noise, and the daily life of pigs is mainly static behaviors such as lying down and sleeping, the present invention proposes and adopts a particle filter algorithm based on momentum correction to preliminarily estimate the spatial position of individual pigs at equal intervals. The steps are as follows. Figure 3 As shown, the specific description is:

[0098] S1.1. Obtain the target's speed and direction information. First, set the starting position coordinates of the pig target and initialize the particle swarm probability. Then, perform equal-interval integration operations on the motion signals collected by the IMU sensor to obtain the target's speed and direction information.

[0099] S1.2, Error Correction. Considering that the instability of the IMU signal will generate noisy data points and the errors generated by these data points will gradually accumulate over time, the present invention uses the momentum equation to correct the errors in the speed and direction information after the integral calculation. The uncorrected speed and direction values at the current moment are uniformly recorded as Ω t , the updated value after the momentum equation is corrected is recorded as The correction value at the previous moment is expressed as Initialization setting is 0. The specific correction relationship is expressed as follows:

[0100]

[0101] Among them, α represents the degree of correction, and the updated It does not depend only on the current value, but also accumulates historical information of the target's movement trajectory.

[0102] S1.3, perform zero-speed detection on the collected data. Considering that pigs mainly lie down and sleep in their daily life, in order to improve the positioning accuracy, the collected motion signal data is firstly tested for zero speed, and then the zero-speed interval is obtained. The acceleration modulus, angular velocity modulus and angular velocity modulus standard deviation are used for threshold judgment. In this step,

[0103] S1.3.1 Acceleration modulus conditions:

[0104]

[0105] Where, |a i | is the axis acceleration modulus; a xi ,a yi ,a zi are the acceleration measurements of the x, y, and z axes respectively; a min ,a max are the upper and lower limits of the acceleration range respectively.

[0106] S1.3.2. Angular velocity modulus conditions:

[0107]

[0108] Where, |ω i | is the angular velocity modulus; ω xi ,ω yi ,ω ziare the angular velocity measurements of the x, y, and z axes respectively; ω max is the angular velocity interval threshold.

[0109] S1.3.3 Acceleration variance conditions:

[0110]

[0111] Where, For|a i | is the standard deviation of the acceleration modulus at the sampling midpoint, and 2n+1 is the length of the sampling interval; is the mean value of acceleration modulus; σ a is the acceleration modulus standard deviation threshold.

[0112] When C1, C2, and C3 conditions all meet the three conditions at the same time, that is, C1(i)C2(i)C3(i)=1, the sampling point i is determined to be a zero-speed point, and the continuous zero-speed points constitute a zero-speed interval.

[0113] S1.4. Obtain the observation equation of the zero-speed correction algorithm. Take the speed value solved in the zero-speed interval as the observation value, establish the observation equation of the system and discretize it to obtain:

[0114] Z vk =H(X k )+V vk (5)

[0115] Z vk =[V e ,V n ,V u ] (6)

[0116] V vk =[ΔV e ,ΔV n ,ΔV u ] (7)

[0117] Where V e ,V n ,V u are the velocities in the east, north and sky directions respectively; H is the system observation matrix; ΔV e ,ΔV n ,ΔV u are the velocity observation noise in the east, north and sky directions respectively.

[0118] S1.5. Use a particle filter algorithm to calculate the individual pig position coordinates. This method uses a motion model in the zero-speed range to find a set of particles to approximate the probability density function. By continuously updating the particle values and weights, the optimal estimate of the state vector is obtained. The algorithm flow is as follows:

[0119] S1.5.1. Initial particle calculation. At the initial time t = 0, the initial particles x0 ~ p(x0) are sampled according to the sampling distribution p(x0) to obtain the initial particle set The number of particles N is taken as 100, and the initial weight of each particle is 1 / N.

[0120] S1.5.2. Importance sampling updates particles and weights. Constructs importance probability density function Calculate each particle according to the following formula Corresponding weight

[0121]

[0122] Where, is the posterior probability density function; is the likelihood function of the observed quantity; is the state transition probability; is the proposed distribution function, satisfying is the weight at the k-1 moment; is the weight of particle k at time.

[0123] Normalize it and get:

[0124]

[0125] Where, is the weight of the i-th particle at the k-th moment; is the weight sum of all particles at time k; is the value after weight normalization.

[0126] S1.5.3. Calculate the filter value. Based on the normalized particle set, the value at time k can be obtained from the state value at time k-1, which is the optimal estimate. The state estimate is:

[0127]

[0128] Where, is the particle at time k.

[0129] S1.5.4. Particle Resampling. Discard particles with smaller weights from the particle set, re-extract particles with larger weights, and replicate them to create a new particle set. Ensure that the mean of the particle set approaches the mathematical expectation with the greatest probability.

[0130] The weight of particles gradually increases during the iterative process. When using a fixed weight threshold to select particles, if there are particles with small weights among the retained particles, the weight threshold is large, and the particles degenerate faster, resulting in filtering failure. Therefore, the present invention adopts an adaptive threshold method to select new particles, and its specific steps are as follows:

[0131] 1) Select an adaptive threshold M. M satisfies equation (11);

[0132]

[0133] Where: The value range of M is [M a ,M b ]; M b is the upper limit; M a is the initial value of the threshold, t is the number of iterations; N e is the effective particle number; K e is the proportional coefficient; the threshold M is proportional to the number of iterations t and the number of effective particles N e Inversely proportional to N e Satisfy formula (12);

[0134]

[0135] Using the effective number of particles N e To measure the degree of degradation of particle weights, N is the number of particles, is the particle weight variance. N e The larger the value, the more serious the degradation. Substituting (12) into (11) we get:

[0136]

[0137] When M reaches the upper limit M b When M=M a .

[0138] 2) Retain particle judgment. If Keep otherwise

[0139] 3) Make a secondary judgment on particle degradation:

[0140]

[0141] When H1+H2=1, return to step S1.5.1 to resample the particles. The weight of the resampled particles is 1 / N, where N is the number of particles. The initial particle at the next moment is used as the initial particle. Otherwise, return to step S1.5.2 for the next iteration until the loop ends. The position of the particle at the end of the loop is the final target position.

[0142] S2. Establishing an algorithm for tracking individual pigs in groups guided by spatial position priors

[0143] The algorithm proposed in the present invention first uses the preliminary position coordinates of the individual pigs in the visual space obtained above, projects the position prior and the information of the corresponding video frame image into the feature space of the same dimension, and designs a detection frame decoder to model the complementary relationship between the two to promote deep fusion of the two, and realize individual pig detection at the single-frame image level; then, based on the preset individual identity association information between video frames, combined with the detection results of the single-frame image, individual detection and tracking in the video stream is realized. The specific process steps are as follows: Figure 4 shown.

[0144] S2.1. Project the position and image into the feature space of the same dimension. The details are as follows:

[0145] S2.1.1, based on the position coordinates p of the individual pig at a certain moment obtained in step S1 m =(x m ,y m ), m=1,2,…,M, which is input into the learnable position encoder to generate the position features corresponding to M pigs

[0146] The position encoder is composed of a function Y based on a random Fourier feature map and a multilayer perceptron with a ReLU activation function. Function Y is expressed as:

[0147]

[0148] Where, Represents learnable parameters, basis vector w i It is randomly sampled from a Gaussian distribution with 0 as mean and σ as standard deviation, that is, The Fourier eigenmap uses Bochner's theorem to efficiently construct a representation of an approximate translation-invariant kernel, enhancing the model's ability to handle translation invariance and its robustness to image position perception.

[0149] S2.2.2. At the same time, the video frame image of the corresponding moment is selected from the video, divided into N fixed-size image blocks, and a vector representation of dimension N×D is generated after flattening and mapping operations. It is added to the corresponding N×D-dimensional position embedding vector and input into the Transformer encoder to obtain the image feature vector The specific process is as follows:

[0150] S2.2.2.1 Video frame image Divided into N fixed-size (P 2 ×C) patches, denoted as N=HW / P 2, C represents the number of channels. Each patch is flattened into a one-dimensional vector and a set of patch vector representations with dimension D is generated through linear mapping. Vectorize patches to their positions with the same dimensions Add up to get input tokens, denoted as

[0151] S2.2.2.2, then, Input into the Transformer encoder with a total of L layers to obtain a new set of patch codes Recorded as image feature vector Each Transformer layer mainly consists of a multi-headed self-attention (MSA) block and a multilayer perceptron (MLP) block, each of which contains a layer normalization layer and a residual connection:

[0152]

[0153] S2.2, Model the complementary relationship between IMU signals and video images. and location characteristics They are input together into the detection frame decoder based on the cross attention mechanism, and the individual detection frame features corresponding to the M-dimensional position feature vector are output. Achieve deep integration of information between the two. The specific process is as follows:

[0154] P′=LN(Attn(P,P,P)+P) (19)

[0155] P″=LN(Attn(P,Z,Z)+P′) (20)

[0156] P″′=LN(MLP(P″)+P″) (21)

[0157] Z′=LN(Attn(Z,P″′,P″′)+Z) (22)

[0158] B= LN(Attn(P″′,Z′,Z″)+P″′) (23)

[0159] In the equation, LN(·) denotes layer normalization, and MLP(·) denotes multilayer perceptron.

[0160] The operation module in the middle of the decoder includes two position-image attention modules, one image-position attention module and a multi-layer perceptron. The position-image attention module replaces the multi-head self-attention block with a cross attention block based on the Transformer layer. Realization, Q = w Q P, K = w K Z, V = w V Z. Similarly, in the image-position attention module, Q = w′ Q Z, K = w′ K P, V = w′ V ·P, w Q , w K , w V , w′ Q , w′ K , and w′ V Represents a learnable mapping matrix. This decoder, built on an attention mechanism, can flexibly adapt to variable-length inputs of position features. By leveraging prior information about individual pig positions, it implicitly introduces the constraint of the actual number of pigs, helping to reduce missed and false detections. It is also applicable to farming scenarios with varying pig populations, demonstrating significant practical value.

[0161] S2.3, realize the tracking of individual pigs in group breeding. Input to the multi-layer perceptron, and the corresponding output is the key point coordinate pair of each pig's individual bounding box, recorded as and where represents the upper left and lower right corner coordinates of the bounding box of the mth pig, respectively. Thus, we can simultaneously detect and identify individual pigs in a single frame. Based on the individual pig detection results in these single frames, combined with the pre-established association of individual identities between video frames using IMU signals, robust group pig tracking is achieved.

[0162] like Figure 5 As shown, this embodiment discloses a multimodal group pig individual tracking system based on spatial prior guidance, based on the above method embodiment, including the following modules:

[0163] Position acquisition module: A momentum-corrected particle filter algorithm is established based on the high-frequency IMU signal to convert the high-frequency IMU signal into spatial position coordinates, so as to map the IMU signal and the video image to the same space;

[0164] Tracking module: Establish a group pig individual tracking algorithm guided by spatial position prior, extract the features of the obtained spatial position and the features of the corresponding video frame image, construct a detection box decoder based on the cross-attention mechanism, and establish a complementary relationship between the two modal data to achieve continuous and accurate tracking of group pig individuals.

[0165] For other contents of this embodiment, please refer to the above method embodiment.

[0166] The above is a preferred embodiment of the present invention. For ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be considered as the scope of protection of the present invention.

Claims

1. A multimodal group pig individual tracking method based on spatial prior guidance is characterized by: Follow these steps: S1, establishes a momentum-corrected particle filter algorithm based on the high-frequency IMU signal, converts the high-frequency IMU signal into spatial position coordinates, and maps the IMU signal and the video image to the same space; S2: Establish a group pig individual tracking algorithm guided by spatial position priors. Extract the features of the spatial position obtained in step S1 and the features of the corresponding video frame image, construct a detection box decoder based on the cross-attention mechanism, and establish a complementary relationship between the two modal data to achieve continuous and accurate tracking of group pig individuals. Step S2 is specifically as follows: S2.1, project the position and image into the feature space of the same dimension; S2.

2. Establish the complementary relationship between IMU signal and video image; and location characteristics They are input together into the detection frame decoder based on the cross attention mechanism, and the individual detection frame features corresponding to the M-dimensional position feature vector are output. To achieve the integration of information between the two; the specific process is as follows: P′=LN(Attn(P,P,P)+P) (19) P″=LN(Attn(P,Z,Z)+P′) (20) P″′=LN(MLP(P″)+P″) (21) Z′=LN(Attn(Z,P″′,P″′)+Z) (22) B=LN(Attn(P″′,Z′,Z″)+P″′) (23) Where, LN(·) denotes layer normalization, and MLP(·) denotes multilayer perceptron; The operation module in the middle of the decoder contains two position-image attention modules, one image-position attention module and a multi-layer perceptron; among them, the position-image attention module replaces the multi-head self-attention block with a cross attention block based on the Transformer layer. Realization, Q = w Q P, K = w K Z, V = w V Z; in the image-position attention module, Q = w′ Q Z, K = w′ K P, V = w′ V ·P, w Q 、w K 、w V , w′ Q , w′ K and w′ V Represents the mapping matrix that can be learned; S2.

3. Obtained individual detection frame features Input to the multi-layer perceptron, and the corresponding output is the key point coordinate pair of each pig's individual bounding box, recorded as and They represent the upper left corner coordinates and lower right corner coordinates of the m-th pig's bounding box respectively.

2. The multimodal group pig individual tracking method based on spatial prior guidance as claimed in claim 1 is characterized in that: Step S1 is specifically as follows: S1.

1. Set the starting position coordinates of the pig target and initialize the particle swarm probability. Perform equal-interval integration operations on the motion signal collected by the IMU sensor to obtain the target's speed and direction information. S1.

2. The uncorrected speed and direction values at the current moment are uniformly recorded as Ω t , the updated value after the momentum equation is corrected is recorded as The correction value at the previous moment is expressed as Initialization setting is 0; the correction relationship is expressed as follows: Among them, α represents the degree of correction; S1.

3. Perform zero-speed detection on the collected motion signal data to obtain a zero-speed interval; perform threshold determination using the acceleration modulus, angular velocity modulus, and angular velocity modulus standard deviation; S1.

4. Take the velocity value solved in the zero-speed interval as the observation quantity, establish the observation equation and discretize it to obtain: Z vk =H(X k )+V vk (5) With vk =[In e ,In n ,In u ] (6) V vk =[ΔV e ,ΔV n ,ΔV u ] (7) Where V e 、V n 、V u are the velocities in the east, north and sky directions respectively; H is the system observation matrix; ΔV e , ΔV n , ΔV u are the velocity observation noises in the east, north, and sky directions, respectively; S1.

5. Use particle filter algorithm to calculate the individual position coordinates of pigs.

3. The multimodal group pig individual tracking method based on spatial prior guidance as claimed in claim 2 is characterized in that the steps In S1.3: Acceleration modulus conditions: Where, |a i | is the axis acceleration modulus; a xi 、a yi 、a zi are the acceleration measurements of the x, y, and z axes respectively; a min 、a max are the upper and lower limits of the acceleration range respectively; Angular velocity modulus conditions: Where, |ω i | is the angular velocity modulus; ω xi 、ω yi 、ω zi are the angular velocity measurements of the x, y, and z axes respectively; ω max is the angular velocity interval threshold; Acceleration variance condition: Where, For|a i | is the standard deviation of the acceleration modulus at the sampling midpoint, and 2n+1 is the length of the sampling interval; is the mean value of acceleration modulus; σ a is the standard deviation threshold of the acceleration modulus; When the C1, C2, and C3 conditions all meet the three conditions at the same time, that is, C1(i)C2(i)C3(i)=1, the sampling point i is determined to be a zero-speed point, and the continuous zero-speed points constitute a zero-speed interval.

4. The multimodal group pig individual tracking method based on spatial prior guidance as claimed in claim 3 is characterized in that the steps S1.5 is as follows: S1.5.

1. At the initial time t = 0, the initial particles x0 ~ p(x0) are sampled according to the sampling distribution p(x0) to obtain the initial particle set Take the number of particles N = 100, and the initial weight of each particle is 1 / N; S1.5.

2. Constructing the importance probability density function Calculate each particle according to the following formula Corresponding weight Where, is the posterior probability density function; is the likelihood function of the observed quantity; is the state transition probability; is the proposed distribution function, satisfying is the weight at the k-1 moment; is the weight of particle k at time; Normalize it and get: Where, is the weight of the i-th particle at the k-th moment; is the weight sum of all particles at time k; is the value after weight normalization; S1.5.

3. Based on the normalized particle set, the state value at time k is obtained from the state value at time k-1, that is, the optimal estimate. The state estimate is: Where, is the particle at time k; S1.5.

4. Discard particles with weights less than the set value from the particle set, re-extract particles with weights greater than the set value and copy them to establish a new particle set.

5. The multimodal group pig individual tracking method based on spatial prior guidance as claimed in claim 4 is characterized in that: In step S1.5.4, the adaptive threshold method is used to select new particles. The specific steps are as follows: 1) Select an adaptive threshold M; M satisfies equation (11); Where: The value range of M is [M a ,M b ]; M b is the upper limit; M a is the initial value of the threshold, t is the number of iterations; N e is the effective particle number; K e is the proportional coefficient; the threshold M is proportional to the number of iterations t and the number of effective particles N e Inversely proportional to N e Satisfy formula (12); Using the effective number of particles N e To measure the degree of degradation of particle weights, N is the number of particles, is the particle weight variance; Substituting (12) into (11) we get: When M reaches the upper limit M b When M=M a ; 2) Retain particle judgment; if Keep otherwise 3) Make a secondary judgment on particle degradation: When H1+H2=1, return to step S1.5.1 to resample the particles. The weight of the resampled particles is 1 / N, where N is the number of particles. The new particle set As the initial particle of the next moment; Otherwise, return to step S1.5.2 for the next iteration until the loop ends; the position of the corresponding particle at the end of the loop is the final target position.

6. The multimodal group pig individual tracking method based on spatial prior guidance as claimed in claim 1 is characterized in that: Step S2.1 is as follows: S2.1.1, based on the position coordinates p of the individual pig at a certain moment obtained in step S1 m =(x m ,y m ), m=1,2,…,M, and input it into the learnable position encoder to generate the position features corresponding to M pigs The position encoder is composed of a function Y based on a random Fourier feature map and a multilayer perceptron with a ReLU activation function; Function Y is expressed as: Where, Represents the parameters that can be learned, basis vector w i It is randomly sampled from a Gaussian distribution with 0 as mean and σ as standard deviation, that is, S2.2.

2. Select the video frame image of the corresponding moment from the video, divide it into N fixed-size image blocks, generate a vector representation of dimension N×D through flattening and mapping operations, add it to the corresponding N×D-dimensional position embedding vector and input it into the Transformer encoder to obtain the image feature vector 7. The multimodal group pig individual tracking method based on spatial prior guidance as claimed in claim 6 is characterized in that: The specific process of step S2.2.2 is as follows: S2.2.2.1 Video frame image Divided into N fixed-size (P 2 ×C) image block, denoted as N=HW / P 2 , C represents the number of channels; each image block is flattened into a one-dimensional vector, and a set of image block vector representations with dimension D is generated through linear mapping Vectorize the position of the image patch to the same dimension as that of the image patch Add up to get input tokens, denoted as S2.2.2.

2. Input into the Transformer encoder with a total of L layers to obtain a new set of image block codes Recorded as image feature vector Each Transformer layer mainly consists of a multi-head self-attention block and a multi-layer perceptron block, each of which contains a layer normalization layer and a residual connection: Where l=1,…,L.

8. A multimodal group pig individual tracking system based on spatial prior guidance, based on the method according to any one of claims 1 to 7, characterized in that: Includes the following modules: Position acquisition module: A momentum-corrected particle filter algorithm is established based on the high-frequency IMU signal to convert the high-frequency IMU signal into spatial position coordinates, so as to map the IMU signal and the video image to the same space; Tracking module: Establish a group pig individual tracking algorithm guided by spatial position prior, extract the features of the obtained spatial position and the features of the corresponding video frame image, construct a detection box decoder based on the cross-attention mechanism, and establish a complementary relationship between the two modal data to achieve continuous and accurate tracking of group pig individuals.

Citation Information

Patent Citations

  • A three-dimensional wire frame structure method and system fusing a binocular camera and IMU positioning

    CN109166149A

  • Cross-video target tracking method and system, electronic equipment and storage medium

    CN114842028A