A Spatial Pattern Animal Pose Tracking Method and System Based on Proposal Relay

Through the spatial mode animal pose tracking method based on proposal broadcasting, the multi-animal pose tracking model is used to perform joint sequence prediction at the instance level and key point level, solving the accuracy and robustness of multi-objective tracking and pose estimation in complex environments, and improving the processing efficiency and accuracy of animal behavior analysis.

CN119992601BActive Publication Date: 2025-07-22TECH & ENG CENT FOR SPACE UTILIZATION CHINESE ACAD OF SCI
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510465503.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-22
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

The prior art multi-objective tracking and posture estimation methods in animal behavior analysis in complex environments have problems of low accuracy and poor robustness, especially in microgravity environments and high-density distribution scenarios, which are difficult to stably track animal targets.

Method used

The spatial mode animal pose tracking method based on proposal broadcast is adopted. Through the multi-animal pose tracking model, including the proposal generation module, feature extraction module, query decoding module and query interaction module, multi-objective tracking and pose estimation at the instance level and key point level are realized. The convolutional neural network and Transformer encoder are used for feature extraction and self-attention calculation, and the tracking accuracy and robustness of animal targets are improved.

Benefits of technology

In complex scenarios, the accuracy and robustness of multi-target tracking and pose estimation of animal targets is significantly improved, and it can effectively alleviate the tracking instability caused by occlusion and identity switching, promote the development of animal posture tracking technology in microgravity environments, and promote in-depth exploration of animal behavior research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992601B_ABST
    Figure CN119992601B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for tracking the postures of spatial model animals based on proposal relay, which relates to the technical field of animal behavior analysis. The method provided by the present invention obtains consecutive multi-frame images including multiple animal targets, and then determines the proposal queries, image features, and tracking queries for each frame of image; and determines the tracking results according to the proposal queries, image features of each frame of image, and the tracking queries of the previous frame of image. Since the proposal queries and tracking queries include feature information at multiple granularities of instance level and key point level, multi-target tracking and pose estimation of the animal targets included in the consecutive multi-frame images can be realized, effectively improving the accuracy and robustness, enhancing the processing efficiency, and meeting the usage requirements of users in different usage scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of animal behavior analysis, and particularly to a method and system for tracking animal postures in a spatial pattern based on proposal relay. Background Art

[0002] In the research of animal behavior analysis, it is usually necessary to perform multi-object tracking and pose estimation on animal targets in consecutive frame images. Among them, multi-object tracking refers to assigning a unique identifier to each animal target in consecutive frame images and maintaining its consistency. Pose estimation refers to simultaneously estimating and continuously tracking the positions of the skeleton key points of multiple animal targets in a video or image sequence, providing a basis for fine-grained behavior analysis.

[0003] In related technologies, the multi-object tracking method usually first uses an object detector to obtain candidate bounding boxes, and then links the candidate bounding boxes into complete tracking trajectories through an association strategy. However, in complex environments, such as scenarios with dense targets, severe occlusion, or frequent identity switching, the multi-object tracking methods in related technologies may have problems such as tracking interruption and target confusion, resulting in a decrease in the overall tracking robustness. Pose estimation methods usually rely on deep learning frameworks, but the deep learning frameworks in related technologies cannot meet the user's usage requirements in complex scenarios. For example, the mutation of behavior patterns caused by the microgravity environment and the high-density distribution scenario among individuals lead to significant differences in pose features compared with ground experiments; the high-density distribution among individuals increases the identity confusion and tracking difficulty; the severe occlusion problem makes it difficult to stably extract key information, resulting in the problems of low accuracy and poor robustness in the multi-object tracking and pose estimation methods in related technologies, and unable to meet the user's usage requirements.

[0004] Therefore, there is an urgent need for a method and system for tracking animal postures in a spatial pattern based on proposal relay, which can quickly complete multi-object tracking and pose estimation of animal targets in consecutive frame images in complex scenarios, improve accuracy and robustness, and improve processing efficiency. Summary of the Invention

[0005] Embodiments of the present invention provide a method and system for tracking animal postures in a spatial pattern based on proposal relay, which can achieve multi-object tracking and pose estimation of animal targets included in consecutive multi-frame images from multiple granularities at the instance level and the key point level in complex scenarios, effectively improve accuracy and robustness, improve processing efficiency, and meet the user's usage requirements in different usage scenarios.

[0006] To achieve the above object, the embodiments of the present invention adopt the following technical solutions:

[0007] In a first aspect, a method for tracking the postures of spatial model animals based on proposal relay is provided. The method includes: obtaining a plurality of consecutive frames of images, where each frame of the plurality of images includes a plurality of animal targets; based on a multi-animal posture tracking model, determining the tracking result of each frame of image according to the plurality of frames of images, the tracking result including the tracking information of each animal target, and the tracking information including a predicted bounding box, a plurality of key points, and a unique identifier; the multi-animal posture tracking model includes: a proposal generation module, a feature extraction module, a query decoding module, and a query interaction module; the proposal generation module is used to determine the proposal queries of each frame of image, and the proposal queries include a plurality of instance-level queries and key-point-level queries corresponding to each instance-level query. The instance-level queries include the candidate bounding boxes and confidence scores of the animal targets, and the key-point-level queries include a plurality of candidate key points of the animal targets; the feature extraction module is used to: determine the image features of each frame of image, and the image features include local features and global context features; the query decoding module is used to, when the image is the first frame of image, determine the hidden state according to the proposal query and image features of the first frame of image; and is further used to, when the image is the t-th frame of image, determine the hidden state according to the proposal query, image features of the t-th frame of image, and the tracking query of the (t - 1)-th frame of image, and is used to determine the tracking result of each frame of image according to the hidden state of each frame of image. The tracking query of the (t - 1)-th frame of image carries the tracking information of each animal target included in the (t - 1)-th frame of image, and t is a positive integer greater than 1; the query interaction module is used to determine the tracking query of the (t - 1)-th frame of image according to the hidden state corresponding to the (t - 1)-th frame of image.

[0008] In a possible implementation manner of the first aspect, the proposal generation module is specifically used to: determine a plurality of instance-level proposals corresponding to each frame of image; with the center of each candidate bounding box of each instance-level proposal as the center, determine the key-point-level proposals corresponding to each instance-level proposal; generate proposal queries according to the plurality of instance-level proposals corresponding to each frame of image and the key-point-level proposals corresponding to each instance-level proposal;

[0009] For the t-th frame of image, the data structure of the instance-level proposal is:

[0010] ;

[0011] ( ) is the center coordinate of the candidate bounding box of the i-th instance-level proposal; is the width of the candidate bounding box of the i-th instance-level proposal; is the height of the candidate bounding box of the i-th instance-level proposal; is the confidence score of the candidate bounding box of the i-th instance-level proposal; is the number of instance-level proposals;

[0012] The key-point-level proposal The data structure is as follows:

[0013] ;

[0014] = ;

[0015] is the set of key points corresponding to the i-th instance-level proposal, is the number of key points corresponding to the i-th instance-level proposal.

[0016] In a possible implementation of the first aspect, the feature extraction module includes a convolutional neural network and a Transformer encoder; the convolutional neural network is used to extract features from each frame of image to obtain a feature map corresponding to each frame of image, and the feature map carries local features; the Transformer encoder is used to determine the global context feature corresponding to each frame of image through a self-attention mechanism based on the feature map corresponding to each frame of image, and obtain the image feature corresponding to each frame of image.

[0017] In a possible implementation of the first aspect, the query decoding module includes an intra-instance self-attention layer, an inter-instance self-attention layer, and a hierarchical cross-attention layer; the intra-instance self-attention layer is used to perform self-attention calculation based on the key-point level queries included in the proposal query to obtain the correlation relationship between the key points corresponding to multiple animal targets, and obtain the updated key-point level queries; the inter-instance self-attention layer is used to perform self-attention calculation based on the instance level queries included in the proposal query to obtain the global relationship between multiple animal targets, and obtain the updated instance level queries; the hierarchical cross-attention layer is used to perform instance-level cross-attention calculation and key-point level cross-attention calculation on the updated key-point level queries, the updated instance level queries, the image features, and the tracking queries to obtain the hidden state.

[0018] The self-attention of the intra-instance self-attention layer The calculation formula of is:

[0019] ;

[0020] The self-attention of the inter-instance self-attention layer

[0021] ;

[0022] The cross-attention of the hierarchical cross-attention layer

[0023] ;

[0024] They are the query matrix, key matrix, and value matrix for key-point level queries respectively; d is the feature dimension; is the dot product of the key-point level query and the key; They are the query matrix, key matrix, and value matrix for instance level queries respectively; is the dot product of the instance level query and the key; They respectively represent the key matrix and value matrix of the image features, is the number of spatial positions of the image features, is the query matrix for the tracking query and the updated instance level query or key-point level query;

[0025] In a possible implementation manner of the first aspect, the query interaction module is specifically configured to: in the case where the image is the first frame image, generate an instance level trajectory query vector and a key-point level trajectory query vector for each animal target included in the first frame image according to the hidden state of the first frame image; in the case where the image is the t-th frame image, based on the time-correlated attention mechanism, determine the instance level trajectory query vector for each animal target included in the t-th frame image according to the instance level trajectory query vector of each animal target included in the (t - 1)-th frame image and the hidden state of the t-th frame image; based on the time-correlated attention mechanism, determine the key-point level trajectory query vector for each animal target included in the t-th frame image according to the key-point level trajectory query vector of each animal target included in the (t - 1)-th frame image and the hidden state of the t-th frame image; generate a tracking query corresponding to the t-th frame image according to the instance level trajectory query vector and the key-point level trajectory query vector of each animal target included in the t-th frame image.

[0026] In a possible implementation manner of the first aspect, the instance level trajectory query vector of each animal target included in the t-th frame image has the following calculation formula:

[0027] ;

[0028] where is the multi-head self-attention operation, is the instance level trajectory query vector of each animal target included in the (t - 1)-th frame image, is the instance level hidden state of the t-th frame image;

[0029] The key-point level trajectory query vector of each animal target included in the t-th frame image has the following calculation formula:

[0030] ;

[0031] is the key-point level trajectory query vector of each animal target included in the (t - 1)-th frame image, is the hidden state at the key point level of the t-th frame image.

[0032] In a possible implementation of the first aspect, before determining the instance-level trajectory query vector of each animal target included in the t-th frame image based on the time-correlated attention mechanism according to the instance-level trajectory query vector of each animal target included in the (t - 1)-th frame image and the hidden state of the t-th frame image, the above method further includes: constructing a trajectory set according to the hidden state of the first frame image, where the trajectory set includes the predicted bounding box, multiple key points, confidence score, and unique identifier of each animal target; updating the trajectory set according to the hidden state of the t-th frame image to obtain the trajectory set corresponding to the t-th frame image; determining the invalid animal targets included in the trajectory set corresponding to the t-th frame image, where the invalid animal targets do not exist in the multiple animal targets included in the trajectory set corresponding to the (t - 1)-th frame image and the confidence score is less than or equal to a preset threshold; deleting the invalid animal targets from the trajectory set corresponding to the t-th frame image to obtain the filtered trajectory set corresponding to the t-th frame image; generating the hidden state of the t-th frame image according to the filtered trajectory set corresponding to the t-th frame image.

[0033] In a possible implementation of the first aspect, before determining the tracking result of each frame image according to multiple frame images based on the multi-animal pose tracking model, the above method further includes: obtaining a plurality of training samples, where each training sample includes a plurality of consecutive multi-frame images and the tracking result corresponding to each frame image in the plurality of consecutive multi-frame images; constructing an objective loss function; iteratively training the multi-animal pose tracking model based on the objective loss function according to the plurality of training samples to obtain a trained multi-animal pose tracking model.

[0034] Objective loss function is:

[0035] ;

[0036] ;

[0037] ;

[0038] ;

[0039] where represents the total number of targets in the i-th frame, is the number of tracked animal targets; is the number of newly born animal targets; is the loss weight coefficient; is the focal loss of class prediction; is the mean absolute error loss of the bounding box overlap degree; is the mean absolute error loss of the target box overlap; k is the number of key points of the animal target, and respectively represent the coordinates of the th predicted key point and the true key point.

[0040] The beneficial effects of the present invention are as follows: The method provided by the present invention determines the proposal query, image features, and tracking query for each frame of image; and determines the tracking result according to the proposal query, image features of each frame of image, and the tracking query of the previous frame of image. Moreover, since the proposal query and the tracking query include feature information of multiple granularities at the instance level and the key point level, it can realize multi-object tracking and pose estimation of animal targets included in multiple consecutive frames of images, effectively improve the accuracy and robustness, improve the processing efficiency, and meet the user's usage requirements in different usage scenarios.

[0041] It can also be understood as: The multi-animal pose tracking model provided by the present invention can solve the limitations of multi-animal pose tracking in complex scenarios by jointly learning the appearance and position changes of animal targets at the instance level and the key point level; and, the method provided by the present invention models the pose tracking problem as a joint sequence prediction task at the instance level and the key point level, and uses instance-level and key point-level queries for dynamic modeling, significantly improving the processing ability of animal target individuals and pose changes in high-density scenarios; and the method provided by the present invention configures the decoder of self-attention and cross-attention within each animal target and between multiple animal targets through the query decoding module, significantly improving the performance of key point detection and tracking, and effectively alleviating the tracking instability problem caused by target occlusion and identity switching; the method provided by the present invention also realizes the efficient fusion of the overall (instance level) and local (key point level) features of animal targets through the query interaction module based on cross-frame modeling, thereby improving the tracking effect in scenarios with flexible postures and complex movements; finally, the method provided by the present invention can promote the development of animal pose tracking technology in the microgravity environment, thus promoting the in-depth exploration of animal behavior research.

[0042] Second aspect, the present invention provides a spatial pattern animal pose tracking system based on proposal relay. The system includes: an acquisition unit for acquiring a continuous multi-frame image, each frame of the multi-frame image including a plurality of animal targets; a tracking unit for determining the tracking result of each frame of image based on a multi-animal pose tracking model according to the multi-frame image. The tracking result includes the tracking information of each animal target, and the tracking information includes a predicted bounding box, a plurality of key points, and a unique identifier; the multi-animal pose tracking model includes: a proposal generation module, a feature extraction module, a query decoding module, and a query interaction module; the proposal generation module is used to determine the proposal query of each frame of image, and the proposal query includes a plurality of instance-level queries and key-point-level queries corresponding to each instance-level query. The instance-level query includes a candidate bounding box and a confidence score of the animal target, and the key-point-level query includes a plurality of candidate key points of the animal target; the feature extraction module is used to: determine the image feature of each frame of image, and the image feature includes local features and global context features; the query decoding module is used to, when the image is the first frame of image, determine the hidden state according to the proposal query and the image feature of the first frame of image; and is also used to, when the image is the t-th frame of image, determine the hidden state according to the proposal query, the image feature of the t-th frame of image, and the tracking query of the (t-1)-th frame of image, and is used to determine the tracking result of each frame of image according to the hidden state of each frame of image. The tracking query of the (t-1)-th frame of image carries the tracking information of each animal target included in the (t-1)-th frame of image, and t is a positive integer greater than 1; the query interaction module is used to determine the tracking query of the (t-1)-th frame of image according to the hidden state corresponding to the (t-1)-th frame of image.

[0043] Third aspect, there is provided an electronic device, which includes a memory and one or more processors; the memory is coupled to the processor; wherein, computer program code is stored in the memory, and the computer program code includes computer instructions. When the computer instructions are executed by the processor, the electronic device is caused to execute the method in any implementation manner of the first aspect.

[0044] Fourth aspect, there is provided a computer-readable storage medium, including computer instructions. When the computer instructions are run on an electronic device, the electronic device is caused to execute the method in any implementation manner of the first aspect.

[0045] Fifth aspect, there is provided a computer program product. When the computer program product is run on a computer, the computer is caused to execute the method in any implementation manner of the first aspect.

[0046] It can be understood that the beneficial effects that can be achieved by the system in the second aspect, the electronic device in the third aspect, the computer-readable storage medium in the fourth aspect, and the computer program product in the fifth aspect provided above can refer to the beneficial effects in the first aspect and any possible design manner thereof, and will not be elaborated here. Description of the Drawings

[0047] Figure 1 Schematic diagram of a structure of an electronic device provided by an embodiment of the present invention;

[0048] Figure 2 Flowchart of a method for tracking animal postures in a spatial pattern based on proposal relay provided by an embodiment of the present invention;

[0049] Figure 3 Schematic diagram of a structure of a multi-animal posture tracking model provided by an embodiment of the present invention;

[0050] Figure 4 Schematic diagram of a structure of a feature extraction module provided by an embodiment of the present invention;

[0051] Figure 5 Schematic diagram of a structure of a query decoding module provided by an embodiment of the present invention;

[0052] Figure 6 Schematic diagram of a structure of a query interaction module provided by an embodiment of the present invention;

[0053] Figure 7 Flowchart of a method for tracking animal postures in a spatial pattern based on proposal relay provided by an embodiment of the present invention;

[0054] Figure 8 Schematic diagram of a structure of a posture tracking system provided by an embodiment of the present invention. Detailed implementation manners

[0055] Next, the technical solutions in the embodiments of the present invention will be described in conjunction with the accompanying drawings in the embodiments of the present invention. Among them, in the description of the present invention, unless otherwise specified, " / " indicates that the objects associated before and after are in an "or" relationship. For example, A / B may represent A or B; the "or" in the present invention is merely a description of the association relationship of the associated objects, indicating that there can be three relationships. For example, A or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. Among them, A and B can be singular or plural. And, in the description of the present invention, unless otherwise specified, "a plurality of" means two or more than two. "At least one (item)" or its similar expression refers to any combination of these items, including any combination of single item (item) or plural items (items).

[0056] In addition, in order to facilitate a clear description of the technical solutions in the embodiments of the present invention, in the embodiments of the present invention, words such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and effects. Those skilled in the art can understand that the words "first" and "second" do not limit the quantity and execution order, and the words "first" and "second" do not necessarily limit to be different.

[0057] Meanwhile, in the embodiments of the present invention, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present invention should not be construed as being superior or more advantageous than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner for easy understanding.

[0058] In animal behavior analysis research, it is usually necessary to perform multi-object tracking and pose estimation on animal targets in consecutive frame images. Among them, multi-object tracking refers to assigning a unique identifier to each animal target in consecutive frame images and maintaining its consistency. Pose estimation refers to simultaneously estimating and continuously tracking the positions of the skeleton key points of multiple animal targets in a video or image sequence, providing a basis for fine-grained behavior analysis.

[0059] In the related art, the multi-object tracking method usually first uses an object detector to obtain candidate bounding boxes, and then links the candidate bounding boxes into complete tracking trajectories through an association strategy. However, in complex environments, such as scenarios with dense targets, severe occlusion, or frequent identity switching, the multi-object tracking method in the related art may have problems such as tracking interruption and target confusion, resulting in a decrease in the overall tracking robustness. The pose estimation method is usually based on a deep learning framework, but the deep learning framework in the related art cannot meet the user's usage requirements in complex scenarios. For example, the mutation of the behavior pattern caused by the microgravity environment and the scenario of high-density distribution among individuals lead to significant differences in pose features from ground experiments; the high-density distribution among individuals increases the identity confusion and tracking difficulty; the severe occlusion problem makes it difficult to stably extract key information, resulting in problems of low accuracy and poor robustness in the multi-object tracking and pose estimation methods in the related art, and unable to meet the user's usage requirements.

[0060] In view of this, an embodiment of the present invention provides a spatial pattern animal pose tracking method based on proposal relay. The above method includes: obtaining a plurality of consecutive frames of images, each frame of the plurality of frames of images including a plurality of animal targets; based on a multi-animal pose tracking model, determining the tracking result of each frame of image according to the plurality of frames of images, the tracking result including the tracking information of each animal target, the tracking information including a predicted bounding box, a plurality of key points, and a unique identifier; the multi-animal pose tracking model includes: a proposal generation module, a feature extraction module, a query decoding module, and a query interaction module; the proposal generation module is used to determine the proposal query of each frame of image, the proposal query including a plurality of instance-level queries and a key-point-level query corresponding to each instance-level query, the instance-level query including a candidate bounding box and a confidence score of the animal target, and the key-point-level query including a plurality of candidate key points of the animal target; the feature extraction module is used to: determine the image feature of each frame of image, the image feature including a local feature and a global context feature; the query decoding module is used to, when the image is the first frame of image, determine the hidden state according to the proposal query and the image feature of the first frame of image; and is further used to, when the image is the t-th frame of image, determine the hidden state according to the proposal query, the image feature of the t-th frame of image, and the tracking query of the (t-1)-th frame of image, and is used to determine the tracking result of each frame of image according to the hidden state of each frame of image, the tracking query of the (t-1)-th frame of image carrying the tracking information of each animal target included in the (t-1)-th frame of image, t being a positive integer greater than 1; the query interaction module is used to determine the tracking query of the (t-1)-th frame of image according to the hidden state corresponding to the (t-1)-th frame of image.

[0061] The method provided by the present invention determines the proposal query, the image feature, and the tracking query of each frame of image; and determines the tracking result according to the proposal query, the image feature, and the tracking query of the previous frame of image of each frame of image. And because the proposal query and the tracking query include feature information of multiple granularities at the instance level and the key-point level, it can realize multi-target tracking and pose estimation of animal targets included in a plurality of consecutive frames of images, effectively improve the accuracy and robustness, improve the processing efficiency, and meet the user's usage requirements in different usage scenarios.

[0062] It can also be understood that the multi - animal pose tracking model provided by the present invention can solve the limitations of multi - animal pose tracking in complex scenarios by jointly learning the appearance and position changes of animal targets at the instance - level and key - point level; moreover, the method provided by the present invention models the pose tracking problem as a joint sequence prediction task at the instance - level and key - point level, and dynamically models using instance - level and key - point level queries, significantly improving the processing ability of animal target individuals and pose changes in high - density scenarios; and the method provided by the present invention configures the decoders of self - attention and cross - attention within each animal target and between multiple animal targets through a query decoding module, significantly improving the performance of key - point detection and tracking, and effectively alleviating the tracking instability problem caused by target occlusion and identity switching; the method provided by the present invention also realizes the efficient fusion of the overall (instance - level) and local (key - point level) features of animal targets through a query interaction module based on cross - frame modeling, thereby improving the tracking effect in scenarios with flexible poses and complex movements; finally, the method provided by the present invention can promote the development of animal pose tracking technology in a microgravity environment, thus facilitating the in - depth exploration of animal behavior research.

[0063] In some embodiments, a spatial - pattern animal pose tracking method based on proposal propagation provided by the embodiments of the present invention can be executed by a spatial - pattern animal pose tracking system 100 based on proposal propagation (hereinafter referred to as the pose tracking system 100). As an example, the pose tracking system 100 can be any electronic device 200 with data - processing capabilities, such as a general - purpose computer, a personal computer, a laptop computer, a switch, or a tablet computer, etc. The specific implementation manner of the pose tracking system 100 is not limited herein.

[0064] Figure 1 The hardware structure schematic diagram of the electronic device provided by the embodiments of the present invention is shown. The electronic device 200 includes a processor 210, a memory 220, and a communication interface 230.

[0065] The processor 210 may include one or more processing cores. The processor 210 is connected to various parts within the electronic device 200 through various interfaces and circuits. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 220, and by invoking data stored in the memory 220, it performs various functions of the electronic device 200 and processes data. Optionally, the processor 210 may be implemented in at least one hardware form such as a Central Processing Unit (CPU), a graphics processing unit (GPU), a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA).

[0066] The memory 220 may include a random access memory (RAM), and may also include a read-only memory (ROM). Optionally, the memory 220 includes a non-transitory computer-readable medium. The memory 220 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 220 may include a storage program area. Among them, the storage program area may store instructions for implementing an operating system, instructions for implementing at least one function (such as an image acquisition function, an object tracking function, etc.), and instructions for implementing the above various method embodiments.

[0067] The communication interface 230 is used to communicate with other devices, equipment, or communication networks, such as data storage devices, image processing devices, or Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc.

[0068] In terms of physical implementation, the above-mentioned various devices (such as the processor 210, the memory 220, and the communication interface 230) may respectively be devices in the same device (such as a laptop computer). Or, at least two of them may be arranged in the same device, that is, as different devices in a device, similar to the deployment method of devices or components in a distributed system.

[0069] It can be understood that the structure illustrated in this embodiment does not constitute a specific limitation on the electronic device 200. In other embodiments of the present invention, the electronic device 200 may include more or fewer components than those shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0070] The following will describe a spatial pattern animal pose tracking method provided by an embodiment of the present invention with reference to the accompanying drawings of the specification.

[0071] Figure 2 It is a flowchart of a spatial pattern animal pose tracking method provided by an embodiment of the present invention. Optionally, this method may be executed by Figure 1 the illustrated electronic device 200. This method may include the following steps:

[0072] S1. Obtain a series of consecutive frames of images, where each frame of the series of consecutive frames of images includes multiple animal targets.

[0073] Specifically, the series of consecutive frames of images is continuous video data composed of multiple frames of images. The animal targets may be nematodes, fruit flies, zebrafish, or small animals such as mice and macaques, may be large wild animals such as horses and tigers, or may also be human targets. The specific types and quantities of the animal targets are not particularly limited in the embodiments of the present invention.

[0074] In one example, the series of consecutive frames of images includes 5 frames of images, and each frame of images includes 3 nematodes.

[0075] S2. Based on a multi-animal pose tracking model, determine the tracking result of each frame of image according to the series of consecutive frames of images. The tracking result includes the tracking information of each animal target, and the tracking information includes a predicted bounding box, multiple key points, and a unique identifier.

[0076] Specifically, according to the tracking result of each frame of image, the movement trajectories of the key points of each animal target can be analyzed, and key information such as movement speed, direction, acceleration, and swing frequency can be statistically analyzed. In addition, based on the tracking result of each frame of image, the behavior categories of the model animals can be further analyzed by unsupervised methods such as clustering and behavior recognition methods based on key points, and a behavior spectrum study can be carried out to provide necessary and important observation information for research in brain science, animal ethology, life science, etc.

[0077] Combined with the above example, when each frame of image includes 3 nematodes, the tracking result of each frame of image includes the predicted bounding box and multiple key points of nematode 1, the predicted bounding box and multiple key points of nematode 2, and the predicted bounding box and multiple key points of nematode 3. Among them, nematode 1, nematode 2, and nematode 3 are the unique identifiers of the 3 nematodes respectively, and can also be called the target indexes of the 3 nematodes.

[0078] Specifically, refer to Figure 3 , Figure 3 which is a schematic structural diagram of a multi-animal pose tracking model provided by an embodiment of the present invention. The multi-animal pose tracking model 300 includes: a proposal generation module 310, a feature extraction module 320, a query decoding module 330, and a query interaction module 340. The proposal generation module 310 is used to determine a proposal query for each frame of image. The proposal query includes a plurality of instance-level queries and key-point level queries corresponding to each instance-level query. The instance-level query includes a candidate bounding box and a confidence score of an animal target, and the key-point level query includes a plurality of candidate key points of the animal target. The feature extraction module 320 is used to: determine the image features of each frame of image, and the image features include local features and global context features. The query decoding module 330 is used to, when the image is the first frame of image, determine the hidden state according to the proposal query and the image features of the first frame of image; and is further used to, when the image is the t-th frame of image, determine the hidden state according to the proposal query, the image features of the t-th frame of image, and the tracking query of the (t - 1)-th frame of image, and is used to determine the tracking result of each frame of image according to the hidden state of each frame of image. The tracking query of the (t - 1)-th frame of image carries the tracking information of each animal target included in the (t - 1)-th frame of image, and t is a positive integer greater than 1. The query interaction module 340 is used to determine the tracking query of the (t - 1)-th frame of image according to the hidden state corresponding to the (t - 1)-th frame of image.

[0079] In one example, the proposal generation module 310 may be a Proposal Generation (PG), the feature extraction module 320 may be a Backbone&Encoder, the query decoding module 330 may be a Joint Query Decoder (JQD), and the query interaction module 340 is a Multi-Granularity Query Interaction Module (MGQIM). It should be understood that the above proposal generation module 310, feature extraction module 320, query decoding module 330, and query interaction module 340 may also be implemented in other ways. The above is only an exemplary illustration, and the embodiments of the present invention do not particularly limit the specific implementation manners of the proposal generation module 310, feature extraction module 320, query decoding module 330, and query interaction module 340.

[0080] In a possible implementation, the proposal generation module 310 is specifically configured to: determine multiple instance-level proposals corresponding to each frame of image; with each candidate bounding box of each instance-level proposal as the center, determine the key-point-level proposals corresponding to each instance-level proposal; generate a proposal query based on the multiple instance-level proposals corresponding to each frame of image and the key-point-level proposals corresponding to each instance-level proposal.

[0081] For the t-th frame of image, the data structure of the instance-level proposal is:

[0082] ;

[0083] ( ) is the center coordinate of the candidate bounding box of the i-th instance-level proposal; is the width of the candidate bounding box of the i-th instance-level proposal; is the height of the candidate bounding box of the i-th instance-level proposal; is the confidence score of the candidate bounding box of the i-th instance-level proposal; is the number of instance-level proposals;

[0084] The data structure of the key-point-level proposal is:

[0085] ;

[0086] = ;

[0087] is the set of key points corresponding to the i-th instance-level proposal, is the number of key points corresponding to the i-th instance-level proposal. is the coordinate of the k-th key point corresponding to the i-th instance-level proposal.

[0088] Specifically, the proposal generation module 310 generates proposals at two granularities of instance level and key-point level (instance-level proposals and key-point-level proposals) during the training and inference phases of the multi-animal pose tracking model 300. For each frame of image included in a series of consecutive frames of images, the proposal generation module 310 generates multiple instance-level proposals, including the center coordinate, width, height, and confidence score of the animal target (equivalent to the candidate bounding box). At the same time, the key-point-level proposals are initialized with the center points of the candidate bounding boxes of each instance-level proposal.

[0089] In one example, the proposal generation module 310 determines multiple instance-level proposals and key-point-level proposals corresponding to each frame of image according to the vision-language large model Grounding DINO, where Grounding DINO is based on DINO (Detection Transformer with Improved deNoising anchOr boxes) and semantic localization (grounding) technology. It can achieve accurate localization of specific semantic animal targets in each frame of image in complex scenes, and obtain multiple instance-level proposals and key-point-level proposals corresponding to each frame of image. Thus, it effectively solves the problem of the lack of detection models for special targets in space science experiments. Based on the vision-language large model (VLM), this method can achieve zero-shot proposal generation, making it more general and adaptable.

[0090] The following provides an exemplary description of the generation process of the instance-level query included in the proposal query provided by the embodiment of the present invention. First, the proposal generation module 310 constructs a shared query vector . Then, the shared query vector is broadcast to a preset shape to match the number of instance-level proposals. Subsequently, the confidence score is encoded using sine-cosine position encoding to generate a score embedding , and it is added to the shared query vector to form the final instance-level query.

[0091] The instance-level query is determined by the formula:

[0092] + ;

[0093] Furthermore, before the proposal generation module 310 generates the instance-level query included in the proposal query, the method provided by the embodiment of the present invention further includes:

[0094] The proposal generation module 310 generates M learnable anchors, and splices the M learnable anchors with the instance-level proposals to optimize the instance-level proposals.

[0095] The method provided by the embodiment of the present invention can optimize the instance-level proposals through the splicing of M learnable anchors and the instance-level proposals, enhance the recall ability of animal targets, that is, it can ensure the robustness of the multi-animal pose tracking model 300 in the case of missed detection by the proposal generation module 310.

[0096] In some embodiments, refer to Figure 4 , Figure 4A structural schematic diagram of a feature extraction module provided by an embodiment of the present invention. The feature extraction module 320 includes a convolutional neural network 3201 and a Transformer encoder 3202. The convolutional neural network 3201 is used to extract features from each frame of image to obtain a feature map corresponding to each frame of image, and the local features are carried in the feature map. The Transformer encoder 3202 is used to determine the global context features corresponding to each frame of image through a self-attention mechanism according to the feature map corresponding to each frame of image, and obtain the image features corresponding to each frame of image.

[0097] In some embodiments, referring to Figure 5 , the query decoding module 330 includes an intra-instance self-attention layer 3301, an inter-instance self-attention layer 3302, and a hierarchical cross-attention layer 3303. The intra-instance self-attention layer 3301 is used to perform self-attention calculation according to the key-point level queries included in the proposal query to obtain the association relationship between the key points corresponding to multiple animal targets, and obtain the updated key-point level queries. The inter-instance self-attention layer 3302 is used to perform self-attention calculation according to the instance level queries included in the proposal query to obtain the global relationship between multiple animal targets, and obtain the updated instance level queries. The hierarchical cross-attention layer 3303 is used to perform instance-level cross-attention calculation and key-point level cross-attention calculation on the updated key-point level queries, the updated instance level queries, the image features, and the tracking queries to obtain the hidden state.

[0098] The self-attention of the intra-instance self-attention layer 3301 The calculation formula of is:

[0099] ;

[0100] The self-attention of the inter-instance self-attention layer 3302

[0101] ;

[0102] The cross-attention of the hierarchical cross-attention layer 3303

[0103] ;

[0104] are the query matrix, key matrix, and value matrix of the key-point level queries respectively; d is the feature dimension; is the dot product of the key-point level query and the key; are the query matrix, key matrix, and value matrix of the instance level queries respectively; is the dot product of the instance level query and the key; respectively represent the key matrix and value matrix of the image features, is the number of spatial positions of the image features, is the query matrix for tracking queries and the updated instance-level queries or keypoint-level queries.

[0105] Specifically, the within-instance self-attention layer 3301 acts on each animal target based on the within-instance self-attention mechanism, enabling attention calculation for each animal target and the keypoints of each animal target, and capturing the correlation between each animal target and the keypoints. This attention mechanism is based on the query-key-value (QKV) calculation method of the Transformer. The across-instance self-attention layer 3302 acts between the instance-level queries of different animal targets based on the across-instance self-attention mechanism to model the global relationship between multiple animal targets, thereby enhancing the robustness of the multi-animal pose tracking model 300 in crowded or occluded scenarios.

[0106] Furthermore, the hierarchical cross-attention layer 3303 is used to connect high-level queries (candidate queries and tracking queries) with low-level visual features (image features), enabling direct interaction between candidate queries, tracking queries, and image features to achieve information fusion. The hierarchical cross-attention layer 3303 includes an instance-level cross-attention (ILCA) mechanism and a keypoints-level cross-attention (KLCA) mechanism. The instance-level cross-attention mechanism enables interaction between instance-level queries and global image features to optimize the class prediction and bounding box regression of animal targets. The keypoints-level cross-attention enables interaction between keypoints-level queries and local image features to accurately locate the keypoint positions.

[0107] In a possible implementation, the query interaction module 340 is specifically configured to: in the case where the image is the first-frame image, generate an instance-level trajectory query vector and a keypoints-level trajectory query vector for each animal target included in the first-frame image according to the hidden state of the first-frame image.

[0108] When the image is the t-th frame image, based on the temporal correlation attention mechanism, the instance-level trajectory query vector of each animal target included in the t-th frame image is determined according to the instance-level trajectory query vector of each animal target included in the (t - 1)-th frame image and the hidden state of the t-th frame image; based on the temporal correlation attention mechanism, the key-point-level trajectory query vector of each animal target included in the t-th frame image is determined according to the key-point-level trajectory query vector of each animal target included in the (t - 1)-th frame image and the hidden state of the t-th frame image; the tracking query corresponding to the t-th frame image is generated according to the instance-level trajectory query vector and the key-point-level trajectory query vector of each animal target included in the t-th frame image.

[0109] Further, the instance-level trajectory query vector of each animal target included in the t-th frame image is calculated as:

[0110] ;

[0111] where is the multi-head self-attention operation, is the instance-level trajectory query vector of each animal target included in the (t - 1)-th frame image, is the instance-level hidden state of the t-th frame image;

[0112] The key-point-level trajectory query vector of each animal target included in the t-th frame image is calculated as:

[0113] ;

[0114] is the key-point-level trajectory query vector of each animal target included in the (t - 1)-th frame image, is the key-point-level hidden state of the t-th frame image.

[0115] In an example, see Figure 6, the query interaction module 340 includes an Instance-wise TAN (Instance-wise Time Aggregation Network) 3401 and a Keypoints-wise TAN (Keypoints-wise Time Aggregation Network) 3402. Among them, the Instance-wise TAN 3401 and the Keypoints-wise TAN 3402 are connected in parallel. The Instance-wise TAN 3401 determines the instance-level trajectory query vector of each animal target included in the t-th frame image based on the time-correlated attention mechanism, according to the instance-level trajectory query vector of each animal target included in the (t - 1)-th frame image and the hidden state of the t-th frame image. The Keypoints-wise TAN 3402 determines the keypoints-level trajectory query vector of each animal target included in the t-th frame image based on the time-correlated attention mechanism, according to the keypoints-level trajectory query vector of each animal target included in the (t - 1)-th frame image and the hidden state of the t-th frame image.

[0116] For the specific implementation manners of the Instance-wise TAN 3401 and the Keypoints-wise TAN 3402, refer to the above calculation processes of the instance-level trajectory query vector and the keypoints-level trajectory query vector, which will not be elaborated here.

[0117] In a possible implementation manner, before determining the instance-level trajectory query vector of each animal target included in the t-th frame image based on the time-correlated attention mechanism, according to the instance-level trajectory query vector of each animal target included in the (t - 1)-th frame image and the hidden state of the t-th frame image, the method provided by the embodiments of the present invention further includes:

[0118] Construct a trajectory set according to the hidden state of the first frame image. The trajectory set includes the predicted bounding box, multiple keypoints, confidence score, and unique identifier of each animal target; update the trajectory set according to the hidden state of the t-th frame image to obtain the trajectory set corresponding to the t-th frame image; determine the invalid animal targets included in the trajectory set corresponding to the t-th frame image. The invalid animal targets do not exist among the multiple animal targets included in the trajectory set corresponding to the (t - 1)-th frame image and the confidence score is less than or equal to a preset threshold; delete the invalid animal targets from the trajectory set corresponding to the t-th frame image to obtain the filtered trajectory set corresponding to the t-th frame image; generate the hidden state of the t-th frame image according to the filtered trajectory set corresponding to the t-th frame image.

[0119] Specifically, the input of the query interaction module 340 is the hidden state output by the query decoding module 330 and the instance-level prediction score of the hidden state. The query interaction module 340 filters active trajectories through the confidence score and identity consistency.

[0120] In an example, in combination with Figure 6When the input is the hidden state of the first frame, the trajectory screening layer 3403 is configured to construct a trajectory set according to the hidden state of the first-frame image. When the input is the hidden state of the t-th frame (other frames except the first frame), the trajectory screening layer 3403 updates the trajectory set according to the hidden state of the t-th frame image to obtain the trajectory set corresponding to the t-th frame image; determine the invalid animal targets included in the trajectory set corresponding to the t-th frame image, where the invalid animal targets do not exist in the multiple animal targets included in the trajectory set corresponding to the (t - 1)-th frame image and the confidence score is less than or equal to a preset threshold; delete the invalid animal targets from the trajectory set corresponding to the t-th frame image to obtain the filtered trajectory set corresponding to the t-th frame image; generate the hidden state of the t-th frame image according to the filtered trajectory set corresponding to the t-th frame image.

[0121] It can also be understood that: the trajectory screening layer 3403 first constructs a trajectory set according to the hidden state of the first-frame image. For the t-th frame image, the trajectory screening layer 3403 updates the trajectory set of the previous frame image (the (t - 1)-th frame) according to the hidden state of the t-th frame image to delete the invalid animal targets in the trajectory set of the previous frame image, obtaining the completed screened trajectory set, and then inputs the hidden state corresponding to the completed screened trajectory set (equivalent to the optimized hidden state of the t-th frame image) into the instance-level temporal aggregation network 3401 and the key-point level temporal aggregation network 3402. Then, the instance-level temporal aggregation network 3401 and the key-point level temporal aggregation network 3402 respectively output the tracking query of the t-th frame image according to the corresponding instance-level part, key-point level part in the hidden state and the tracking query of the previous frame image.

[0122] In one example, the specific implementation manner for the trajectory screening layer 3403 to screen active trajectories through confidence scores and identity consistency is as follows:

[0123] First, the trajectory screening layer 3403 constructs a trajectory set according to the hidden state , where each trajectory has an object index of the animal target (equivalent to a unique identifier), a confidence score and the IoU value with the animal target of the current frame image . The query interaction module 340 adopts different strategies for trajectory screening in the training and inference stages of the multi-animal pose tracking model 300.

[0124] In the training stage, the trajectory screening layer 3403 only retains the trajectories with valid object indices or confidence scores higher than the preset threshold , that is:

[0125]

[0126] Among them, is a preset threshold. In one example, the preset threshold is 0.5. In addition, if the IoU of the trajectory and the current detection is lower than the set threshold , the object index of the animal target is reset to invalid (equivalent to determining an invalid animal target):

[0127]

[0128] In the inference stage, the trajectory screening layer 3403 only retains the trajectories with valid object indices, that is:

[0129]

[0130] Through the above screening strategy, the trajectory screening layer 3403 can effectively remove the trajectories with low confidence or un-matched for a long time, complete the screening and optimization of the hidden state of each frame of image, so that the instance-level time aggregation network 3401 and the key-point-level time aggregation network 3402 can obtain the tracking query of the current frame according to the optimized hidden state and the tracking query of the previous frame of image, which can effectively improve the accuracy of the model, thereby improving the robustness of the tracking of animal targets. Especially in complex scenarios, it can effectively reduce the risk of drift and incorrect association.

[0131] In some embodiments, referring to Figure 7 , before determining the tracking result of each frame of image based on the multi-animal pose tracking model, the method provided by the embodiments of the present invention further includes:

[0132] S71. Obtain a plurality of training samples, each training sample including a plurality of consecutive multi-frame images and the tracking result corresponding to each frame of image in the plurality of consecutive multi-frame images.

[0133] Specifically, the training samples are crucial for the temporal modeling of trajectories. Similar to the baseline method, this method learns the temporal change characteristics of the target pose in a data-driven manner, rather than relying on heuristics designed manually such as Kalman filtering. The common training strategy based on adjacent two frames is difficult to effectively cover the target motion patterns with long time spans. This study follows the baseline method, takes consecutive multi-frame images as the input of training samples, and at the same time introduces key-point-level training samples, combining the multi-granularity learning strategies at the instance level and the key-point level to enhance the model's ability to model long-term motion. Through this improvement, the model can more accurately capture the local structural changes of the target, thereby improving the stability of the key-point trajectories and the robustness of the overall tracking.

[0134] S72. Construct a target loss function.

[0135] S73. Iteratively train the multi-animal pose tracking model based on a target loss function using multiple training samples to obtain a trained multi-animal pose tracking model.

[0136] Target loss function is:

[0137] ;

[0138] ;

[0139] ;

[0140] ;

[0141] where represents the total number of targets in the i-th frame, is the number of tracked animal targets; is the number of newborn animal targets; is the loss weight coefficient; is the focal loss for class prediction; is the mean absolute error loss of the bounding box overlap; is the mean absolute error loss of the target box overlap; k is the number of key points of the animal target, and respectively represent the coordinates of the th predicted key point and the true key point.

[0142] The method provided in the embodiments of the present invention can effectively improve the model's ability to model long-term motion by introducing the Pose-Aware Collective Average Loss (PCAL). Especially in the accurate tracking of target key points. Due to considering the local motion details of animal targets, PCAL not only enhances the temporal learning ability of the overall target motion, but also improves the capture effect of long-time-span target motion patterns under the multi-granularity learning strategy, thus significantly improving the robustness and accuracy of the target tracking of the multi-animal pose tracking model 300.

[0143] The beneficial effects of a spatial pattern animal pose tracking method based on proposal relay provided in the embodiments of the present invention are exemplarily described below with an example.

[0144] Exemplarily, a validation dataset is obtained. The validation dataset includes three different animal targets: Caenorhabditis elegans, Drosophila, and Zebrafish. Each type of animal target has different numbers and types of key points to accurately characterize its movement and behavior patterns. The validation dataset is divided into three subsets. The Caenorhabditis elegans subset contains 127 videos, 21,502 frames, and 47,320 annotation instances, with 5 key points for each instance; the Zebrafish subset contains 10 videos, a total of 1,711 frames, and 6,844 annotation instances, with 10 key points for each instance; the Drosophila subset contains 10 videos, a total of 1,082 frames, and 11,972 annotation instances, with 26 key points for each Drosophila. This dataset is obtained in complex scenarios, such as changes in movement patterns caused by the influence of microgravity environment, occlusion and overlap caused by high individual density, key point drift and blur affecting detection accuracy, and the requirement of cross-species generalization ability, which requires the model to have strong adaptability.

[0145] In this example, the method provided by the embodiments of the present invention is verified based on the evaluation metrics MOTAkeypoints, MOTAtotal, and recall rate. MOTAkeypoints and MOTAtotal are evaluation metrics in multi-object tracking (MOT) and key point detection tasks, and are usually used to measure the overall performance of the model in the multi-object key point tracking task. They are extended based on the MOTA (Multi-Object Tracking Accuracy) metric.

[0146] Among them, MOTAkeypoints calculates the MOTA for specific key points (Keypoints) in the multi-object tracking task to evaluate whether the model can accurately and stably track the key points of each target.

[0147] The calculation formula is as follows:

[0148] ;

[0149] Among them, GT keypoints is the total number of all ground truth trajectories, FN keypoints is the number of missed key point detections, FP keypoints is the number of false key point detections, and IDSW keypoints is the number of incorrect key point identity switches. The value range of MOTA is , and the closer it is to 1, the better the tracking performance.

[0150] MOTA total calculates the average value of MOTA for each target level and key point level.

[0151] First, the proposal generator 310 uses the vision - language model Grounding - DINO in MMDetection to generate multiple instance - level proposals and corresponding keypoint - level proposals for each instance - level proposal. The pre - trained weights and hyperparameters are used, and the prompt words are set to "worm", "fish", and "fruit fly", corresponding to C. elegans, zebrafish, and Drosophila melanogaster respectively. To improve the recall rate of the proposals, we retain all Grounding - DINO predicted bounding boxes with a confidence higher than 0.2 as instance - level proposals, and initialize the keypoints based on the center points inferred from the predicted bounding boxes to obtain keypoint - level proposals. Among them, C. elegans, zebrafish, and Drosophila melanogaster contain 5, 10, and 26 keypoints respectively.

[0152] The implementation of the multi - animal pose tracking model 300 is based on PyTorch. The baseline method selects MOTRv2, and a convolutional neural network with ResNet50 as the feature extraction module is used. The batch size is set to 1, and the optimizer is AdamW.

[0153] In terms of training configuration, the training epoch of the C. elegans dataset is 50 rounds, the initial learning rate is set to 2e - 4, and it is decayed by 10 times at the 40th round. The training epochs of the zebrafish and Drosophila melanogaster datasets are 200 rounds, and the initial learning rate is also set to 2e - 4, but it is decayed by 10 times at the 180th round.

[0154] In terms of the video clip (Clip) setting, the Clip size of the C. elegans dataset is set to 5, and the frame sampling step within each clip is randomly selected between 1 and 5. In contrast, the Clip size of the zebrafish and Drosophila melanogaster datasets is also set to 5, but the frame sampling step is fixed to 1.

[0155] All models are initialized from the COCO pre - trained weights of Deformable DETR. Since we have improved the keypoint decoder on the basis of the baseline method, the model only loads the parts with the same network structure, and the rest are trained from scratch. Therefore, a longer training epoch is required to ensure convergence and performance optimization.

[0156] Refer to Table 1. Table 1 is a comparison data table of evaluation metrics for the nematode dataset of the method provided in the embodiments of the present invention and the methods in the related art. Among them, each nematode includes 5 key points, namely MOTAhead, MOTAfront, MOTAmiddle, MOTAback, and MOTAtail. Among them, we evaluated the detection-based tracking methods Bytetrack and OC-SORT to analyze the combined performance of different pose estimation methods and tracking algorithms, and generated detection boxes according to the characteristics of different methods. For the top-down method, Related Art 1 is the method combining HRNet and OC-SORT, and Related Art 2 is the method combining VITPose and OC-SORT. For the bottom-up method, Related Art 3 is the method combining AE and OC-SORT. Related Art 4 is the method combining DEKR and OC-SORT. Related Art 5 is the method combining CID and OC-SORT. Related Art 6 is the method combining HRNet and Bytetrack, and Related Art 7 is the method combining VITPose and Bytetrack. For the bottom-up method, Related Art 8 is the method combining AE and Bytetrack. Related Art 9 is the method combining DEKR and Bytetrack. Related Art 10 is the method combining CID and Bytetrack.

[0157] Table 1

[0158]

[0159] As can be seen from Table 1, the method provided by the present invention has achieved the best performance in both MOTA and MOTA_total of the key points, with a 2.75% improvement compared to VItPose+OC-SORT. The results show that the method provided by the present invention can achieve good performance in the space nematode with intense movement and crowded scenes, meeting the requirements of subsequent research and analysis.

[0160] The above mainly introduced the solution of the embodiments of the present invention from the perspective of the method. It can be understood that in order to implement the above functions, the pose tracking system 100 includes at least one of the corresponding hardware structures and software modules for executing each function. Those skilled in the art should easily realize that, combining the units and algorithm steps of each example described in the embodiments disclosed in this article, the embodiments of the present invention can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present invention.

[0161] In the embodiments of the present invention, the pose tracking system 100 can be divided into functional units according to the above method examples. For example, each functional unit corresponding to each function of the pose tracking system 100 can be divided, or two or more functions can be integrated into one processing unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. It should be noted that the division of units in the embodiments of the present invention is illustrative, and is only a logical function division. There may be other division methods in actual implementation.

[0162] Exemplarily, Figure 8 FIG. shows a schematic hardware structure diagram of a pose tracking system provided by an embodiment of the present invention. The pose tracking system 100 includes: an acquisition unit 110, configured to acquire a plurality of consecutive frames of images, each frame of the plurality of frames of images including a plurality of animal targets; a tracking unit 120, configured to determine a tracking result of each frame of image based on a multi-animal pose tracking model according to the plurality of frames of images, the tracking result including tracking information of each animal target, the tracking information including a predicted bounding box, a plurality of key points, and a unique identifier; the multi-animal pose tracking model includes: a proposal generation module, a feature extraction module, a query decoding module, and a query interaction module; the proposal generation module is configured to determine a proposal query for each frame of image, the proposal query including a plurality of instance-level queries and a key-point level query corresponding to each instance-level query, the instance-level query including a candidate bounding box and a confidence score of an animal target, and the key-point level query including a plurality of candidate key points of the animal target; the feature extraction module is configured to: determine an image feature of each frame of image, the image feature including a local feature and a global context feature; the query decoding module is configured to, when the image is the first frame of image, determine a hidden state according to the proposal query and the image feature of the first frame of image; and is further configured to, when the image is the t-th frame of image, determine a hidden state according to the proposal query, the image feature of the t-th frame of image, and the tracking query of the (t-1)-th frame of image, and is configured to determine a tracking result of each frame of image according to the hidden state of each frame of image, the tracking query of the (t-1)-th frame of image carrying the tracking information of each animal target included in the (t-1)-th frame of image, where t is a positive integer greater than 1; the query interaction module is configured to determine the tracking query of the (t-1)-th frame of image according to the hidden state corresponding to the (t-1)-th frame of image.

[0163] It should be understood that the specific descriptions of the above optional manners can be referred to the foregoing method embodiments, and will not be repeated here. In addition, the explanations and descriptions of the beneficial effects of any of the above-provided pose tracking systems 100 can refer to the corresponding method embodiments above, and will not be repeated.

[0164] An embodiment of the present invention also provides a computer-readable storage medium, in which at least one computer instruction is stored, and the at least one computer instruction is loaded and executed by a processor to implement the methods of the above various embodiments. For the explanations and beneficial effects descriptions of the relevant content in any of the above provided computer-readable storage media, reference can be made to the corresponding embodiments above, and details are not repeated here.

[0165] An embodiment of the present invention also provides a chip. The chip integrates a control circuit and one or more ports for implementing the functions of the above-mentioned attitude tracking system 100. Optionally, the functions supported by the chip can be referred to above, and details are not repeated here.

[0166] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a random access memory, etc. The above-mentioned processing unit or processor can be a central processing unit, a general-purpose processor, an application specific integrated circuit (ASIC), a digital signal processor (DSP), a field programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof.

[0167] Embodiments of the present invention also provide a computer program product containing instructions. When the instructions run on a computer, the computer is caused to execute any one of the methods in the above embodiments. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present invention are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from a website, a computer, a server, or a data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more available media integrated. The available medium may be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as an SSD), etc.

[0168] It should be noted that the devices for storing computer instructions or computer programs provided in the embodiments of the present invention, such as but not limited to, the above-mentioned memory, computer-readable storage medium, and communication chip, etc., are all non-transitory. Those skilled in the art should be able to realize that in the above one or more examples, the functions described in the embodiments of the present invention can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions may be stored in a computer-readable storage medium or transmitted as one or more instructions or codes on a computer-readable storage medium. The computer-readable storage medium includes a computer storage medium and a communication medium, where the communication medium includes any medium that facilitates the transmission of a computer program from one place to another. The storage medium may be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0169] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. A spatial model animal posture tracking method based on proposal relay, characterized in that The method includes: Obtaining a plurality of consecutive frames of images, each frame of the plurality of frames of images including a plurality of animal targets; Based on a multi-animal pose tracking model, determining the tracking result of each frame of image according to the plurality of frames of images, the tracking result including the tracking information of each animal target, the tracking information including a predicted bounding box, a plurality of key points, and a unique identifier; The multi-animal pose tracking model includes: a proposal generation module, a feature extraction module, a query decoding module, and a query interaction module; the proposal generation module is used to determine the proposal query of each frame of image, the proposal query including a plurality of instance-level queries and a key-point-level query corresponding to each instance-level query, the instance-level query including a candidate bounding box and a confidence score of an animal target, and the key-point-level query including a plurality of candidate key points of the animal target; the feature extraction module is used to: determine the image feature of each frame of image, the image feature including a local feature and a global context feature; the query decoding module is used to, when the image is the first frame of image, determine the hidden state according to the proposal query and the image feature of the first frame of image; and is also used to, when the image is the t-th frame of image, determine the hidden state according to the proposal query, the image feature of the t-th frame of image, and the tracking query of the (t-1)-th frame of image, and is used to determine the tracking result of each frame of image according to the hidden state of each frame of image, the tracking query of the (t-1)-th frame of image carrying the tracking information of each animal target included in the (t-1)-th frame of image, where t is a positive integer greater than 1; the query interaction module is used to determine the tracking query of the (t-1)-th frame of image according to the hidden state corresponding to the (t-1)-th frame of image; The query interaction module is specifically used for: When the image is the first frame of image, generating an instance-level trajectory query vector and a key-point-level trajectory query vector for each animal target included in the first frame of image according to the hidden state of the first frame of image; When the image is the t-th frame of image, determining the instance-level trajectory query vector for each animal target included in the t-th frame of image based on a time-correlated attention mechanism according to the instance-level trajectory query vector of each animal target included in the (t-1)-th frame of image and the hidden state of the t-th frame of image; Based on a time-correlated attention mechanism, determining the key-point-level trajectory query vector for each animal target included in the t-th frame of image according to the key-point-level trajectory query vector of each animal target included in the (t-1)-th frame of image and the hidden state of the t-th frame of image; Generating a tracking query corresponding to the t-th frame of image according to the instance-level trajectory query vector and the key-point-level trajectory query vector of each animal target included in the t-th frame of image; The instance-level trajectory query vector of each animal target included in the t-th frame image The calculation formula is as follows: ; Among them, is the multi-head self-attention operation, is the instance-level trajectory query vector for each animal target included in the (t-1)-th frame image, is the instance-level hidden state of the t-th frame image; The key-point level trajectory query vector of each animal target included in the t-th frame image The calculation formula is as follows: ; is the key-point level trajectory query vector for each animal target included in the (t - 1)-th frame image, is the key-point level hidden state of the t-th frame image.

2. The method according to claim 1, characterized in that The proposal generation module is specifically used for: Determining a plurality of instance-level proposals corresponding to each frame of image; Taking each candidate bounding box of each instance-level proposal as the center, determining the key-point-level proposal corresponding to each instance-level proposal; Generating a proposal query according to the plurality of instance-level proposals corresponding to each frame of image and the key-point-level proposal corresponding to each instance-level proposal; For the t-th frame image, the instance-level proposal has the following data structure: ; ( is the center coordinate of the candidate bounding box for the i-th instance-level proposal; is the width of the candidate bounding box for the i-th instance-level proposal; is the height of the candidate bounding box for the i-th instance-level proposal; is the confidence score of the candidate bounding box for the i-th instance-level proposal; is the number of instance-level proposals; Key point level proposal The data structure is as follows: ; = ; is the set of key points corresponding to the i-th instance-level proposal, is the number of key points corresponding to the i-th instance-level proposal.

3. The method according to claim 2, characterized in that, The feature extraction module includes a convolutional neural network and a Transformer encoder; The convolutional neural network is used to extract features from each frame of image, obtaining a feature map corresponding to each frame of image, and the local features are carried in the feature map; The Transformer encoder is used to determine the global context features corresponding to each frame of image through the self-attention mechanism according to the feature map corresponding to each frame of image, obtaining the image features corresponding to each frame of image.

4. The method according to claim 3, characterized in that The query decoding module includes an intra-instance self-attention layer, an inter-instance self-attention layer, and a hierarchical cross-attention layer; The intra-instance self-attention layer is used to perform self-attention calculation according to the key-point level queries included in the proposal query, obtaining the correlation relationships between the key points corresponding to multiple animal targets, and obtaining the updated key-point level queries; The inter-instance self-attention layer is used to perform self-attention calculation according to the instance level queries included in the proposal query, obtaining the global relationships between multiple animal targets, and obtaining the updated instance level queries; The hierarchical cross-attention layer is used to perform instance-level cross-attention calculation and key-point level cross-attention calculation on the updated key-point level queries, the updated instance level queries, the image features, and the tracking queries, obtaining a hidden state; The self-attention of the internal self-attention layer of the instance has the following calculation formula: ; The self-attention of the cross-instance self-attention layer ; The cross-attention of the hierarchical cross-attention layer ; They are the query matrix, key matrix, and value matrix for key-point level queries respectively; d is the feature dimension; It is the dot product of the key-point level query and the key; They are the query matrix, key matrix, and value matrix for instance level queries respectively; It is the dot product of the instance level query and the key; They respectively represent the key matrix and value matrix of the image features, It is the number of spatial positions of the image features, It is the query matrix of the tracking query and the updated instance level query or key-point level query.

5. The method according to claim 4, wherein Before determining the instance-level trajectory query vectors of each animal target included in the t-th frame of image based on the time-correlated attention mechanism according to the instance-level trajectory query vectors of each animal target included in the (t - 1)-th frame of image and the hidden state of the t-th frame of image, the method further includes: Constructing a trajectory set according to the hidden state of the first frame of image, where the trajectory set includes the predicted bounding boxes, multiple key points, confidence scores, and unique identifiers of each animal target; Updating the trajectory set according to the hidden state of the t-th frame of image, obtaining the trajectory set corresponding to the t-th frame of image; Determining the invalid animal targets included in the trajectory set corresponding to the t-th frame of image, where the invalid animal targets do not exist in the multiple animal targets included in the trajectory set corresponding to the (t - 1)-th frame of image and the confidence scores are less than or equal to a preset threshold; Deleting the invalid animal targets from the trajectory set corresponding to the t-th frame of image, obtaining the filtered trajectory set corresponding to the t-th frame of image; Generating the hidden state of the t-th frame of image according to the filtered trajectory set corresponding to the t-th frame of image.

6. The method according to claim 5, wherein Before determining the tracking results of each frame of image based on the multi-animal pose tracking model according to the multiple frames of image, the method further includes: Obtaining a plurality of training samples, each training sample including a plurality of consecutive frames of images and the tracking results corresponding to each frame of image in the plurality of consecutive frames of images; Constructing an objective loss function; Based on the objective loss function, iteratively training the multi-animal pose tracking model according to the plurality of training samples, obtaining a trained multi-animal pose tracking model; The target loss function is as follows: ; ; ; ; Among them, represents the total number of targets in the i-th frame, is the number of tracked animal targets; is the number of newborn animal targets; is the loss weight coefficient; is the focal loss for class prediction; is the mean absolute error loss of the bounding box overlap; is the mean absolute error loss of the target box overlap; k is the number of key points of the animal target, and respectively represent the coordinates of the th predicted key point and the true key point.

7. A spatial model animal pose tracking system based on proposal relay, characterized in that, The system includes: An acquisition unit, configured to acquire a plurality of consecutive frames of images, where each frame of image in the plurality of frames of images includes a plurality of animal targets; A tracking unit, configured to determine the tracking results of each frame of image based on the multi-animal pose tracking model according to the plurality of frames of image, where the tracking results include the tracking information of each animal target, and the tracking information includes a predicted bounding box, multiple key points, and a unique identifier; The multi-animal pose tracking model includes: a proposal generation module, a feature extraction module, a query decoding module, and a query interaction module; the proposal generation module is used to determine the proposal queries for each frame of image, the proposal queries include a plurality of instance-level queries and key-point-level queries corresponding to each instance-level query, the instance-level queries include candidate bounding boxes and confidence scores of animal targets, and the key-point-level queries include a plurality of candidate key points of animal targets; the feature extraction module is used to: determine the image features of each frame of image, the image features include local features and global context features; the query decoding module is used to, when the image is the first frame of image, determine the hidden state according to the proposal queries and image features of the first frame of image; and is also used to, when the image is the t-th frame of image, determine the hidden state according to the proposal queries, image features of the t-th frame of image, and the tracking queries of the (t - 1)-th frame of image, and is used to determine the tracking results of each frame of image according to the hidden state of each frame of image, the tracking queries of the (t - 1)-th frame of image carry the tracking information of each animal target included in the (t - 1)-th frame of image, and t is a positive integer greater than 1; the query interaction module is used to determine the tracking queries of the (t - 1)-th frame of image according to the hidden state corresponding to the (t - 1)-th frame of image; The query interaction module is specifically used for: When the image is the first frame of image, generating an instance-level trajectory query vector and a key-point-level trajectory query vector for each animal target included in the first frame of image according to the hidden state of the first frame of image; When the image is the t-th frame of image, based on the temporal correlation attention mechanism, determining the instance-level trajectory query vector for each animal target included in the t-th frame of image according to the instance-level trajectory query vector of each animal target included in the (t - 1)-th frame of image and the hidden state of the t-th frame of image; Based on the temporal correlation attention mechanism, determining the key-point-level trajectory query vector for each animal target included in the t-th frame of image according to the key-point-level trajectory query vector of each animal target included in the (t - 1)-th frame of image and the hidden state of the t-th frame of image; Generating the tracking query corresponding to the t-th frame of image according to the instance-level trajectory query vector and the key-point-level trajectory query vector of each animal target included in the t-th frame of image; The instance-level trajectory query vector of each animal target included in the t-th frame image The calculation formula is as follows: ; Among them, is the multi-head self-attention operation, is the instance-level trajectory query vector for each animal target included in the (t-1)-th frame image, is the instance-level hidden state of the t-th frame image; The key-point level trajectory query vector of each animal target included in the t-th frame image The calculation formula is as follows: ; is the key-point level trajectory query vector for each animal target included in the (t - 1)-th frame image, is the key-point level hidden state of the t-th frame image.

8. An electronic device, characterized in that, Includes: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the instructions to implement the method for tracking the animal pose in the spatial pattern based on proposal relay according to any one of claims 1-6.

Citation Information

Patent Citations

  • Multi-target tracking method for endowing tracking proposal propagation by diffusion model

    CN117893570A

  • Key point extracting and tracking method for movement of space nematodes in space station

    CN118379327A

  • Deep learning method for multiple object tracking from video

    US20240144489A1