Spatial mode animal posture tracking method and system based on proposal rebroadcasting
By using spatial mode animal pose tracking methods and systems based on proposal broadcasting in complex scenarios, the accuracy and robustness of multi-objective tracking and pose estimation in the prior art are solved, and efficient animal target tracking and pose estimation are achieved.
Patent Information
- Application Number
- CN202510465503.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-15
AI Technical Summary
The prior art is difficult to achieve multi-objective tracking and pose estimation of animal targets in continuous frame images in complex scenarios, and there are problems such as tracking interruption and target confusion, resulting in low accuracy and poor robustness.
The spatial mode animal pose tracking method and system based on proposal broadcasting is adopted, and the multi-objective tracking model, including the proposal generation module, feature extraction module, query decoding module and query interaction module, is used to realize multi-objective tracking and pose estimation of animal targets.
In complex scenarios, it effectively improves the accuracy and robustness of multi-objective tracking and pose estimation, improves processing efficiency, and meets the needs of different usage scenarios.
Smart Images

Figure CN119992601A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of animal behavior analysis, and in particular to a spatial pattern animal posture tracking method and system based on proposal relay. Background Art
[0002] In animal behavior analysis research, it is usually necessary to perform multi-target tracking and pose estimation of animal targets in continuous frame images. Multi-target tracking refers to assigning a unique identifier to each animal target in continuous frame images and maintaining its consistency. Pose estimation refers to simultaneously estimating and continuously tracking the positions of skeleton key points of multiple animal targets in a video or image sequence, providing a basis for fine-grained behavior analysis.
[0003] In the related art, the multi-target tracking method usually first uses a target detector to obtain a candidate bounding box, and then links the candidate bounding boxes into a complete tracking trajectory through an association strategy. However, in complex environments, such as scenarios with dense targets, severe occlusion, or frequent identity switching, the multi-target tracking method in the related art may have problems such as tracking interruption and target confusion, resulting in a decrease in the overall tracking robustness. The posture estimation method is usually based on a deep learning framework, but the deep learning framework in the related art cannot meet the user's usage needs in complex scenarios. For example, the sudden change in behavior patterns caused by the microgravity environment and the high-density distribution of individuals lead to significant differences in posture characteristics from ground experiments; the high-density distribution of individuals increases identity confusion and tracking difficulty; the serious occlusion problem makes it difficult to stably extract key information, resulting in low accuracy and poor robustness in the multi-target tracking and posture estimation methods in the related art, which cannot meet the user's usage needs.
[0004] Therefore, there is an urgent need for a spatial pattern animal posture tracking method and system based on proposal relay, which can quickly complete multi-target tracking and posture estimation of animals in continuous frame images in complex scenes, improve accuracy and robustness, and improve processing efficiency. Summary of the invention
[0005] The embodiments of the present invention provide a spatial pattern animal posture tracking method and system based on proposal relay, which can realize multi-target tracking and posture estimation of animal targets included in continuous multi-frame images from multiple granularities at the instance level and key point level in complex scenes, effectively improve accuracy and robustness, improve processing efficiency, and meet user needs in different usage scenarios.
[0006] To achieve the above object, the embodiments of the present invention adopt the following technical solutions: In a first aspect, a spatial pattern animal posture tracking method based on proposal relay is provided, the method comprising: acquiring multiple consecutive frames of images, each of the multiple frames including multiple animal targets; based on a multi-animal posture tracking model, determining the tracking result of each frame according to the multiple frames, the tracking result including the tracking information of each animal target, the tracking information including a predicted bounding box, multiple key points and a unique identifier; the multi-animal posture tracking model comprises: a proposal generation module, a feature extraction module, a query decoding module and a query interaction module; the proposal generation module is used to determine a proposal query for each frame of the image, the proposal query includes multiple instance-level queries and a key point-level query corresponding to each instance-level query, the instance-level query includes a candidate bounding box and a confidence score of the animal target, and the key point-level query includes the animal target multiple candidate key points; the feature extraction module is used to: determine the image features of each frame of image, the image features include local features and global context features; the query decoding module is used to determine the hidden state according to the proposal query and image features of the first frame of image when the image is the first frame of image; it is also used to determine the hidden state according to the proposal query, image features and tracking query of the t-1 frame of image when the image is the t-th frame of image, and to determine the tracking result of each frame of image according to the hidden state of each frame of image, the tracking query of the t-1 frame of image carries the tracking information of each animal target included in the t-1 frame of image, and t is a positive integer greater than 1; the query interaction module is used to determine the tracking query of the t-1 frame of image according to the hidden state corresponding to the t-1 frame of image.
[0007] In a possible implementation of the first aspect, the proposal generation module is specifically used to: determine multiple instance-level proposals corresponding to each frame of the image; determine a key point-level proposal corresponding to each instance-level proposal with each candidate bounding box of each instance-level proposal as the center; generate a proposal query according to the multiple instance-level proposals corresponding to each frame of the image and the key point-level proposal corresponding to each instance-level proposal; For the t-th frame image, instance-level proposal The data structure is: ; ( ) is the center coordinate of the candidate bounding box of the i-th instance-level proposal; is the width of the candidate bounding box of the i-th instance-level proposal; is the height of the candidate bounding box of the i-th instance-level proposal; is the confidence score of the candidate bounding box of the i-th instance-level proposal; is the number of instance-level proposals; Keypoint-level proposals The data structure is: ; = ; is the set of key points corresponding to the i-th instance-level proposal, is the number of key points corresponding to the i-th instance-level proposal.
[0008] In a possible implementation of the first aspect, the feature extraction module includes a convolutional neural network and a Transformer encoder; the convolutional neural network is used to extract features from each frame of image to obtain a feature map corresponding to each frame of image, and the feature map carries local features; the Transformer encoder is used to determine the global context features corresponding to each frame of image through a self-attention mechanism based on the feature map corresponding to each frame of image, and obtain the image features corresponding to each frame of image.
[0009] In a possible implementation of the first aspect, the query decoding module includes an internal instance self-attention layer, a cross-instance self-attention layer and a hierarchical cross-attention layer; the internal instance self-attention layer is used to perform self-attention calculations based on the key point level queries included in the proposal queries, obtain the association relationship between the key points corresponding to multiple animal targets, and obtain an updated key point level query; the cross-instance self-attention layer is used to perform self-attention calculations based on the instance level queries included in the proposal queries, obtain the global relationship between multiple animal targets, and obtain an updated instance level query; the hierarchical cross-attention layer is used to perform instance-level cross-attention calculations and key point-level cross-attention calculations on the updated key point level queries, the updated instance level queries, the image features, and the tracking queries to obtain the hidden state.
[0010] Self-attention of self-attention layer within instance The calculation formula is: ; Self-Attention across Instance Self-Attention Layers ; Cross-Attention Layer Cross-Attention ; are the query matrix, key matrix and value matrix of key point level query respectively; d is the feature dimension; is the dot product of the keypoint-level query and the key; They are the query matrix, key matrix, and value matrix for instance-level queries respectively; is the dot product of the instance-level query and the key; Represent the key matrix and value matrix of image features respectively, is the number of spatial locations of image features, A query matrix for tracking query and updated instance-level query or keypoint-level query; In a possible implementation of the first aspect, the query interaction module is specifically used to: when the image is the first frame image, generate an instance-level trajectory query vector and a key point-level trajectory query vector for each animal target included in the first frame image according to the hidden state of the first frame image; when the image is the t-th frame image, based on the time-related attention mechanism, determine the instance-level trajectory query vector for each animal target included in the t-1 frame image according to the hidden state of the t-th frame image; based on the time-related attention mechanism, determine the key point-level trajectory query vector for each animal target included in the t-1 frame image according to the hidden state of the t-th frame image; generate a tracking query corresponding to the t-th frame image according to the instance-level trajectory query vector and the key point-level trajectory query vector for each animal target included in the t-th frame image.
[0011] In a possible implementation of the first aspect, the instance-level trajectory query vector of each animal target included in the t-th frame image is The calculation formula is: ; in, For multi-head self-attention operation, is the instance-level trajectory query vector for each animal target included in the t-1 frame image, is the instance-level hidden state of the t-th frame image; The key point level trajectory query vector of each animal target included in the t-th frame image The calculation formula is: ; is the key point level trajectory query vector for each animal target included in the t-1 frame image, is the hidden state of the key point level of the t-th frame image.
[0012] In a possible implementation of the first aspect, before determining the instance-level trajectory query vector of each animal target included in the t-1 frame image based on the time-correlated attention mechanism and according to the instance-level trajectory query vector of each animal target included in the t-1 frame image and the hidden state of the t frame image, the above method also includes: constructing a trajectory set according to the hidden state of the 1st frame image, the trajectory set including a predicted bounding box, multiple key points, a confidence score and a unique identifier for each animal target; updating the trajectory set according to the hidden state of the t-th frame image to obtain a trajectory set corresponding to the t-th frame image; determining invalid animal targets included in the trajectory set corresponding to the t-1 frame image, the invalid animal targets do not exist in the multiple animal targets included in the trajectory set corresponding to the t-1 frame image and the confidence score is less than or equal to a preset threshold; deleting the invalid animal targets from the trajectory set corresponding to the t-th frame image to obtain a filtered trajectory set corresponding to the t-th frame image; and generating the hidden state of the t-th frame image according to the filtered trajectory set corresponding to the t-th frame image.
[0013] In a possible implementation of the first aspect, before determining the tracking result of each frame image according to multiple frames of images based on the multi-animal posture tracking model, the method further includes: obtaining multiple training samples, each training sample including multiple consecutive multi-frame images and the tracking result corresponding to each frame image in the multiple consecutive multi-frame images; constructing a target loss function; based on the target loss function, iteratively training the multi-animal posture tracking model according to the multiple training samples to obtain a trained multi-animal posture tracking model; Objective loss function for: ; ; ; ; in, represents the number of all targets in the i-th frame, The target number of animals tracked; Target number for newborn animals; is the loss weight coefficient; Focal loss for class prediction; is the mean absolute error loss of bounding box overlap; is the mean absolute error loss of target frame overlap; k is the number of key points of the animal target, and Respectively represent The coordinates of the predicted key points and the true key points.
[0014] The beneficial effects of the present invention are as follows: the method provided by the present invention determines the proposal query, image features and tracking query of each frame image; and determines the tracking result according to the proposal query, image features and tracking query of the previous frame image of each frame image, and because the proposal query and the tracking query include feature information of multiple granularities at the instance level and the key point level, it is possible to realize multi-target tracking and posture estimation of animal targets included in multiple consecutive frames of images, effectively improve accuracy and robustness, improve processing efficiency, and meet the user's usage needs in different usage scenarios.
[0015] It can also be understood that: the multi-animal posture tracking model provided by the present invention can solve the limitations of multi-animal posture tracking in complex scenes in the related art by jointly learning the appearance and position changes of animal targets at the instance level and the key point level; and the method provided by the present invention models the posture tracking problem as a joint sequence prediction task at the instance level and the key point level, and uses instance level and key point level queries for dynamic modeling, which significantly improves the processing ability of individual animal targets and posture changes in high-density scenes; and the method provided by the present invention configures the decoders of self-attention and cross-attention within each animal target and between multiple animal targets through the query decoding module, which significantly improves the performance of key point detection and tracking, and effectively alleviates the tracking instability problem caused by target occlusion and identity switching; the method provided by the present invention also realizes the efficient fusion of the overall (instance level) and local (key point level) features of the animal target based on cross-frame modeling through the query interaction module, thereby improving the tracking effect in scenes with flexible and changeable postures and complex movements; finally, the method provided by the present invention can be applied to outer space animal experimental scenes, thereby promoting the development of animal posture tracking technology in microgravity environments, thereby promoting in-depth exploration of outer space animal behavior research.
[0016] In the second aspect, the present invention provides a spatial pattern animal posture tracking system based on proposal relay, the system comprising: an acquisition unit for acquiring continuous multiple frames of images, each frame of the multiple frames including multiple animal targets; a tracking unit for determining the tracking result of each frame of image based on the multiple frames of images based on a multi-animal posture tracking model, the tracking result including the tracking information of each animal target, the tracking information including a predicted bounding box, multiple key points and a unique identifier; the multi-animal posture tracking model comprises: a proposal generation module, a feature extraction module, a query decoding module and a query interaction module; the proposal generation module is used to determine the proposal query for each frame of image, the proposal query includes multiple instance-level queries and a key point-level query corresponding to each instance-level query, the instance-level query includes a candidate bounding box and a confidence score of the animal target, the key points The query level includes multiple candidate key points of the animal target; the feature extraction module is used to: determine the image features of each frame of the image, the image features include local features and global context features; the query decoding module is used to determine the hidden state according to the proposal query and image features of the first frame of the image when the image is the first frame of the image; it is also used to determine the hidden state according to the proposal query, image features and tracking query of the t-1 frame of the image when the image is the t-th frame of the image, and to determine the tracking result of each frame of the image according to the hidden state of each frame of the image, the tracking query of the t-1 frame of the image carries the tracking information of each animal target included in the t-1 frame of the image, and t is a positive integer greater than 1; the query interaction module is used to determine the tracking query of the t-1 frame of the image according to the hidden state corresponding to the t-1 frame of the image.
[0017] According to a third aspect, an electronic device is provided, comprising a memory and one or more processors; the memory is coupled to the processor; wherein the memory stores computer program code, the computer program code comprises computer instructions, and when the computer instructions are executed by the processor, the electronic device executes a method as in any implementation of the first aspect.
[0018] According to a fourth aspect, a computer-readable storage medium is provided, comprising computer instructions. When the computer instructions are executed on an electronic device, the electronic device executes the method in any implementation of the first aspect.
[0019] According to a fifth aspect, a computer program product is provided. When the computer program product is run on a computer, the computer is enabled to execute the method in any implementation of the first aspect.
[0020] It can be understood that the beneficial effects that can be achieved by the system of the second aspect, the electronic device of the third aspect, the computer-readable storage medium of the fourth aspect, and the computer program product of the fifth aspect provided above can be referred to the beneficial effects in the first aspect and any possible design method thereof, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention; Figure 2 A flowchart of a spatial pattern animal posture tracking method based on proposal relay provided by an embodiment of the present invention; Figure 3 A schematic diagram of the structure of a multi-animal posture tracking model provided by an embodiment of the present invention; Figure 4 A structural schematic diagram of a feature extraction module provided by an embodiment of the present invention; Figure 5 A schematic diagram of the structure of a query decoding module provided by an embodiment of the present invention; Figure 6 A schematic diagram of the structure of a query interaction module provided by an embodiment of the present invention; Figure 7 A flowchart of a spatial pattern animal posture tracking method based on proposal relay provided by an embodiment of the present invention; Figure 8 A schematic diagram of the structure of a posture tracking system provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0022] The technical solution in the embodiment of the present invention will be described below in conjunction with the accompanying drawings in the embodiment of the present invention. In the description of the present invention, unless otherwise specified, " / " indicates that the objects associated before and after are in an "or" relationship. For example, A / B can represent A or B; the "or" in the present invention is only a description of the association relationship of the associated objects, indicating that there can be three relationships. For example, A or B can represent: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. In addition, in the description of the present invention, unless otherwise specified, "multiple" means two or more than two. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items.
[0023] In addition, in order to clearly describe the technical solutions of the embodiments of the present invention, in the embodiments of the present invention, words such as "first" and "second" are used to distinguish the same or similar items with substantially the same functions and effects. Those skilled in the art can understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit the difference.
[0024] Meanwhile, in the embodiments of the present invention, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present invention should not be interpreted as being better or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a concrete way for easy understanding.
[0025] In animal behavior analysis research, it is usually necessary to perform multi-target tracking and pose estimation of animal targets in continuous frame images. Multi-target tracking refers to assigning a unique identifier to each animal target in continuous frame images and maintaining its consistency. Pose estimation refers to simultaneously estimating and continuously tracking the positions of skeleton key points of multiple animal targets in a video or image sequence, providing a basis for fine-grained behavior analysis.
[0026] In the related art, the multi-target tracking method usually first uses a target detector to obtain a candidate bounding box, and then links the candidate bounding boxes into a complete tracking trajectory through an association strategy. However, in complex environments, such as scenarios with dense targets, severe occlusion, or frequent identity switching, the multi-target tracking method in the related art may have problems such as tracking interruption and target confusion, resulting in a decrease in the overall tracking robustness. The posture estimation method is usually based on a deep learning framework, but the deep learning framework in the related art cannot meet the user's usage needs in complex scenarios. For example, the sudden change in behavior patterns caused by the microgravity environment and the high-density distribution of individuals lead to significant differences in posture characteristics from ground experiments; the high-density distribution of individuals increases identity confusion and tracking difficulty; the serious occlusion problem makes it difficult to stably extract key information, resulting in low accuracy and poor robustness in the multi-target tracking and posture estimation methods in the related art, which cannot meet the user's usage needs.
[0027] In view of this, an embodiment of the present invention provides a spatial pattern animal posture tracking method based on proposal relay, the method comprising: acquiring multiple consecutive frames of images, each frame of the multiple frames including multiple animal targets; based on a multi-animal posture tracking model, determining the tracking result of each frame of the image according to the multiple frames of the image, the tracking result including the tracking information of each animal target, the tracking information including a predicted bounding box, multiple key points and a unique identifier; the multi-animal posture tracking model comprises: a proposal generation module, a feature extraction module, a query decoding module and a query interaction module; the proposal generation module is used to determine a proposal query for each frame of the image, the proposal query includes multiple instance-level queries and a key point-level query corresponding to each instance-level query, the instance-level query includes a candidate bounding box and a confidence score of the animal target, and the key point-level query includes the animal The invention relates to a plurality of candidate key points of a target; the feature extraction module is used to determine the image features of each frame of the image, wherein the image features include local features and global context features; the query decoding module is used to determine the hidden state according to the proposal query and image features of the first frame of the image when the image is the first frame of the image; the module is also used to determine the hidden state according to the proposal query, image features and tracking query of the t-1 frame of the image when the image is the t-th frame of the image, and to determine the tracking result of each frame of the image according to the hidden state of each frame of the image, wherein the tracking query of the t-1 frame of the image carries the tracking information of each animal target included in the t-1 frame of the image, and t is a positive integer greater than 1; the query interaction module is used to determine the tracking query of the t-1 frame of the image according to the hidden state corresponding to the t-1 frame of the image.
[0028] The method provided by the present invention determines the proposal query, image features and tracking query of each frame image; and determines the tracking result according to the proposal query, image features and tracking query of the previous frame image. Moreover, since the proposal query and tracking query include feature information of multiple granularities at the instance level and the key point level, it is possible to realize multi-target tracking and posture estimation of animal targets included in multiple consecutive frames of images, effectively improve accuracy and robustness, improve processing efficiency, and meet the user's usage needs in different usage scenarios.
[0029] It can also be understood that: the multi-animal posture tracking model provided by the present invention can solve the limitations of multi-animal posture tracking in complex scenes in the related art by jointly learning the appearance and position changes of animal targets at the instance level and the key point level; and the method provided by the present invention models the posture tracking problem as a joint sequence prediction task at the instance level and the key point level, and uses instance level and key point level queries for dynamic modeling, which significantly improves the processing ability of individual animal targets and posture changes in high-density scenes; and the method provided by the present invention configures the decoders of self-attention and cross-attention within each animal target and between multiple animal targets through the query decoding module, which significantly improves the performance of key point detection and tracking, and effectively alleviates the tracking instability problem caused by target occlusion and identity switching; the method provided by the present invention also realizes the efficient fusion of the overall (instance level) and local (key point level) features of the animal target based on cross-frame modeling through the query interaction module, thereby improving the tracking effect in scenes with flexible and changeable postures and complex movements; finally, the method provided by the present invention can be applied to outer space animal experimental scenes, thereby promoting the development of animal posture tracking technology in microgravity environments, thereby promoting in-depth exploration of outer space animal behavior research.
[0030] In some embodiments, a spatial pattern animal posture tracking method based on proposal relay provided by an embodiment of the present invention can be performed by a spatial pattern animal posture tracking system 100 based on proposal relay (hereinafter referred to as posture tracking system 100). As an example, the posture tracking system 100 can be any electronic device 200 with data processing capabilities, such as a general-purpose computer, a personal computer, a laptop, a switch or a tablet computer, etc. The specific implementation method of the posture tracking system 100 is not limited here.
[0031] Figure 1 The hardware structure diagram of the electronic device provided by the embodiment of the present invention is shown. The electronic device 200 includes a processor 210, a memory 220 and a communication interface 230.
[0032] The processor 210 may include one or more processing cores. The processor 210 uses various interfaces and lines to connect various parts in the electronic device 200, and executes various functions and processes data of the electronic device 200 by running or executing instructions, programs, code sets or instruction sets stored in the memory 220, and calling data stored in the memory 220. Optionally, the processor 210 can be implemented in at least one hardware form of a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processing (DSP), a field-programmable gate array (FPGA), and a programmable logic array (PLA).
[0033] The memory 220 may include a random access memory (RAl) or a read-only memory (ROL). Optionally, the memory 220 includes a non-transitory computer-readable storage medium (non-transitory colputer-readable storage lediul). The memory 220 may be used to store instructions, programs, codes, code sets or instruction sets. The memory 220 may include a program storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as an image acquisition function, a target tracking function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.
[0034] The communication interface 230 is used to communicate with other devices, equipment or communication networks, such as data storage devices, image processing equipment or Ethernet, radio access network (RAN), wireless local area networks (WLAN), etc.
[0035] In physical implementation, the above-mentioned components (such as processor 210, memory 220 and communication interface 230) can be components in the same device (such as a laptop computer). Alternatively, at least two of the components can be set in the same device, that is, as different components in a device, such as a deployment method similar to devices or components in a distributed system.
[0036] It is to be understood that the structure illustrated in this embodiment does not constitute a specific limitation on the electronic device 200. In other embodiments of the present invention, the electronic device 200 may include more or fewer components than those illustrated, or combine certain components, or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0037] The following describes a spatial pattern animal posture tracking method based on proposal relay provided by an embodiment of the present invention in conjunction with the accompanying drawings of the specification.
[0038] Figure 2 A flowchart of a spatial pattern animal posture tracking method based on proposal relay provided by an embodiment of the present invention. Optionally, the method can be Figure 1 The electronic device 200 shown is executed. The method may include the following steps: S1. Acquire multiple continuous image frames, each of which includes multiple animal targets.
[0039] Specifically, the continuous multi-frame images are continuous video data composed of multiple frames of images. The animal target may be a nematode, a fruit fly, a zebrafish, or a small animal such as a mouse or a macaque, or a large wild animal such as a horse or a tiger, or a human target. The embodiment of the present invention does not impose any particular restrictions on the specific types and quantities of the animal targets.
[0040] In one example, the continuous multiple frames of images include 5 frames of images, and each frame of image includes 3 nematodes.
[0041] S2. Based on the multi-animal posture tracking model, the tracking result of each frame image is determined according to the multiple frames of images, and the tracking result includes the tracking information of each animal target, and the tracking information includes a predicted bounding box, multiple key points and a unique identifier.
[0042] Specifically, based on the tracking results of each frame of the image, the motion trajectory of the key points of each animal target can be analyzed, and key information such as motion speed, direction, acceleration, swing frequency, etc. can be statistically analyzed. In addition, based on the tracking results of each frame of the image, the behavior categories of model animals can be further analyzed based on unsupervised methods such as clustering and key point-based behavior recognition methods, and behavioral spectrum research can be carried out to provide necessary and important observation information for research in brain science, animal behavior, life science, etc.
[0043] In combination with the above example, when each frame image includes three nematodes, the tracking result of each frame image includes the predicted bounding box and multiple key points of nematode 1, the predicted bounding box and multiple key points of nematode 2, and the predicted bounding box and multiple key points of nematode 3. Among them, nematode 1, nematode 2 and nematode 3 are unique identifiers of the three nematodes, which can also be called target indexes of the three nematodes.
[0044] For details, see Figure 3 , Figure 3 A structural schematic diagram of a multi-animal posture tracking model provided by an embodiment of the present invention, the multi-animal posture tracking model 300 includes: a proposal generation module 310, a feature extraction module 320, a query decoding module 330 and a query interaction module 340; the proposal generation module 310 is used to determine a proposal query for each frame of an image, the proposal query includes multiple instance-level queries and key point-level queries corresponding to each instance-level query, the instance-level query includes a candidate bounding box and a confidence score of an animal target, and the key point-level query includes multiple candidate key points of the animal target; the feature extraction module 320 is used to: determine the image features of each frame of the image, the image features include local features and global context Features; the query decoding module 330 is used to determine the hidden state according to the proposal query and image features of the 1st frame image when the image is the 1st frame image; it is also used to determine the hidden state according to the proposal query, image features and tracking query of the t-1th frame image when the image is the t-th frame image, and to determine the tracking result of each frame image according to the hidden state of each frame image, the tracking query of the t-1th frame image carries the tracking information of each animal target included in the t-1th frame image, and t is a positive integer greater than 1; the query interaction module 340 is used to determine the tracking query of the t-1th frame image according to the hidden state corresponding to the t-1th frame image.
[0045] In one example, the proposal generation module 310 may be a proposal generator (PG), the feature extraction module 320 may be a backbone network and encoder (Backbone & Encoder), the query decoding module 330 may be a joint query decoder (Joint Query Decoder, JQD), and the query interaction module 340 may be a multi-granularity query interaction module (Multi-Granularity Query Interaction Module, MGQIM). It should be understood that the above proposal generation module 310, feature extraction module 320, query decoding module 330 and query interaction module 340 can also be implemented in other ways. The above is only an exemplary description. The embodiment of the present invention does not particularly limit the specific implementation of the proposal generation module 310, feature extraction module 320, query decoding module 330 and query interaction module 340.
[0046] In a possible implementation, the proposal generation module 310 is specifically used to: determine multiple instance-level proposals corresponding to each frame of the image; determine the key point-level proposal corresponding to each instance-level proposal with each candidate bounding box of each instance-level proposal as the center; generate a proposal query according to the multiple instance-level proposals corresponding to each frame of the image and the key point-level proposal corresponding to each instance-level proposal; For the t-th frame image, instance-level proposal The data structure is: ; ( ) is the center coordinate of the candidate bounding box of the i-th instance-level proposal; is the width of the candidate bounding box of the i-th instance-level proposal; is the height of the candidate bounding box of the i-th instance-level proposal; is the confidence score of the candidate bounding box of the i-th instance-level proposal; is the number of instance-level proposals; Keypoint-level proposals The data structure is: ; = ; is the set of key points corresponding to the i-th instance-level proposal, is the number of key points corresponding to the i-th instance-level proposal. is the coordinate of the kth key point corresponding to the i-th instance-level proposal.
[0047] Specifically, the proposal generation module 310 generates proposals at two granularities (instance-level proposals and key-point-level proposals) at the instance-level and key-point-level during the training and reasoning phases of the multi-animal posture tracking model 300. For each frame of the continuous multi-frame image, the proposal generation module 310 generates multiple instance-level proposals, including the center coordinates, width, height, and confidence score (equivalent to the candidate bounding box) of the animal target. At the same time, the key-point-level proposal is initialized with the center point of the candidate bounding box of each instance-level proposal.
[0048] In one example, the proposal generation module 310 determines multiple instance-level proposals and key-point-level proposals corresponding to each frame of the image based on the visual language large model Grounding DINO, where Grounding DINO is based on DINO (Detection Transformer with Improved deNoising anchOr boxes) and semantic positioning (grounding) technology. It can accurately locate animal targets with specific semantics in each frame of the image in complex scenes, and obtain multiple instance-level proposals and key-point-level proposals corresponding to each frame of the image. This effectively solves the problem of lack of detection models for special targets in space science experiments. Based on the visual language large model (VLM), this method can achieve zero-sample proposal generation, making it more versatile and adaptable.
[0049] The following is an exemplary description of the generation process of the instance-level query included in the proposal query provided in the embodiment of the present invention. First, the proposal generation module 310 constructs a shared query vector The shared query vector is then broadcast to the preset shape , to match the number of instance-level proposals. Then, sine-cosine positional encoding is used to encode the confidence scores. Encode and generate score embedding , and add it to the shared query vector to form the final instance-level query.
[0050] Instance-level query The formula for determining is: + ; Furthermore, before the proposal generation module 310 generates the instance-level query included in the proposal query, the method provided in the embodiment of the present invention further includes: The proposal generation module 310 generates M learnable anchor points, and concatenates the M learnable anchor points with the instance-level proposal to optimize the instance-level proposal.
[0051] The method provided by the embodiment of the present invention can optimize instance-level proposals and enhance the recall capability of animal targets by splicing M learnable anchor points with instance-level proposals. That is, it can ensure the robustness of the multi-animal posture tracking model 300 when the proposal generation module 310 misses detection.
[0052] In some embodiments, see Figure 4 , Figure 4 A structural schematic diagram of a feature extraction module provided for an embodiment of the present invention, the feature extraction module 320 includes a convolutional neural network 3201 and a Transformer encoder 3202; the convolutional neural network 3201 is used to extract features from each frame of image to obtain a feature map corresponding to each frame of image, wherein the feature map carries local features; the Transformer encoder 3202 is used to determine the global context features corresponding to each frame of image through a self-attention mechanism based on the feature map corresponding to each frame of image to obtain image features corresponding to each frame of image.
[0053] In some embodiments, see Figure 5The query decoding module 330 includes an internal instance self-attention layer 3301, a cross-instance self-attention layer 3302 and a hierarchical cross-attention layer 3303; the internal instance self-attention layer 3301 is used to perform self-attention calculations according to the key point-level queries included in the proposal query, obtain the association relationship between the key points corresponding to multiple animal targets, and obtain an updated key point-level query; the cross-instance self-attention layer 3302 is used to perform self-attention calculations according to the instance-level queries included in the proposal query, obtain the global relationship between multiple animal targets, and obtain an updated instance-level query; the hierarchical cross-attention layer 3303 is used to perform instance-level cross-attention calculations and key point-level cross-attention calculations on the updated key point-level queries, the updated instance-level queries, the image features and the tracking queries to obtain the hidden state.
[0054] Self-attention of self-attention layer 3301 within the instance The calculation formula is: ; Self-attention across instance self-attention layer 3302 ; Cross-Attention Layer 3303 Cross-Attention ; are the query matrix, key matrix and value matrix of key point level query respectively; d is the feature dimension; is the dot product of the keypoint-level query and the key; They are the query matrix, key matrix, and value matrix for instance-level queries, respectively; is the dot product of the instance-level query and the key; Represent the key matrix and value matrix of image features respectively, is the number of spatial locations of image features, is the query matrix of the tracking query and the updated instance-level query or keypoint-level query.
[0055] Specifically, the instance-internal self-attention layer 3301 acts on each animal target based on the within-instance self-attention mechanism, so that each animal target and the key points of each animal target are calculated with attention, and the relationship between each animal target and the key points is captured. This attention mechanism is based on the Transformer's query-key-value (QKV) calculation method. The cross-instance self-attention layer 3302 acts on instance-level queries of different animal targets based on the across-instance self-attention mechanism to model the global relationship between multiple animal targets, thereby enhancing the robustness of the multi-animal posture tracking model 300 in crowded or occluded scenes.
[0056] Furthermore, the hierarchical cross-attention layer 3303 is used to connect high-level queries (candidate queries and tracking queries) with underlying visual features (image features), so that candidate queries, tracking queries and image features interact directly to achieve information fusion. The hierarchical cross-attention layer 3303 includes an instance-level cross-attention (ILCA) mechanism and a keypoint-level cross-attention (KLCA) mechanism. The instance-level cross-attention mechanism enables instance-level queries to interact with global image features to optimize the category prediction and bounding box regression of animal targets. The keypoint-level cross-attention enables keypoint-level queries to interact with local image features to accurately locate keypoint positions.
[0057] In a possible implementation, the query interaction module 340 is specifically used to: when the image is the first frame image, generate an instance-level trajectory query vector and a key point-level trajectory query vector for each animal target included in the first frame image according to the hidden state of the first frame image.
[0058] When the image is the t-th frame image, based on the time-related attention mechanism, the instance-level trajectory query vector of each animal target included in the t-1-th frame image and the hidden state of the t-th frame image are determined; based on the time-related attention mechanism, the key point-level trajectory query vector of each animal target included in the t-1-th frame image and the hidden state of the t-th frame image are determined; based on the time-related attention mechanism, the key point-level trajectory query vector of each animal target included in the t-1-th frame image and the hidden state of the t-th frame image are determined; and the tracking query corresponding to the t-th frame image is generated according to the instance-level trajectory query vector and the key point-level trajectory query vector of each animal target included in the t-th frame image.
[0059] Furthermore, the instance-level trajectory query vector of each animal target included in the t-th frame image is The calculation formula is: ; in, For multi-head self-attention operation, is the instance-level trajectory query vector for each animal target included in the t-1 frame image, is the instance-level hidden state of the t-th frame image; The key point level trajectory query vector of each animal target included in the t-th frame image The calculation formula is: ; is the key point level trajectory query vector for each animal target included in the t-1 frame image, is the hidden state of the key point level of the t-th frame image.
[0060] In one example, see Figure 6 The query interaction module 340 includes an instance-wise TAN 3401 and a keypoint-wise TAN 3402, wherein the instance-wise TAN 3401 and the keypoint-wise TAN 3402 are connected in parallel. The instance-wise TAN 3401 determines the instance-level trajectory query vector of each animal target included in the t-1 frame image based on the time-related attention mechanism according to the instance-level trajectory query vector of each animal target included in the t-1 frame image and the hidden state of the t-frame image, and the keypoint-wise TAN 3402 determines the keypoint-level trajectory query vector of each animal target included in the t-1 frame image based on the time-related attention mechanism according to the keypoint-level trajectory query vector of each animal target included in the t-1 frame image and the hidden state of the t-frame image.
[0061] The specific implementation methods of the instance-level time aggregation network 3401 and the key point-level time aggregation network 3402 are detailed in the calculation process of the above instance-level trajectory query vector and the key point-level trajectory query vector, which will not be repeated here.
[0062] In a possible implementation, before determining the instance-level trajectory query vector of each animal target included in the t-th frame image according to the instance-level trajectory query vector of each animal target included in the t-1th frame image and the hidden state of the t-th frame image based on the time-correlated attention mechanism, the method provided by the embodiment of the present invention further includes: A trajectory set is constructed according to the hidden state of the first frame image, wherein the trajectory set includes a predicted bounding box, multiple key points, a confidence score and a unique identifier of each animal target; the trajectory set is updated according to the hidden state of the t-th frame image to obtain a trajectory set corresponding to the t-th frame image; invalid animal targets included in the trajectory set corresponding to the t-th frame image are determined, wherein the invalid animal targets do not exist in the multiple animal targets included in the trajectory set corresponding to the t-1-th frame image and the confidence score is less than or equal to a preset threshold; the invalid animal targets are deleted from the trajectory set corresponding to the t-th frame image to obtain a filtered trajectory set corresponding to the t-th frame image; and the hidden state of the t-th frame image is generated according to the filtered trajectory set corresponding to the t-th frame image.
[0063] Specifically, the input of the query interaction module 340 is the hidden state and the instance-level prediction score of the hidden state output by the query decoding module 330. The query interaction module 340 filters the active tracks by confidence score and identity consistency.
[0064] In one example, combining Figure 6 The query interaction module 340 also includes a trajectory screening layer 3403, which is used to: when the input is the hidden state of the first frame, construct a trajectory set according to the hidden state of the first frame image; when the input is the hidden state of the tth frame (other frames except the first frame), update the trajectory set according to the hidden state of the tth frame image to obtain the trajectory set corresponding to the tth frame image; determine the invalid animal target included in the trajectory set corresponding to the tth frame image, the invalid animal target does not exist in the multiple animal targets included in the trajectory set corresponding to the t-1th frame image and the confidence score is less than or equal to a preset threshold; delete the invalid animal target from the trajectory set corresponding to the tth frame image to obtain the trajectory set corresponding to the tth frame image after screening; generate the hidden state of the tth frame image according to the trajectory set corresponding to the t-1th frame image.
[0065] It can also be understood as: the trajectory screening layer 3403 first constructs a trajectory set according to the hidden state of the first frame image. For the t-th frame image, the trajectory screening layer 3403 updates the trajectory set of the previous frame image (t-1 frame) according to the hidden state of the t-th frame image to delete invalid animal targets in the trajectory set of the previous frame image to obtain a screened trajectory set. Then, the hidden state corresponding to the screened trajectory set (equivalent to the optimized hidden state of the t-th frame image) is input into the instance-level time aggregation network 3401 and the key point-level time aggregation network 3402. Then, the instance-level time aggregation network 3401 and the key point-level time aggregation network 3402 output the tracking query of the t-th frame image according to the corresponding instance-level part, key point-level part and tracking query of the previous frame image in the hidden state.
[0066] In one example, the specific implementation of the track screening layer 3403 for screening active tracks by confidence score and identity consistency is as follows: First, the trajectory screening layer 3403 constructs a trajectory set based on the hidden state , where each trajectory Object index with animal target (equivalent to a unique identifier), confidence score And the IoU value with the animal target in the current frame image The query interaction module 340 adopts different strategies for trajectory screening during the training and reasoning stages of the multi-animal posture tracking model 300 .
[0067] During the training phase, the trajectory filtering layer 3403 only retains the object index that is valid Or trajectories with confidence scores above a preset threshold ,Right now: in, is a preset threshold. In one example, the preset threshold is 0.5. In addition, if the IoU between the trajectory and the current detection is lower than the set threshold , the object index of the animal target is reset to invalid (equivalent to being determined as an invalid animal target): During the inference phase, the trajectory filtering layer 3403 only retains trajectories with valid object indices, i.e.: Through the above-mentioned screening strategy, the trajectory screening layer 3403 can effectively remove trajectories with low confidence or long-term unmatched trajectories, and complete the screening optimization of the hidden state of each frame image, so that the instance-level time aggregation network 3401 and the key point-level time aggregation network 3402 can obtain the tracking query of the current frame based on the optimized hidden state and the tracking query of the previous frame image, which can effectively improve the accuracy of the model, thereby improving the robustness of animal target tracking, especially in complex scenes, and can effectively reduce the risk of drift and false association.
[0068] In some embodiments, see Figure 7 Before determining the tracking result of each frame of image according to the multiple frames of image based on the multi-animal posture tracking model, the method provided by the embodiment of the present invention further includes: S71. Acquire multiple training samples, each training sample including multiple consecutive multi-frame images and tracking results corresponding to each frame of the multiple consecutive multi-frame images.
[0069] Specifically, training samples are crucial for temporal modeling of trajectories. Similar to the baseline method, this method learns the temporal variation characteristics of the target posture in a data-driven manner, rather than relying on hand-designed heuristic methods such as Kalman filtering. The common training strategy based on two adjacent frames is difficult to effectively cover the target motion pattern over a long time span. This study follows the baseline method, using continuous multi-frame images as training sample input, while introducing key point-level training samples, and combining instance-level and key point-level multi-granularity learning strategies to enhance the model's ability to model long-term motion. Through this improvement, the model can more accurately capture the local structural changes of the target, thereby improving the stability of the key point trajectory and the robustness of the overall tracking.
[0070] S72. Construct a target loss function.
[0071] S73. Based on the target loss function, the multi-animal posture tracking model is iteratively trained according to multiple training samples to obtain a trained multi-animal posture tracking model.
[0072] Objective loss function for: ; ; ; ; in, represents the number of all targets in the i-th frame, The target number of animals tracked; Target number for newborn animals; is the loss weight coefficient; Focal loss for class prediction; is the mean absolute error loss of bounding box overlap; is the mean absolute error loss of target frame overlap; k is the number of key points of the animal target, and Respectively represent The coordinates of the predicted key points and the true key points.
[0073] The method provided by the embodiment of the present invention can effectively improve the modeling ability of the model for long-term motion by introducing the posture-aware collective average loss (PCAL), especially in the accurate tracking of target key points. By taking into account the local motion details of the animal target, PCAL not only enhances the temporal learning ability of the overall target motion, but also improves the capture effect of the long-term target motion pattern under the multi-granularity learning strategy, thereby significantly improving the robustness and accuracy of the target tracking of the multi-animal posture tracking model 300.
[0074] The following is an example to illustrate the beneficial effects of a spatial pattern animal posture tracking method based on proposal relay provided by an embodiment of the present invention.
[0075] Exemplarily, a validation dataset is obtained, which includes three different animal targets: Caenorhabditis elegans, Drosophila, and zebrafish. Each type of animal target has different numbers and types of key points to accurately characterize its movement and behavior patterns. The validation dataset is divided into three subsets, of which the nematode subset contains 127 videos, 21,502 frames, 47,320 annotated instances, and 5 key points for each instance; the zebrafish subset contains 10 videos, a total of 1,711 frames, 6,844 annotated instances, and 10 key points for each instance; the Drosophila subset contains 10 videos, a total of 1,082 frames, 11,972 annotated instances, and 26 key points for each fruit fly. This dataset is obtained in complex scenarios, such as the microgravity environment causing changes in movement patterns, high individual density causing occlusion and overlap, key point drift and blur affecting detection accuracy, and the need for cross-species generalization, requiring the model to have strong adaptability.
[0076] In this example, the method provided by the embodiment of the present invention is verified based on the evaluation indicators MOTAkeypoints, MOTAtotal and recall rate. MOTAkeypoints and MOTAtotal are evaluation indicators in multi-object tracking (MOT) and key point detection tasks, and are usually used to measure the overall performance of the model in multi-object key point tracking tasks. They are extended based on the MOTA (Multi-Object Tracking Accuracy) indicator.
[0077] Among them, MOTAkeypoints calculates the MOTA of specific key points (Keypoints) in multi-target tracking tasks to evaluate whether the model can accurately and stably track the key points of each target.
[0078] The calculation formula is as follows: ; Among them, GT keypoints is the total number of all ground truth trajectories, FN keypoints is the number of missed key points, FP keypoints is the number of false detections of key points, IDSW keypoints is the number of key point identity switching errors. The value range of MOTA is , the closer it is to 1, the better the tracking performance.
[0079] MOTAtotal The MOTA is averaged for each object level and keypoint level.
[0080] First, the proposal generator 310 uses the visual language model Grounding-DINO in MMDetection to generate multiple instance-level proposals and key point-level proposals corresponding to each instance-level proposal, and uses pre-trained weights and hyperparameters. The prompt words are set to "worm", "fish" and "fruit fly", corresponding to nematodes, zebrafish and fruit flies. In order to improve the recall rate of the proposal, we retain all Grounding-DINO predicted bounding boxes with confidence levels higher than 0.2 as instance-level proposals, and initialize the key points based on the center points inferred from the predicted bounding boxes to obtain key point-level proposals. Among them, nematodes, zebrafish and fruit flies contain 5, 10 and 26 key points, respectively.
[0081] The implementation of the multi-animal posture tracking model 300 is based on PyTorch. The baseline method selects MOTRv2, and ResNet50 is used as the convolutional neural network of the feature extraction module. The batch size is set to 1, and the optimizer selects AdamW.
[0082] In terms of training configuration, the training cycle of the nematode dataset is 50 rounds, the initial learning rate is set to 2e-4, and it is decayed by 10 times in the 40th round. The training cycle of the zebrafish and fruit fly datasets is 200 rounds, and the initial learning rate is also set to 2e-4, but it is decayed by 10 times in the 180th round.
[0083] In the video clip settings, the Clip size of the C. elegans dataset is set to 5, and the frame sampling step in each clip is randomly selected between 1 and 5. In contrast, the Clip size of the zebrafish and fruit fly datasets is also set to 5, but the frame sampling step is fixed to 1.
[0084] All models are initialized from the COCO pre-trained weights of Deformable DETR. Since we have improved the key point decoder based on the baseline method, the model only loads the part with the same network structure and trains the rest from scratch, so a longer number of training rounds is required to ensure convergence and performance optimization.
[0085] See Table 1, which is a comparative data table of evaluation indicators of the nematode dataset of the method provided in the embodiment of the present invention and the method in the related art. Among them, each nematode includes 5 key points, namely MOTAhead, MOTAfront, MOTAmiddle, MOTAback, and MOTAtail. Among them, we evaluated the detection-based tracking methods Bytetrack and OC-SORT to analyze the combined performance of different posture estimation methods and tracking algorithms, and generated detection boxes according to the characteristics of different methods. For the top-down method, related technology 1 is a method combining HRNet and OC-SORT, and related technology 2 is a method combining VITPose and OC-SORT. For the bottom-up method, related technology 3 is a method combining AE and OC-SORT. Related technology 4 is a method combining DEKR and OC-SORT. Related technology 5 is a method combining CID and OC-SORT. Related technology 6 is a method combining HRNet and Bytetrack, and related technology 7 is a method combining VITPose and Bytetrack. For the bottom-up method, related technology 8 is a method combining AE and Bytetrack. Related technology 9 is a method combining DEKR and Bytetrack. Related technology 10 is a method of combining CID and Bytetrack.
[0086] Table 1 As shown in Table 1, the method provided by the present invention achieves the best performance in both MOTA and MOTA_total at key points, which is 2.75% higher than VItPose+OC-SORT. The results show that the method provided by the present invention can achieve better performance in the violent and crowded scenes such as space nematodes, meeting the needs of subsequent research and analysis.
[0087] The above mainly introduces the scheme of the embodiment of the present invention from the perspective of the method. It can be understood that in order to realize the above functions, the posture tracking system 100 includes at least one of the hardware structure and software modules corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiment disclosed in this article, the embodiment of the present invention can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiment of the present invention.
[0088] The embodiment of the present invention can divide the posture tracking system 100 into functional units according to the above method example. For example, the posture tracking system 100 can be divided into functional units corresponding to various functions, or two or more functions can be integrated into one processing unit. The above integrated unit can be implemented in the form of hardware or in the form of software functional units. It should be noted that the division of units in the embodiment of the present invention is schematic and is only a logical functional division. There may be other division methods in actual implementation.
[0089] For example, Figure 8 A hardware structure diagram of a posture tracking system provided by an embodiment of the present invention is shown. The posture tracking system 100 includes: an acquisition unit 110, which is used to acquire continuous multiple frames of images, each of which includes multiple animal targets; a tracking unit 120, which is used to determine the tracking result of each frame of the image based on the multiple frames of the image based on the multi-animal posture tracking model, and the tracking result includes the tracking information of each animal target, and the tracking information includes a predicted bounding box, multiple key points and a unique identifier; the multi-animal posture tracking model includes: a proposal generation module, a feature extraction module, a query decoding module and a query interaction module; the proposal generation module is used to determine the proposal query for each frame of the image, and the proposal query includes multiple instance-level queries and key point-level queries corresponding to each instance-level query, and the instance-level query includes a candidate bounding box and a confidence score of the animal target, and the key point-level query includes multiple candidate key points; the feature extraction module is used to: determine the image features of each frame of the image, the image features include local features and global context features; the query decoding module is used to determine the hidden state according to the proposal query and image features of the first frame of the image when the image is the first frame of the image; it is also used to determine the hidden state according to the proposal query, image features and tracking query of the t-1 frame of the image when the image is the t-th frame of the image, and to determine the tracking result of each frame of the image according to the hidden state of each frame of the image, the tracking query of the t-1 frame of the image carries the tracking information of each animal target included in the t-1 frame of the image, and t is a positive integer greater than 1; the query interaction module is used to determine the tracking query of the t-1 frame of the image according to the hidden state corresponding to the t-1 frame of the image.
[0090] It should be understood that the specific description of the above optional methods can refer to the above method embodiments, which will not be repeated here. In addition, the explanation of any posture tracking system 100 provided above and the description of the beneficial effects can refer to the above corresponding method embodiments, which will not be repeated here.
[0091] The embodiment of the present invention further provides a computer-readable storage medium, in which at least one computer instruction is stored, and the at least one computer instruction is loaded and executed by a processor to implement the methods of the above embodiments. For the explanation of the relevant contents and the description of the beneficial effects in any of the above-mentioned computer-readable storage media, reference can be made to the above-mentioned corresponding embodiments, which will not be repeated here.
[0092] The embodiment of the present invention further provides a chip. The chip integrates a control circuit and one or more ports for implementing the functions of the above-mentioned posture tracking system 100. Optionally, the functions supported by the chip can be referred to above and will not be described in detail here.
[0093] Those skilled in the art will appreciate that all or part of the steps of the above embodiments can be implemented by instructing the relevant hardware through a program, and the program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a random access memory, etc. The above-mentioned processing unit or processor can be a central processing unit, a general-purpose processor, a specific circuit structure (application specific integrated circuit, ASIC), a microprocessor (digital signal processor, DSP), a field programmable gate array (field prograllable gatearray, FPGA) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof.
[0094] The embodiment of the present invention also provides a computer program product including instructions, which, when executed on a computer, enables the computer to perform any of the methods in the above embodiments. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present invention is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media integrated therein. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., an SSD), etc.
[0095] It should be noted that the above-mentioned devices for storing computer instructions or computer programs provided in the embodiments of the present invention, such as but not limited to the above-mentioned memories, computer-readable storage media, and communication chips, etc., are all non-transitory. Those skilled in the art should be aware that in one or more of the above examples, the functions described in the embodiments of the present invention can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable storage medium or transmitted as one or more instructions or codes on a computer-readable storage medium. Computer-readable storage media include computer storage media and communication media, wherein the communication medium includes any medium that facilitates the transmission of computer programs from one place to another. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0096] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present invention. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present invention.
Claims
1. A spatial pattern animal posture tracking method based on proposal relay, characterized in that: The method comprises: Acquire a plurality of continuous frames of images, wherein each frame of the plurality of frames of images includes a plurality of animal targets; Based on a multi-animal posture tracking model, determining a tracking result of each frame of the image according to the multiple frames of the image, the tracking result including tracking information of each animal target, the tracking information including a predicted bounding box, a plurality of key points and a unique identifier; The multi-animal posture tracking model includes: a proposal generation module, a feature extraction module, a query decoding module and a query interaction module; the proposal generation module is used to determine the proposal query for each frame of the image, the proposal query includes multiple instance-level queries and key point-level queries corresponding to each instance-level query, the instance-level query includes the candidate bounding box and confidence score of the animal target, and the key point-level query includes multiple candidate key points of the animal target; the feature extraction module is used to: determine the image features of each frame of the image, the image features include local features and global context features; the query decoding module is used to In the case of 1 frame image, the hidden state is determined according to the proposal query and image features of the 1st frame image; it is also used to determine the hidden state according to the proposal query, image features and tracking query of the t-1th frame image when the image is the t-th frame image, and to determine the tracking result of each frame image according to the hidden state of each frame image, the tracking query of the t-1th frame image carries the tracking information of each animal target included in the t-1th frame image, and t is a positive integer greater than 1; the query interaction module is used to determine the tracking query of the t-1th frame image according to the hidden state corresponding to the t-1th frame image.
2. The method according to claim 1, characterized in that The proposal generation module is specifically used for: Determine multiple instance-level proposals corresponding to each frame image; Centering each candidate bounding box of each instance-level proposal, determine the keypoint-level proposal corresponding to each instance-level proposal; Generate a proposal query based on multiple instance-level proposals corresponding to each frame image and the key point-level proposal corresponding to each instance-level proposal; For the t-th frame image, instance-level proposal The data structure is: ; ( ) is the center coordinate of the candidate bounding box of the i-th instance-level proposal; is the width of the candidate bounding box of the i-th instance-level proposal; is the height of the candidate bounding box of the i-th instance-level proposal; is the confidence score of the candidate bounding box of the i-th instance-level proposal; is the number of instance-level proposals; Keypoint-level proposals The data structure is: ; = ; is the set of key points corresponding to the i-th instance-level proposal, is the number of key points corresponding to the i-th instance-level proposal.
3. The method according to claim 2, characterized in that The feature extraction module includes a convolutional neural network and a Transformer encoder; The convolutional neural network is used to extract features from each frame of image to obtain a feature map corresponding to each frame of image, wherein the feature map carries local features; The Transformer encoder is used to determine the global context features corresponding to each frame of the image through a self-attention mechanism according to the feature map corresponding to each frame of the image, so as to obtain the image features corresponding to each frame of the image.
4. The method according to claim 3, characterized in that The query decoding module includes an intra-instance self-attention layer, a cross-instance self-attention layer, and a hierarchical cross-attention layer; The self-attention layer within the instance is used to perform self-attention calculation according to the key point level query included in the proposal query, obtain the association relationship between the key points corresponding to the multiple animal targets, and obtain an updated key point level query; The cross-instance self-attention layer is used to perform self-attention calculation according to the instance-level query included in the proposal query, obtain the global relationship between the multiple animal targets, and obtain an updated instance-level query; The hierarchical cross attention layer is used to perform instance-level cross attention calculation and key-point-level cross attention calculation on the updated key-point-level query, the updated instance-level query, the image feature and the tracking query to obtain a hidden state; The self-attention of the self-attention layer within the instance The calculation formula is: ; The self-attention of the cross-instance self-attention layer ; The cross attention layer ; are the query matrix, key matrix and value matrix of key point level query respectively; d is the feature dimension; is the dot product of the keypoint-level query and the key; They are the query matrix, key matrix, and value matrix for instance-level queries respectively; is the dot product of the instance-level query and the key; Represent the key matrix and value matrix of image features respectively, is the number of spatial locations of image features, is the query matrix of the tracking query and the updated instance-level query or keypoint-level query.
5. The method according to claim 4, characterized in that The query interaction module is specifically used for: In the case where the image is the first frame image, an instance-level trajectory query vector and a key-point-level trajectory query vector of each animal target included in the first frame image are generated according to the hidden state of the first frame image; In the case where the image is the t-th frame image, based on the time-correlated attention mechanism, the instance-level trajectory query vector of each animal target included in the t-1-th frame image and the hidden state of the t-th frame image are determined; Based on the time-correlated attention mechanism, a key point-level trajectory query vector of each animal target included in the t-1 frame image is determined according to the key point-level trajectory query vector of each animal target included in the t-1 frame image and the hidden state of the t frame image; A tracking query corresponding to the t-th frame image is generated according to the instance-level trajectory query vector and the key-point-level trajectory query vector of each animal target included in the t-th frame image.
6. The method according to claim 5, characterized in that The instance-level trajectory query vector of each animal target included in the t-th frame image The calculation formula is: ; in, For multi-head self-attention operation, is the instance-level trajectory query vector for each animal target included in the t-1 frame image, is the instance-level hidden state of the t-th frame image; The key point level trajectory query vector of each animal target included in the t-th frame image The calculation formula is: ; is the key point level trajectory query vector for each animal target included in the t-1 frame image, is the hidden state of the key point level of the t-th frame image.
7. The method according to claim 6, characterized in that Before determining the instance-level trajectory query vector of each animal target included in the t-th frame image according to the instance-level trajectory query vector of each animal target included in the t-1th frame image and the hidden state of the t-th frame image based on the time-related attention mechanism, the method further includes: Constructing a trajectory set according to the hidden state of the first frame image, wherein the trajectory set includes a predicted bounding box, a plurality of key points, a confidence score, and a unique identifier for each animal target; Update the trajectory set according to the hidden state of the t-th frame image to obtain the trajectory set corresponding to the t-th frame image; Determine an invalid animal target included in the trajectory set corresponding to the t-th frame image, where the invalid animal target does not exist in the multiple animal targets included in the trajectory set corresponding to the t-1-th frame image and the confidence score is less than or equal to a preset threshold; The invalid animal target is deleted from the trajectory set corresponding to the t-th frame image, and the trajectory set corresponding to the t-th frame image after screening is obtained; The hidden state of the t-th frame image is generated according to the trajectory set corresponding to the filtered t-th frame image.
8. The method according to claim 7, characterized in that Before determining the tracking result of each frame of image according to the multiple frames of image based on the multi-animal posture tracking model, the method further includes: Acquire multiple training samples, each training sample including multiple consecutive multi-frame images and a tracking result corresponding to each frame of the multiple consecutive multi-frame images; Construct the target loss function; Based on the target loss function, the multi-animal posture tracking model is iteratively trained according to multiple training samples to obtain a trained multi-animal posture tracking model; The objective loss function for: ; ; ; ; in, represents the number of all targets in the i-th frame, The target number of animals tracked; Target number for newborn animals; is the loss weight coefficient; Focal loss for class prediction; is the mean absolute error loss of bounding box overlap; is the mean absolute error loss of target frame overlap; k is the number of key points of the animal target, and Respectively represent The coordinates of the predicted key points and the true key points.
9. A spatial pattern animal posture tracking system based on proposal relay, characterized in that: The system comprises: An acquisition unit, used for acquiring a plurality of consecutive frames of images, wherein each frame of the plurality of frames of images includes a plurality of animal targets; A tracking unit, configured to determine a tracking result of each frame of image according to the multiple frames of image based on a multi-animal posture tracking model, wherein the tracking result includes tracking information of each animal target, and the tracking information includes a predicted bounding box, a plurality of key points, and a unique identifier; The multi-animal posture tracking model includes: a proposal generation module, a feature extraction module, a query decoding module and a query interaction module; the proposal generation module is used to determine the proposal query for each frame of the image, the proposal query includes multiple instance-level queries and key point-level queries corresponding to each instance-level query, the instance-level query includes the candidate bounding box and confidence score of the animal target, and the key point-level query includes multiple candidate key points of the animal target; the feature extraction module is used to: determine the image features of each frame of the image, the image features include local features and global context features; the query decoding module is used to In the case of 1 frame image, the hidden state is determined according to the proposal query and image features of the 1st frame image; it is also used to determine the hidden state according to the proposal query, image features and tracking query of the t-1th frame image when the image is the t-th frame image, and to determine the tracking result of each frame image according to the hidden state of each frame image, the tracking query of the t-1th frame image carries the tracking information of each animal target included in the t-1th frame image, and t is a positive integer greater than 1; the query interaction module is used to determine the tracking query of the t-1th frame image according to the hidden state corresponding to the t-1th frame image.
10. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the instructions to implement the spatial pattern animal posture tracking method based on proposal relay as described in any one of claims 1-8.
Citation Information
Patent Citations
Animal tracking and attitude estimation method and device, electronic equipment and storage medium
CN116543006A
Video multi-target tracking method based on multi-scale channel feature aggregation
CN117173217A
Multi-target tracking method for endowing tracking proposal propagation by diffusion model
CN117893570A
Key point extracting and tracking method for movement of space nematodes in space station
CN118379327A
Structure for preventing the spread of fire in parking lots
KR102968787B1