A Drosophila multi-object tracking method and system based on space science experiment videos
By extracting optical flow information from the fruit fly space scientific experimental video and performing multimodal feature fusion, the trajectory prediction failure and identity switching problems of fruit fly multi-target tracking in microgravity environment are solved, and accurate and stable tracking of fruit fly is achieved.
Patent Information
- Application Number
- CN202510541648.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-28
AI Technical Summary
The existing multi-objective tracking technology cannot stably and accurately achieve multi-objective tracking of fruit fly in microgravity environments, especially in nonlinear motion, frequent occlusion and high density, resulting in trajectory prediction failure and identity switching.
Using a multimodal feature fusion method based on optical flow information, the frame image optical flow information of the fruit fly space scientific experiment video is extracted, the microgravity motion characteristics are decoupled, and the multimodal progressive fusion is carried out in combination with static appearance characteristics, conditioned prediction and offset prediction association are performed, and the fruit fly motion trajectory is updated.
Accurate and stable multi-objective tracking of fruit flies in microgravity environments is achieved, trajectory accuracy and identity consistency are improved, and false detection and missed detection are reduced.
Smart Images

Figure CN120070510B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target tracking, and particularly relates to a Drosophila multi-target tracking method and system based on space science experiment videos. Background Art
[0002] With the rapid development of space station technology, large-scale multidisciplinary space science research can be carried out in the space station. For example, Drosophila observation experiments can be conducted through the life ecology experiment cabinet carried in the space station laboratory. Specifically, according to the obtained Drosophila space science experiment videos, multi-target tracking technology can be used to locate multiple Drosophila targets, maintain their identity labels, and generate the respective movement trajectories of the Drosophila targets, thereby providing reliable technical support for in-depth analysis of the behavioral characteristics and gene expression laws of organisms under space microgravity and submagnetic environments.
[0003] Generally, multi-target tracking technology can adopt a tracking method of detecting first and then associating and a method of jointly detecting and tracking. Among them, the tracking method of detecting first and then associating decouples the tracking task into two stages of target detection and data association, and improves the overall tracking performance by independently optimizing the detection accuracy and association strategy. For example, the target center can be located through a detection box, and on this basis, the linear movement trajectory can be predicted using Kalman filtering. The method of jointly detecting and tracking converts the detector into a tracker and combines the two tasks of detection and tracking in the same detection framework. For example, low-dimensional re-identification features can be extracted by sharing the backbone network.
[0004] However, the Drosophila observation experiments conducted in the space station have the following characteristics: the movement under microgravity environment is nonlinear, random, and highly dynamic; the Drosophila targets have high density and frequent occlusion; the appearance similarity of the Drosophila targets is high; and there is a stable scene under a fixed field of view. For the Drosophila space science experiment videos obtained in the space station, the above-mentioned related technologies of multi-target tracking do not model the microgravity nonlinear movement, ignore the significant differences between the characteristics of Drosophila targets and the data distribution of Drosophila space science experiment videos, resulting in the deviation of the multi-target detection center, causing feature ambiguity, identity switching in frequently occluded scenes, and the failure of nonlinear trajectory prediction. As a result, it is impossible to stably and accurately achieve multi-target tracking of Drosophila, and it is difficult to meet the scientific requirements of trajectory accuracy and long-term identity consistency for Drosophila behavior analysis in the space station. Summary of the Invention
[0005] The technical problem to be solved by the present invention is the problem of being unable to stably and accurately achieve multi-target tracking of Drosophila.
[0006] To solve the above technical problem, the present invention provides a Drosophila multi-target tracking method and system based on space science experiment videos, and specifically adopts the following technical solutions:
[0007] In a first aspect, the present invention provides a method for multi-object tracking of fruit flies based on a space science experiment video, including: First, extract the t-th frame image, the (t + 1)-th frame image, and the (t - 1)-th frame image from the fruit fly space science experiment video, where t is a positive integer greater than 1. Then, extract optical flow information based on the t-th frame image and the (t + 1)-th frame image to determine the first optical flow image corresponding to the t-th frame image; extract optical flow information based on the t-th frame image and the (t - 1)-th frame image to determine the second optical flow image corresponding to the (t - 1)-th frame image. Next, perform feature extraction on the t-th frame image, the first optical flow image, the (t - 1)-th frame image, and the second optical flow image respectively to determine the first static appearance feature, the first moving optical flow feature, the second static appearance feature, the second moving optical flow feature, and the heatmap feature. The heatmap feature is used to characterize the probability distribution of the target center points in the (t - 1)-th frame image, and the target center points are used to characterize the center points of the target fruit flies. Further, perform multi-modal feature progressive fusion on the first static appearance feature, the first moving optical flow feature, the second static appearance feature, the second moving optical flow feature, and the heatmap feature to obtain the target fusion feature. Secondly, perform conditional prediction through a target detection network model based on the target fusion feature to determine the heatmap prediction result, the offset prediction result, and the target size prediction result of the target center points in the t-th frame image. Finally, perform offset prediction association based on the heatmap prediction result, the offset prediction result, and the heatmap feature. In the case where the offset prediction association is successful, update the motion trajectory information of the target fruit flies corresponding to the t-th frame image according to the offset prediction result and the target size prediction result.
[0008] This method first realizes the decoupling of microgravity motion characteristics by extracting the optical flow information of the frame images in the fruit fly space science experiment video. Then, perform multi-modal feature progressive fusion on the static appearance feature, the moving optical flow feature, and the heatmap feature to obtain the target fusion feature. Next, perform conditional prediction based on the target fusion feature to determine the heatmap prediction result, the offset prediction result, and the target size prediction result of the target center points. Finally, perform offset prediction association based on the prediction results. In the case where the offset prediction association is successful, update the motion trajectory information of the target fruit flies according to the prediction results, so as to realize accurate multi-object tracking of fruit flies. This method is based on an appearance-motion dual-modal fusion fruit fly multi-object tracking framework. By extracting optical flow information to realize the decoupling of motion characteristics and performing multi-modal progressive feature fusion with the static appearance feature, it can achieve accurate and stable multi-object tracking of fruit flies.
[0009] Combined with the first aspect, in an alternative implementation, the above-mentioned extraction of optical flow information based on the t-th frame image and the (t + 1)-th frame image to determine the first optical flow image corresponding to the t-th frame image includes: First, preprocess the t-th frame image and the (t + 1)-th frame image respectively to obtain the corresponding first preprocessed image and second preprocessed image. Then, perform an image subtraction process based on the first preprocessed image and the second preprocessed image to determine the first difference image. Next, according to the first difference image, determine the significant feature points through a corner detection algorithm. According to the first preprocessed image, the second preprocessed image, and the significant feature points, determine the first optical flow vector corresponding to the pixel points in the t-th frame image. Finally, generate the first optical flow image according to the first optical flow vector corresponding to the pixel points in the t-th frame image.
[0010] In this implementation, by preprocessing the t-th frame image and the (t + 1)-th frame image, performing an image subtraction process, and determining the first optical flow vector, finally, the first optical flow image corresponding to the t-th frame image can be accurately and effectively determined according to the first optical flow vector.
[0011] Combined with the first aspect, in an alternative implementation, the above-mentioned determination of the first optical flow vector corresponding to the pixel points in the t-th frame image according to the first preprocessed image, the second preprocessed image, and the significant feature points includes: First, according to the first preprocessed image and the second preprocessed image, determine the displacement vector of the significant feature points by minimizing the optical flow error function, linearizing by first-order Taylor expansion, and solving by the least squares method. Then, according to the displacement vector of the significant feature points, determine the displacement vector corresponding to the pixel points in the t-th frame image by bilinear interpolation. Next, according to the displacement vector corresponding to the pixel points in the t-th frame image, perform optical flow field superposition by constructing an image pyramid with a preset number of layers to determine the superposed optical flow vector. Perform median filtering on the superposed optical flow vector to determine the filtered optical flow vector. Finally, perform filtering processing on the filtered optical flow vector according to a preset speed amplitude threshold to obtain the first optical flow vector.
[0012] In this implementation, as an implementation of determining the first optical flow vector corresponding to the pixel points in the t-th frame image, the displacement vector of the significant feature points is determined by constructing and solving the optical flow error function. Then, through bilinear interpolation, optical flow field superposition by constructing an image pyramid, median filtering, and filtering processing in sequence, the first optical flow vector can be accurately obtained.
[0013] In combination with the first aspect, in an alternative implementation, the above multi-modal feature progressive fusion of the first static appearance feature, the first motion optical flow feature, the second static appearance feature, the second motion optical flow feature, and the heat map feature to obtain the target fusion feature specifically includes: performing a first feature fusion on the first static appearance feature and the second static appearance feature based on the cross-attention mechanism to obtain a static appearance fusion feature. Performing a second feature fusion on the first motion optical flow feature and the second motion optical flow feature based on the cross-attention mechanism to obtain a motion optical flow fusion feature. Determining the static appearance time feature through a first feed-forward neural network according to the static appearance fusion feature and the heat map feature. Determining the motion optical flow time feature through a second feed-forward neural network according to the motion optical flow fusion feature and the heat map feature. Adding the static appearance time feature and the motion optical flow time feature to obtain an initial multi-modal feature. Performing a third feature fusion based on the cross-attention mechanism with the initial multi-modal feature as the key K and value V, and the static appearance time feature as the query Q to obtain a first multi-modal feature. Performing a fourth feature fusion based on the cross-attention mechanism with the initial multi-modal feature as the Q, and the static appearance time feature as the K and V to obtain a second multi-modal feature. Performing a fifth feature fusion based on the cross-attention mechanism with the initial multi-modal feature as the K and V, and the motion optical flow time feature as the query Q to obtain a third multi-modal feature. Performing a sixth feature fusion based on the cross-attention mechanism with the initial multi-modal feature as the Q, and the motion optical flow time feature as the K and V to obtain a fourth multi-modal feature. Performing feature concatenation on the first multi-modal feature, the second multi-modal feature, the third multi-modal feature, and the fourth multi-modal feature to obtain a multi-modal concatenated feature. Determining the target fusion feature according to the multi-modal concatenated feature through a third feed-forward neural network.
[0014] In combination with the first aspect, in an alternative implementation, the expression of the above heat map prediction result is:
[0015] 。
[0016] Among them, represents the heat map prediction result, represents the Sigmoid activation function, represents the heat map output operation of the convolutional network, represents the target fusion feature. The expression of the target size prediction result is:
[0017] 。
[0018] Among them, represents the target size prediction result, represents the rectified linear unit function, represents the size output operation of the convolutional network. The expression of the offset prediction result is:
[0019] 。
[0020] Among them, represents the offset prediction result, represents the offset output operation of the convolutional network.
[0021] Combined with the first aspect, in an alternative implementation, the loss expression of the above object detection network model is:
[0022] 。
[0023] 。
[0024] 。
[0025] 。
[0026] Among them, represents the loss of the object detection network model; represents the prediction loss of the heatmap, represents the prediction loss of the object size, represents the prediction loss of the offset; represents the number of positive samples in the training set, represents the ground truth of the heatmap, the predicted value of the heatmap, represents the weight coefficient of the positive sample, represents the weight coefficient of the negative sample in the training set; represents the predicted value of the width in the object size, represents the ground truth of the width in the object size, represents the predicted value of the height in the object size, represents the ground truth of the height in the object size; represents the predicted value of the width offset, represents the ground truth of the width offset, represents the predicted value of the height offset, represents the ground truth of the height offset.
[0027] Combined with the first aspect, in an alternative implementation, the above offset prediction association based on the heatmap prediction result, offset prediction result and heatmap feature includes: First, extract the local extreme points from the heatmap prediction result to determine the detection points corresponding to the t-th frame image. Then, determine the backtracking position corresponding to the detection points according to the offset prediction result, and the expression of the backtracking position is:
[0028] 。
[0029] Among them, represents the backtracking position, represents the abscissa of the detection point in the t-th frame image, represents the predicted value of the offset of the abscissa and ordinate of the detection point, represents the ordinate of the detection point in the t-th frame image. Finally, according to the heatmap feature, determine the heatmap response corresponding to the backtracking position, and determine whether the heatmap response corresponding to the backtracking position meets the preset response threshold.
[0030] Combined with the first aspect, in an alternative implementation, the target fruit fly motion trajectory information includes: the abscissa of the center point corresponding to the target fruit fly, the ordinate of the center point, the width, the height, and the trajectory number. In the case where the offset prediction is successfully associated, update the target fruit fly motion trajectory information corresponding to the t-th frame image according to the offset prediction result and the target size prediction result, including: superimpose the offset prediction result, the target size prediction result with the target fruit fly motion trajectory information corresponding to the (t - 1)-th frame image to obtain the target fruit fly motion trajectory information corresponding to the t-th frame image; the expression of the target fruit fly motion trajectory information corresponding to the t-th frame image is:
[0031] .
[0032] = + .
[0033] = + .
[0034] Among them, represents the -th target fruit fly motion trajectory information corresponding to the t-th frame image, represents the -th target fruit fly motion trajectory information corresponding to the (t - 1)-th frame image, represents the abscissa of the center point of the -th target fruit fly, represents the ordinate of the center point of the -th target fruit fly, represents the predicted value of the width in the target size prediction result of the -th target fruit fly, represents the predicted value of the height in the target size prediction result of the -th target fruit fly, represents the trajectory number of the -th target fruit fly, represents the abscissa of the center point of the -th target fruit fly in the (t - 1)-th frame image, represents the predicted value of the horizontal offset in the offset prediction result of the th target fruit fly, represents the vertical coordinate of the center point of the th target fruit fly in the (t - 1)-th frame image, represents the predicted value of the vertical offset in the offset prediction result of the th target fruit fly.
[0035] In a second aspect, the present invention provides a multi-target tracking system for fruit flies based on a space science experiment video. The system includes: an image extraction module, an optical flow information extraction module, a feature extraction module, a multi-modal feature progressive fusion module, a conditional prediction module, an offset prediction association module, and a trajectory generation module. Among them, the image extraction module can be used to extract the t-th frame image, the (t + 1)-th frame image, and the (t - 1)-th frame image from the fruit fly space science experiment video, where t is a positive integer greater than 1. The optical flow information extraction module can be used to extract optical flow information based on the t-th frame image and the (t + 1)-th frame image to determine the first optical flow image corresponding to the t-th frame image; extract optical flow information based on the t-th frame image and the (t - 1)-th frame image to determine the second optical flow image corresponding to the (t - 1)-th frame image. The feature extraction module can be used to extract features from the t-th frame image, the first optical flow image, the (t - 1)-th frame image, and the second optical flow image respectively to determine the first static appearance feature, the first motion optical flow feature, the second static appearance feature, the second motion optical flow feature, and the heat map feature. The heat map feature is used to characterize the probability distribution of the target center point in the (t - 1)-th frame image, and the target center point is used to represent the center point of the target fruit fly. The multi-modal feature progressive fusion module can be used to perform multi-modal feature progressive fusion on the first static appearance feature, the first motion optical flow feature, the second static appearance feature, the second motion optical flow feature, and the heat map feature to obtain the target fusion feature. The conditional prediction module can be used to perform conditional prediction through a target detection network model based on the target fusion feature to determine the heat map prediction result, the offset prediction result, and the target size prediction result of the target center point in the t-th frame image. The offset prediction association module can be used to perform offset prediction association based on the heat map prediction result, the offset prediction result, and the heat map feature. The trajectory generation module can be used to update the target fruit fly motion trajectory information corresponding to the t-th frame image according to the offset prediction result and the target size prediction result when the offset prediction association is successful.
[0036] Combined with the second aspect, in an alternative implementation, the above-mentioned multi-modal feature progressive fusion module includes a multi-modal feature progressive fusion model. The multi-modal feature progressive fusion model specifically includes: a first fusion module, a second fusion module, a first feed-forward neural network, a second feed-forward neural network, a feature addition module, a third fusion module, a fourth fusion module, a fifth fusion module, a sixth fusion module, a feature splicing module, and a third feed-forward neural network. Among them, the first fusion module is used to perform a first feature fusion on the first static appearance feature and the second static appearance feature based on the cross-attention mechanism to obtain a static appearance fusion feature. The second fusion module is used to perform a second feature fusion on the first motion optical flow feature and the second motion optical flow feature based on the cross-attention mechanism to obtain a motion optical flow fusion feature. The first feed-forward neural network is used to determine the static appearance time feature according to the static appearance fusion feature and the heat map feature. The second feed-forward neural network is used to determine the motion optical flow time feature according to the motion optical flow fusion feature and the heat map feature. The feature addition module is used to add the static appearance time feature and the motion optical flow time feature to obtain an initial multi-modal feature. The third fusion module is used to perform a third feature fusion based on the cross-attention mechanism with the initial multi-modal feature as K and V, and the static appearance time feature as Q to obtain a first multi-modal feature. The fourth fusion module is used to perform a fourth feature fusion based on the cross-attention mechanism with the initial multi-modal feature as Q, and the static appearance time feature as K and V to obtain a second multi-modal feature. The fifth fusion module is used to perform a fifth feature fusion based on the cross-attention mechanism with the initial multi-modal feature as K and V, and the motion optical flow time feature as Q to obtain a third multi-modal feature. The sixth fusion module is used to perform a sixth feature fusion based on the cross-attention mechanism with the initial multi-modal feature as Q, and the motion optical flow time feature as K and V to obtain a fourth multi-modal feature. The feature splicing module is used to splice the first multi-modal feature, the second multi-modal feature, the third multi-modal feature, and the fourth multi-modal feature to obtain a multi-modal splicing feature. The third feed-forward neural network is used to determine the target fusion feature according to the multi-modal splicing feature.
[0037] In a third aspect, the present invention provides an electronic device, including: a memory, one or more processors; the memory is coupled to the processor; wherein, the memory stores computer program code, and the computer program code includes computer instructions, when the computer instructions are executed by the processor, the electronic device is caused to execute the method provided in the first aspect and any of its alternative implementations as described above.
[0038] In a fourth aspect, the present invention provides a computer-readable storage medium, including computer instructions, when the computer instructions are run on an electronic device, the electronic device is caused to execute the method provided in the first aspect and any of its alternative implementations as described above.
[0039] Understandably, for the beneficial effects that can be achieved by the Drosophila multi-object tracking system provided in the second aspect above, the electronic device in the third aspect, and the computer-readable storage medium in the fourth aspect, reference can be made to the beneficial effects in the first aspect and any of its possible design manners, which will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 It is a schematic diagram of the principle of the Drosophila multi-object tracking method based on the space science experiment video provided by an embodiment of the present application;
[0041] Figure 2 It is a schematic flowchart of the Drosophila multi-object tracking method based on the space science experiment video provided by an embodiment of the present application;
[0042] Figure 3 It is a schematic flowchart of the method for determining the first optical flow image provided by an embodiment of the present application;
[0043] Figure 4 It is a schematic diagram of the principle of progressive fusion of multi-modal features provided by an embodiment of the present application;
[0044] Figure 5 It is a schematic structural diagram of the Drosophila multi-object tracking system based on the space science experiment video provided by an embodiment of the present application;
[0045] Figure 6 It is a schematic structural diagram of the multi-modal feature progressive fusion model provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] The embodiments will be described in detail below, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following examples do not represent all embodiments consistent with the present application. They are only examples of systems and methods consistent with some aspects of the present application detailed in the claims.
[0047] With the rapid development of space station technology, large-scale multi-disciplinary space science research can be carried out in the space station. For example, through the life ecology experiment cabinet carried in the space station laboratory, Drosophila observation experiments can be conducted. Specifically, according to the obtained Drosophila space science experiment video, multi-object tracking technology can be used to locate multiple Drosophila targets, maintain their identity identifiers, and generate the respective movement trajectories of the Drosophila targets, thereby providing reliable technical support for in-depth analysis of the behavioral characteristics and gene expression laws of organisms under space microgravity and sub-magnetic environments.
[0048] Generally, multi-object tracking techniques can adopt the tracking method of detecting first and then associating, and the method of jointly detecting and tracking. Among them, the tracking method of detecting first and then associating decouples the tracking task into two stages: object detection and data association, and improves the overall tracking performance by independently optimizing the detection accuracy and association strategy. For example, the center of the object can be located by the detection box, and on this basis, the Kalman filter is used to predict the linear motion trajectory. The method of jointly detecting and tracking converts the detector into a tracker and combines the two tasks of detection and tracking in the same detection framework. For example, low-dimensional re-identification features can be extracted by sharing the backbone network.
[0049] Exemplarily, related technology one proposed a Multiple Object Tracking and Detection (MOTDT) method based on the framework of detecting first and then associating. By combining the detection and tracking results into candidates and selecting the optimal candidate based on a deep neural network, it can cope with the interference of unreliable detections in tracking. Further, a hierarchical data association strategy is designed to improve the tracking performance. This method specifically includes: real-time object classification, trajectory confidence scoring, appearance feature representation, and hierarchical data association based on the Drosophila space science experiment video. Finally, the position trajectory is obtained through the tracking results.
[0050] Related technology two proposed the FairMOT method based on the framework of jointly detecting and tracking, which converts multi-object tracking into a pixel-level key point estimation and identity classification problem on a high-resolution feature map through a single-stage end-to-end architecture and feature alignment design. This method specifically includes: high-resolution feature extraction based on the Drosophila space science experiment video, and then detection and re-identification are performed through the detection branch and the re-identification branch respectively. Next, bounding box temporal association is carried out. Finally, the position trajectory is obtained through the tracking results.
[0051] However, the Drosophila observation experiment conducted in the space station has the following characteristics: the motion in the microgravity environment is nonlinear, random, and highly dynamic; there are high-density and frequent occlusions of Drosophila targets; the appearance similarity of Drosophila targets is high; and there is a stable scene under a fixed field of view.
[0052] The related technology one relies on the Kalman filter to predict the trajectory candidate boxes. However, the non-linear motion of Drosophila in the microgravity environment (such as high-frequency turning and hovering) will cause a significant increase in the prediction deviation of the motion model, making it difficult for the trajectory confidence calculation to accurately reflect the true tracking state. Especially in the case of long-term occlusion, it is easy to cause trajectory drift or loss. Secondly, although the MOTDT method alleviates the identity switching problem through re-identification features, the highly similar appearance features among Drosophila individuals will greatly weaken the discriminative ability of the feature space, resulting in frequent ID switching. Moreover, this method adopts an online processing mode and cannot use future frame information for trajectory backtracking and correction. It is difficult to recover the interrupted trajectory in the high-density occlusion scenario of the target, exacerbating trajectory fragmentation. In addition, the candidate box scoring mechanism of the MOTDT method depends on the quality of the detection box, and the fast dynamic motion of Drosophila is likely to cause inaccurate positioning or missed detection of the detection box, which may wrongly suppress valid tracking candidate boxes and cause trajectory initialization failure.
[0053] The related technology two adopts a linear motion prediction model based on the Kalman filter, which is not optimized for the stable background of the fixed field of view, may introduce irrelevant noise, lacks prior motion modeling based on physical constraints, and is difficult to adapt to the non-linear and high-dynamic flight trajectories of Drosophila in the microgravity environment, and there are biases in the trajectory prediction deviation. And its re-identification branch only uses 128-dimensional low-dimensional features, which are difficult to capture subtle differences when the appearances of Drosophila are highly similar, and are prone to false associations. At the same time, due to the feature conflict problem between the detection and re-identification tasks, combined with the design that only relies on the alignment of the target center point features, it will lead to feature ambiguity and identity switching in the frequent occlusion scenario.
[0054] It can be seen that for the Drosophila space science experiment videos obtained in the space station, the above-mentioned related technologies of multi-object tracking do not model the non-linear motion in microgravity, ignore the significant differences between the characteristics of Drosophila targets and the data distribution of Drosophila space science experiment videos, resulting in the deviation of the multi-object detection center, causing feature ambiguity, identity switching in the frequent occlusion scenario, and the failure of non-linear trajectory prediction. As a result, it is impossible to stably and accurately achieve multi-object tracking of Drosophila, and it is difficult to meet the scientific requirements of trajectory accuracy and long-term identity consistency for Drosophila behavior analysis in the space station.
[0055] To solve the above problems, the embodiments of the present application provide a Drosophila multi-object tracking method and system based on space science experiment videos. This method and system can be applied to Drosophila observation experiment videos collected in a stable scenario under microgravity and a fixed field of view. For example, it can be applied to Drosophila space science experiment videos collected during Drosophila observation experiments in the space station. Specifically, Figure 1 is a schematic diagram of the principle of the Drosophila multi-object tracking method based on space science experiment videos provided by the embodiments of the present application, as Figure 1As shown, this method can extract optical flow information from the image frames in the Drosophila space science experiment video to achieve microgravity motion decoupling. Then, feature extraction is performed based on the image frames and optical flow information in the Drosophila space science experiment video to determine static appearance features and motion optical flow features. Next, multimodal feature progressive fusion is performed based on the static appearance features and motion optical flow features to obtain the fused target fusion features. Further, conditional prediction is performed based on the target fusion features to obtain the heatmap prediction result, offset prediction result, and target size prediction result. Finally, offset prediction association is performed based on the prediction results. In the case where the offset prediction association is successful, the motion trajectory information of the target Drosophila is updated based on the prediction results to obtain the multi-target tracking result, thereby achieving accurate multi-target tracking of Drosophila. This method is based on an appearance-motion bimodal fusion Drosophila multi-target tracking framework. By extracting optical flow information, motion feature decoupling is achieved, and multimodal progressive feature fusion is performed with static appearance features, enabling precise and stable multi-target tracking of Drosophila.
[0056] Next, the solution provided in the embodiments of the present application will be introduced with reference to the accompanying drawings.
[0057] Specifically, Figure 2 is a schematic flowchart of the Drosophila multi-target tracking method based on the space science experiment video provided in the embodiments of the present application. As Figure 2 shown, the Drosophila multi-target tracking method based on the space science experiment video provided in the embodiments of the present application includes the following steps S101-S107:
[0058] S101. Extract the t-th frame image, the (t + 1)-th frame image, and the (t - 1)-th frame image from the Drosophila space science experiment video.
[0059] In the embodiments of the present application, the Drosophila space science experiment video is pre-acquired video data. Exemplarily, the Drosophila space science experiment video can be an experimental observation video collected from a Drosophila observation experiment carried out in the Life and Ecological Experiment Cabinet on the space station. This Drosophila space science experiment video can be used to characterize the movement process of multiple target Drosophila in the Life and Ecological Experiment Cabinet.
[0060] Specifically, the Drosophila space science experiment video is composed of multiple consecutive image frames. In the embodiments of the present application, taking the multi-target tracking of the target Drosophila in the t-th frame image corresponding to the t-th moment as an example, first, the t-th frame image, and the previous frame image and the subsequent frame image adjacent to the t-th frame image, that is, the (t - 1)-th frame image and the (t + 1)-th frame image, need to be extracted from the Drosophila space science experiment video. Among them, t is a positive integer greater than 1. The t-th frame image, the (t + 1)-th frame image, and the (t - 1)-th frame image can all be RGB frame images.
[0061] S102. Extract optical flow information based on the t-th frame image and the (t + 1)-th frame image to determine the first optical flow image corresponding to the t-th frame image; extract optical flow information based on the t-th frame image and the (t - 1)-th frame image to determine the second optical flow image corresponding to the (t - 1)-th frame image.
[0062] Specifically, the first optical flow image corresponding to the t-th frame image can be determined by extracting optical flow information from the t-th frame image and the (t + 1)-th frame image. The first optical flow image includes the optical flow information (i.e., optical flow vectors, including: direction and magnitude of motion) corresponding to the pixel points in the t-th frame image.
[0063] The second optical flow image corresponding to the (t - 1)-th frame image can be determined by extracting optical flow information from the t-th frame image and the (t - 1)-th frame image. The second optical flow image includes the optical flow information corresponding to the pixel points in the (t - 1)-th frame image.
[0064] In the embodiments of the present application, in view of the characteristics of the video background of the Drosophila space science experiment video being stable and without dynamic interference, optical flow information can be extracted from the frame images of the Drosophila space science experiment video, that is, the microgravity motion of the Drosophila space science experiment video can be decoupled through optical flow estimation based on frame difference, and the non-linear motion characteristics in the microgravity environment can be separated, such as: hovering, acceleration mutation, etc. In this way, the dynamic characteristics of the motion trajectory of the target Drosophila can be directly modeled through optical flow information, thereby effectively solving the problem of deviation in motion trajectory prediction caused by microgravity drift of the target Drosophila.
[0065] In some embodiments, Figure 3 is a schematic flowchart of the method for determining the first optical flow image provided by the embodiments of the present application. As Figure 3 shown, taking the extraction of optical flow information from the t-th frame image and the (t + 1)-th frame image to determine the first optical flow image corresponding to the t-th frame image as an example, S102 may specifically include the following steps S1021 - S1025:
[0066] S1021. Preprocess the t-th frame image and the (t + 1)-th frame image respectively to obtain the corresponding first preprocessed image and second preprocessed image.
[0067] Specifically, the preprocessing may include: image grayscale processing, Gaussian blur processing, etc. In this way, the computational complexity of subsequent image processing can be reduced, and the noise influence in the t-th frame image and the (t + 1)-th frame image can be smoothed, thereby improving the accuracy of subsequent image processing.
[0068] S1022. Perform image subtraction processing based on the first preprocessed image and the second preprocessed image to determine the first difference image.
[0069] Then, the first preprocessed image and the second preprocessed image can be subjected to image subtraction processing. For example, the frame difference method can be used for image subtraction processing to determine the first difference image. Among them, the first difference image includes the gray value difference corresponding to each pixel point.
[0070] S1023. Determine the significant feature points according to the first difference image through a corner detection algorithm.
[0071] Furthermore, whether there is a moving target (i.e., the target Drosophila) can be determined according to the first difference image, that is, the gray value difference. Specifically, the feature points with obvious structural changes (i.e., significant feature points) in the first difference image can be determined through a corner detection algorithm, and the significant feature points can be used for the target Drosophila.
[0072] Exemplarily, first, the Shi-Tomasi corner detection algorithm can be used to detect the corner response value corresponding to the pixel points in the first difference image. The expression of the corner response value is:
[0073] .
[0074] Among them, represents the corner response value, and represent the eigenvalues of the image gradient matrix in the first difference image, and the smaller of the two eigenvalues is taken as the corner response value.
[0075] Then, the pixel points corresponding to the corner response values in the first difference image that satisfy the preset corner response threshold can be determined as significant feature points.
[0076] S1024. Determine the first optical flow vector corresponding to the pixel points in the t-th frame image according to the first preprocessed image, the second preprocessed image, and the significant feature points.
[0077] Furthermore, according to the first preprocessed image and the second preprocessed image, the displacement vector of the significant feature points can be determined, and it can be extended to the pixel-level displacement of the entire image of the t-th frame through bilinear interpolation. Then, through constructing an image pyramid, filtering, and filtering processing, the first optical flow vector can be obtained.
[0078] Specifically, in some embodiments, the above S1024 may specifically include the following steps S10241 - S10245:
[0079] S10241. Determine the displacement vector of the significant feature points according to the first preprocessed image and the second preprocessed image by minimizing the optical flow error function, linearizing by first-order Taylor expansion, and solving by the least squares method.
[0080] Specifically, for the significant feature points determined in S1023, with the condition that the pixel points within the first preset neighborhood window of the significant feature points have the same motion vector, the displacement vector of the significant feature points can be determined by minimizing the optical flow error function. The expression of the displacement vector of the significant feature points is:
[0081] .
[0082] Among them, represents the displacement vector of the significant feature points, represents the first preset neighborhood window. For example: can be a 3×3 or 5×5 square region. represents the gray value of the pixel point ( ) in the first preprocessed image, represents the gray value of the pixel point ( ) in the second preprocessed image.
[0083] Then, based on the above expression of the displacement vector of the significant feature points, through first-order Taylor expansion for linearization, an approximate error function is obtained:
[0084] .
[0085] .
[0086] .
[0087] .
[0088] Among them, and represent the image space gradient, represents the time gradient.
[0089] Next, through least squares solution, a closed-form solution of the displacement vector of the significant feature points is obtained:
[0090] .
[0091] S10242. Determine the displacement vector corresponding to the pixel points in the t-th frame image according to the displacement vector of the significant feature points through bilinear interpolation.
[0092] Specifically, through bilinear interpolation, the displacement vector of the significant feature points can be extended to the pixel-level displacement of the entire image of the t-th frame image, and the displacement vector corresponding to the pixel points in the t-th frame image is obtained.
[0093] .
[0094] .
[0095] Among them, represents the displacement vector corresponding to the pixel point in the t-th frame image corresponding displacement vector, represents the nearest significant feature point, represents the significant feature point displacement vector of is the interpolation weight.
[0096] S10243. According to the displacement vector corresponding to the pixel point in the t-th frame image, perform optical flow field superposition by constructing an image pyramid with a preset number of layers, and determine the superimposed optical flow vector.
[0097] Furthermore, according to the displacement vector corresponding to the pixel point in the t-th frame image determined in S10242, an image pyramid can be constructed, and the optical flow field extracted can be optimized from coarse to fine using multi-scale representation. Calculate the optical flow field at each layer and use the result as the initial value of the optical flow field of the next layer.
[0098] Exemplarily, the expression of the optical flow field of the layer can be:
[0099] .
[0100] Among them, is the optical flow field of the layer, that is, the displacement vector at the pixel point , which can be determined according to to obtain, represents the optical flow correction amount of the optical flow field of the layer, represents the bilinear upsampling operation.
[0101] The higher the layer number ( the larger), the lower the resolution of the image, and the processing result is from coarse to fine. Therefore, first calculate the optical flow field of the top layer ( is the preset number of layers), and then upsample it to the layer, add the optical flow correction amount of this layer, and so on, until the bottom layer . The final expression of the superimposed optical flow vector after optical flow field superposition can be recursively expanded as:
[0102] .
[0103] Among them, represents the superimposed optical flow vector, represents the preset number of layers. For example: in the embodiments of the present application, It can be set to 3. Indicates the cumulative upsampling operation from the layer to layer 0, is the optical flow correction amount for the layer, and its coordinates need to be mapped to the original resolution space.
[0104] S10244. Perform median filtering on the superimposed optical flow vectors to determine the filtered optical flow vectors.
[0105] Furthermore, median filtering can be performed on the horizontal and vertical components of the superimposed optical flow vectors to remove noise and obtain the filtered optical flow vectors.
[0106] .
[0107] Among them, represents the filtered optical flow vector, represents the second preset neighborhood window, which can be the same as the first preset neighborhood window. For example: it can be a square area of 3×3 or 5×5. represents the median operation, and the horizontal and vertical components of the superimposed optical flow vectors within are sorted respectively and the middle values are taken.
[0108] S10245. Filter the filtered optical flow vectors according to a preset velocity magnitude threshold to obtain the first optical flow vectors.
[0109] Finally, the filtered optical flow vectors can be further filtered by a preset velocity magnitude threshold to filter out abnormal movements and obtain the filtered optical flow vectors, that is, the first optical flow vectors. Exemplarily, the expression of the first optical flow vectors is:
[0110] .
[0111] Among them, represents the first optical flow vectors, represents the velocity magnitude of the optical flow, represents the preset velocity magnitude threshold. Exemplarily, it can be the 95th percentile of the optical flow magnitude distribution in the image.
[0112] S1025. Generate a first optical flow image based on the first optical flow vectors corresponding to the pixel points in the t-th frame image.
[0113] Finally, a first optical flow image can be generated based on the first optical flow vectors corresponding to the pixel points in the t-th frame image determined in S1024.
[0114] It should be noted that the principle of the method for determining the second optical flow image corresponding to the (t-1)-th frame image in S102 is the same as that of the method shown in S1021-S1025 above, and will not be elaborated here.
[0115] S103. Feature extraction is respectively performed on the t-th frame image, the first optical flow image, the (t-1)-th frame image, and the second optical flow image to determine the first static appearance feature, the first motion optical flow feature, the second static appearance feature, the second motion optical flow feature, and the heat map feature.
[0116] In the embodiment of the present application, in order to perform multi-modal feature progressive fusion, feature extraction needs to be first performed on the t-th frame image, the first optical flow image, the (t-1)-th frame image, and the second optical flow image. Specifically, the feature extraction network (for example: Resnet network) can be used to perform feature extraction on the t-th frame image, the first optical flow image, the (t-1)-th frame image, and the second optical flow image respectively, to obtain the first static appearance feature corresponding to the t-th frame image, the first motion optical flow feature corresponding to the first optical flow image, the second static appearance feature corresponding to the (t-1)-th frame image, and the second motion optical flow feature corresponding to the second optical flow image. Among them, the static appearance features (the first static appearance feature, the second static appearance feature) can be used to characterize the static features of the pixel points in the frame image (the t-th frame image, the (t-1)-th frame image), such as: color value, gray value, etc. The motion optical flow features (the first motion optical flow feature, the second motion optical flow feature) can be used to characterize the dynamic features of the pixel points in the optical flow image (the first optical flow image, the second optical flow image), such as: motion direction, motion speed, etc.
[0117] Moreover, the heat map feature can also be extracted from the (t-1)-th frame image through the target detection network for feature fusion. Specifically, the heat map feature can be used to characterize the probability distribution of the target center point in the (t-1)-th frame image, and the target center point can be used to characterize the center point of the target fruit fly.
[0118] S104. Multi-modal feature progressive fusion is performed on the first static appearance feature, the first motion optical flow feature, the second static appearance feature, the second motion optical flow feature, and the heat map feature to obtain the target fusion feature.
[0119] In the embodiments of the present application, the first static appearance feature, the first motion optical flow feature, the second static appearance feature, the second motion optical flow feature, and the heat map feature extracted in S103 are fused in multiple levels and progressively. The static appearance feature extraction is guided by the optical flow field to enhance the robustness of the motion blur area detection. The subsequent generation of the conditional prediction heat map is constrained by the temporal consistency of the motion field to suppress false detections and missed detections in high-density occlusion scenarios, so as to provide an appearance-motion joint representation. Specifically, the progressive fusion of multi-modal features is divided into two fusion stages, which respectively fuse the temporal features and the multi-modal features to make full use of the multi-modal information of visual appearance and motion optical flow as well as the temporal information.
[0120] Exemplarily, the expression of the target fusion feature can be:
[0121] 。
[0122] Among them, represents the target fusion feature, represents the temporal feature fusion operation, represents the multi-modal feature fusion operation, represents the first static appearance feature, represents the first motion optical flow feature, represents the second static appearance feature, represents the second motion optical flow feature, represents the heat map feature.
[0123] For multi-object tracking, extracting temporal information is crucial for improving the performance of multi-object tracking. Since the t-th frame image and the (t-1)-th frame image are not strictly aligned in space, it is difficult to effectively integrate the feature information of the (t-1)-th frame image. Therefore, in the temporal feature fusion stage, the embodiments of the present application adopt an attention mechanism that does not rely on the strict spatial alignment of adjacent frames to effectively integrate the temporal information.
[0124] Specifically, in some embodiments, S104 may specifically include the following steps S1041-S1046:
[0125] S1041. Perform the first feature fusion on the first static appearance feature and the second static appearance feature based on the cross-attention mechanism to obtain the static appearance fusion feature; perform the second feature fusion on the first motion optical flow feature and the second motion optical flow feature based on the cross-attention mechanism to obtain the motion optical flow fusion feature.
[0126] Specifically, Figure 4 is the schematic diagram of the principle of the progressive fusion of multi-modal features provided by the embodiments of the present application. As Figure 4 shown, the first static appearance feature can be used as the query Q based on the cross-attention mechanism (i.e., Figure 4 shown in ), the second static appearance feature is used as the key K (i.e., Figure 4 shown in ), and the value V (i.e., Figure 4 shown in ) to perform the first feature fusion to capture the spatio-temporal context information in the static appearance feature, obtaining the static appearance fusion feature.
[0127] Moreover, based on the cross-attention mechanism, the first motion optical flow feature can be used as the query Q (i.e., Figure 4 shown in ), the second motion optical flow feature as the key K (i.e., Figure 4 shown in ), and the value V (i.e., Figure 4 shown in ) to perform the second feature fusion to capture the spatio-temporal context information in the motion optical flow feature, obtaining the motion optical flow fusion feature.
[0128] Exemplarily, the expression of the static appearance fusion feature can be:
[0129] .
[0130] Wherein, represents the static appearance fusion feature, represents the position encoding, represents the cross-attention operation.
[0131] The expression of the motion optical flow fusion feature can be:
[0132] .
[0133] Wherein, represents the motion optical flow fusion feature.
[0134] S1042. According to the static appearance fusion feature and the heatmap feature, determine the static appearance time feature through the first feed-forward neural network; according to the motion optical flow fusion feature and the heatmap feature, determine the motion optical flow time feature through the second feed-forward neural network.
[0135] Then, in order to enhance the positioning ability of target tracking, the heatmap feature of the (t - 1)-th frame image is used as the position condition and integrated into the static appearance fusion feature and the motion optical flow fusion feature, and finally the corresponding time feature is obtained through the feed-forward neural network to enhance the target feature representation.
[0136] Exemplarily, the expression of the static appearance time feature is:
[0137] .
[0138] .
[0139] Among them, represents the static appearance time feature, represents the layer normalization operation, represents the feed-forward neural network, represents the intermediate result of feature processing.
[0140] The expression of the motion optical flow time feature is:
[0141] .
[0142] .
[0143] Among them, represents the motion optical flow time feature.
[0144] S1043. Add the static appearance time feature and the motion optical flow time feature to obtain the initial multimodal feature.
[0145] In the progressive fusion of multimodal features adopted in the embodiments of the present application, the interaction between modalities is crucial for obtaining effective multimodal features. The cross-attention mechanism tends to enhance the similarity information (homogeneous information) between modalities, while possibly ignoring the modality-specific information (heterogeneous information). Therefore, in the multimodal fusion stage, an addition operation is used to obtain the initial multimodal feature, and the initial multimodal feature is used as a bridging feature to interact with the unimodal feature. In this way, the problems brought about by directly interacting with the unimodal feature can be effectively avoided.
[0146] Specifically, the static appearance time feature and the motion optical flow time feature can be integrated through a feature addition operation to obtain an initial multimodal feature representation, that is, the initial multimodal feature is obtained.
[0147] Exemplarily, the expression of the initial multimodal feature is:
[0148] .
[0149] Among them, represents the initial multimodal feature.
[0150] S1044. Using the initial multimodal features as the key K and value V based on the cross-attention mechanism, and the static appearance temporal features as the query Q for the third feature fusion to obtain the first multimodal feature; using the initial multimodal features as Q, the static appearance temporal features as K and V for the fourth feature fusion to obtain the second multimodal feature; using the initial multimodal features as K and V, and the motion optical flow temporal features as Q for the fifth feature fusion to obtain the third multimodal feature; using the initial multimodal features as Q, and the motion optical flow temporal features as K and V for the sixth feature fusion to obtain the fourth multimodal feature.
[0151] Furthermore, the multimodal features (i.e., the initial multimodal features) can be used as K and V, while the two unimodal features (the static appearance temporal features and the motion optical flow temporal features) are used as Q respectively. Based on the cross-attention mechanism, the unimodal features and the multimodal features are interacted to enhance the representation of the modality-specific features. At the same time, the multimodal features can be used as Q, and the two unimodal features are used as K and V respectively. Then, based on the cross-attention mechanism, the multimodal features are further enhanced.
[0152] S1045. Perform feature concatenation on the first multimodal feature, the second multimodal feature, the third multimodal feature, and the fourth multimodal feature to obtain the multimodal concatenated feature.
[0153] Exemplarily, the expression of the multimodal concatenated feature is:
[0154]
[0155] .
[0156] Wherein, represents the multimodal concatenated feature, represents the cross-attention operation, represents the first multimodal feature, represents the second multimodal feature, represents the third multimodal feature, represents the fourth multimodal feature, represents the feature concatenation operation.
[0157] S1046. Determine the target fusion feature according to the multimodal concatenated feature through the third feed-forward neural network.
[0158] Finally, input the multimodal concatenated feature into the third feed-forward neural network to obtain the refined multimodal feature, that is, the target fusion feature. Exemplarily, the target fusion feature can be specifically represented as:
[0159] .
[0160] Among them, represents the target fusion feature.
[0161] S105. Perform conditional prediction through the target detection network model according to the target fusion feature to determine the heatmap prediction result, offset prediction result, and target size prediction result of the target center point in the t-th frame image.
[0162] In the embodiment of the present application, the conditional prediction can introduce the multi-target tracking result, image feature, and motion feature of the previous frame (i.e., the (t - 1)-th frame image) according to the target fusion feature, enhance the detection robustness of the current frame (i.e., the t-th frame image), and achieve the joint optimization of detection and tracking. It not only explicitly uses the temporal information but also fuses the motion information of the optical flow field for joint reasoning.
[0163] In one implementation, the target detection network model can adopt the CenterNet model. Specifically, based on the target fusion feature determined in S104, through three parallel branches of the CenterNet model, the heatmap prediction result and offset prediction result of the target center point in the t-th frame image, and the target size prediction result can be obtained respectively.
[0164] In some embodiments, the expression of the heatmap prediction result is:
[0165] .
[0166] Among them, represents the heatmap prediction result, represents the Sigmoid activation function, represents the heatmap output operation of the convolutional network, represents the target fusion feature.
[0167] The expression of the target size prediction result is:
[0168] .
[0169] Among them, represents the target size prediction result, represents the rectified linear unit function, represents the size output operation of the convolutional network.
[0170] The expression of the offset prediction result is:
[0171] .
[0172] Among them, represents the offset prediction result, represents the offset output operation of the convolutional network.
[0173] In some embodiments, the object detection network model is a network model trained based on a sample training set. Specifically, the loss expression of the object detection network model is:
[0174] .
[0175] .
[0176] .
[0177] .
[0178] Wherein, represents the loss of the object detection network model; represents the prediction loss of the heat map, represents the prediction loss of the object size, represents the prediction loss of the offset; represents the number of positive samples in the training set, represents the ground truth of the heat map, the predicted value of the heat map, represents the weight coefficient of the positive sample, represents the weight coefficient of the negative sample in the training set; represents the predicted value of the width in the object size, represents the ground truth of the width in the object size, represents the predicted value of the height in the object size, represents the ground truth of the height in the object size; represents the predicted value of the width offset, represents the ground truth of the width offset, represents the predicted value of the height offset, represents the ground truth of the height offset.
[0179] S106. Perform offset prediction association according to the heat map prediction result, the offset prediction result, and the heat map feature.
[0180] In the embodiments of the present application, in order to solve the problem of cross-frame object matching ambiguity caused by the vulnerability of distance matching to occlusion interference and the inability to rely on re-identification features due to the similar appearance of target fruit flies. Predict the target displacement through the offset to improve the adaptability to non-linear patterns. Use the previous frame heat map (i.e., the heat map feature) as prior knowledge to constrain the search space for the detection of the t-th frame image, and reduce false detections and missed detections caused by occlusion, etc.
[0181] Specifically, in some embodiments, S106 may specifically include the following steps S1061-S1063:
[0182] S1061. Extract local extreme points from the heatmap prediction results to determine the detection points corresponding to the t-th frame image.
[0183] S1062. Determine the backtracking position corresponding to the detection point according to the offset prediction results.
[0184] Among them, the expression of the backtracking position is:
[0185] .
[0186] Among them, represents the backtracking position, represents the abscissa of the detection point in the t-th frame image, represents the predicted offset value of the abscissa and ordinate of the detection point, represents the ordinate of the detection point in the t-th frame image.
[0187] S1063. Determine the heatmap response corresponding to the backtracking position according to the heatmap features, and determine whether the heatmap response corresponding to the backtracking position meets the preset response threshold.
[0188] Specifically, the heatmap response corresponding to the backtracking position can be determined according to the heatmap features. Further, it can be determined whether the heatmap response meets the preset response threshold, that is, it can be determined whether there is a significant heatmap response near the backtracking position to determine whether the offset prediction association is successful. Among them, the preset response threshold can be preset according to prior knowledge and the requirements of actual applications.
[0189] S107. When the offset prediction association is successful, update the target fruit fly motion trajectory information corresponding to the t-th frame image according to the offset prediction results and the target size prediction results.
[0190] Finally, when the heatmap response corresponding to the backtracking position meets the preset response threshold, it can be determined that the offset prediction association is successful. In this case, the target fruit fly motion trajectory information corresponding to the t-th frame image can be updated according to the offset prediction results and the target size prediction results to obtain the multi-target tracking result of the target fruit fly.
[0191] In some embodiments, the target fruit fly motion trajectory information includes: the abscissa of the center point corresponding to the target fruit fly, the ordinate of the center point, the width, the height, and the trajectory number.
[0192] Then S107 specifically includes: superimposing the offset prediction results, the target size prediction results with the target fruit fly motion trajectory information corresponding to the (t - 1)-th frame image to obtain the target fruit fly motion trajectory information corresponding to the t-th frame image.
[0193] Among them, the expression of the target fruit fly motion trajectory information corresponding to the t-th frame image is:
[0194] 。
[0195] = + 。
[0196] = + 。
[0197] Among them, represents the th target fruit fly motion trajectory information corresponding to the t-th frame image, represents the th target fruit fly motion trajectory information corresponding to the (t - 1)-th frame image, represents the abscissa of the center point of the th target fruit fly, represents the ordinate of the center point of the th target fruit fly, represents the predicted width value in the predicted target size of the th target fruit fly, represents the predicted height value in the predicted target size of the th target fruit fly, represents the motion trajectory number of the th target fruit fly, represents the abscissa of the center point of the th target fruit fly in the (t - 1)-th frame image, represents the predicted abscissa offset value in the predicted offset of the th target fruit fly, represents the ordinate of the center point of the th target fruit fly in the (t - 1)-th frame image, represents the predicted ordinate offset value in the predicted offset of the th target fruit fly.
[0198] Adopting the fruit fly multi-target tracking method based on the space science experiment video provided in the embodiments of the present application, first, the microgravity motion characteristics are decoupled by extracting the optical flow information of the frame images in the fruit fly space science experiment video. Then, the static appearance characteristics, motion optical flow characteristics, and heat map characteristics are gradually fused in a multi-modal manner to obtain the target fusion characteristics. Next, conditional prediction is performed based on the target fusion characteristics to determine the heat map prediction result and offset prediction result of the target center point, as well as the target size prediction result. Finally, offset prediction association is performed based on the prediction results. When the offset prediction association is successful, the target fruit fly motion trajectory information is updated according to the prediction results, thereby realizing the multi-target tracking of fruit flies.
[0199] This method is aimed at the non-linear and random characteristics of Drosophila's movement in a microgravity environment. It adopts an optical flow feature extraction and microgravity movement decoupling scheme. Through multi-scale feature point detection and sub-pixel level positioning optimization, it generates a high-precision optical flow field, captures instantaneous motion vectors such as hovering and sudden acceleration, and provides complementary semantic enhancement for appearance features. Aiming at the problems that Drosophila have high appearance similarity and high-density occlusion, a single modality (such as RGB) is prone to cause identity confusion (IDSW), and directly adding features leads to the superposition of modality noise, this method uses multi-modal feature progressive fusion to enhance the robustness of motion blur region detection. By the temporal consistency constraint of the motion field, a conditional detection heat map is generated subsequently, which can suppress false detections and missed detections in high-density occlusion scenarios and provide an appearance-motion joint representation. In this way, based on the Drosophila multi-object tracking framework of appearance-motion bimodal fusion, this method realizes the decoupling of motion features by extracting optical flow information, and performs multi-modal progressive feature fusion with static appearance features, and can achieve accurate and stable multi-object tracking of Drosophila.
[0200] In some embodiments, to verify the performance and effectiveness of the Drosophila multi-object tracking method based on space science experiment videos provided in the above embodiments of the present application, a verification analysis is carried out.
[0201] Experimental dataset: Representative time periods in the Drosophila life cycle were selected, and a Drosophila multi-object tracking dataset was refined and manually annotated. Drosophila space science experiment video dataset: It contains 20 videos, with a frame rate of 25 frames per second, a total of 2500 frames of images, and a total of 27500 Drosophila instances.
[0202] Evaluation method: The evaluation metrics for the performance of the Drosophila multi-object tracking method in this technical solution include: HOTA (Higher Order Tracking Accuracy), MOTA (Multi-Object Tracking Accuracy), IDF1 (Identity F1 Score), IDs (Number of times of incorrect identity switching). The higher the metric values of HOTA, MOTA, and IDF1, and the lower the metric value of IDs, the better the tracking effect. The specific calculation methods are as follows:
[0203] 。
[0204] 。
[0205] 。
[0206] Among them, represents the detection accuracy rate, represents the association accuracy rate, represents the number of missed detection targets, represents the number of false detection targets, represents the number of times of incorrect identity switching, Indicates the number of target true values, Indicates the proportion of detection boxes that correctly associate target identities, Indicates the proportion of true targets with the same identity that are correctly associated.
[0207] The following respectively use the above-mentioned related technology 1, related technology 2, and the Drosophila multi-object tracking method provided by the embodiments of the present application (hereinafter referred to as: this solution) to conduct a comparison experiment on the accuracy of Drosophila multi-object tracking for the Drosophila space science experiment video dataset. Table 1 shows the results of the Drosophila multi-object tracking accuracy experiment.
[0208] Table 1 Results of Drosophila multi-object tracking accuracy experiment
[0209]
[0210] As shown in Table 1, this solution is superior to related technology 1 and related technology 2 in terms of indicators such as HOTA (83.15), MOTA (89.045), IDF1 (85.174), and IDs (65). The HOTA indicator combines the dual advantages of detection accuracy (DetA) and association robustness (AssA). The score of 83.15 in this solution verifies the synergistic effect of optical flow motion feature decoupling and cross-modal feature fusion. The extremely high value of MOTA (89.045) is based on the targeted design of the experimental scene background stability of the Drosophila space science experiment video in this solution. The significant advantages of IDF1 (85.174) and IDs (65) reflect the key role of motion trajectory decoupling in the identity consistency of target Drosophila. Through the dynamic weight fusion of motion optical flow features and static appearance features, the limitation of relying on a single visual feature in related technologies is broken. Even if the appearances of target Drosophila are highly similar, the uniqueness of the motion trajectory can still ensure the stability of identity association.
[0211] Furthermore, in some embodiments, an ablation experiment analysis was conducted to verify the roles of each step in this solution. By gradually adding steps to introduce motion optical flow features and multi-modal feature progressive fusion steps, the influence of each step on the accuracy of multi-object tracking of target Drosophila was analyzed. Table 2 shows the results of the Drosophila multi-object tracking ablation experiment. Among them, "O" in Table 2 indicates the use of the corresponding step method.
[0212] Table 2 Results of Drosophila multi-object tracking accuracy experiment
[0213]
[0214] As shown in Table 2, after introducing the step of motion optical flow features, HOTA increased from 78.8 to 80.889 (+2.64%), MOTA increased from 82.583 to 85.528 (+3.55%), IDF1 increased from 77.998 to 82.675 (+5.99%), and IDs decreased from 92 to 69 (-25%). This shows that introducing motion optical flow features can effectively model non-linear motions (such as the "floating" and "somersaulting" behaviors of fruit flies) in the microgravity environment, enhance the robustness of trajectory prediction through the motion vector field, and reduce trajectory breaks (FN) and false detections (FP) caused by sudden accelerations. In addition, the spatio-temporal continuity of motion optical flow features supports short-term trajectory extrapolation, maintains the consistency of target identity during occlusion, and significantly reduces the number of IDSW.
[0215] After the superposition of the multi-modal feature progressive fusion step, HOTA increased from 80.889 to 83.15 (+2.8%), MOTA increased from 85.528 to 89.045 (+4.12%), IDF1 increased from 82.675 to 85.174 (+3.03%), and IDs decreased from 69 to 51 (-26%). This shows that multi-modal feature progressive fusion can fuse motion optical flow features and static appearance features in stages through a cross-modal attention mechanism, enhance the positioning ability of blurred targets in the detection stage (heat map generation), suppress the interference of appearance similarity in the association stage (feature matching), automatically adjust the fusion weights of motion optical flow features and static appearance features for different scenarios (such as high-density occlusion, fast motion), and preferentially use the differences in motion trajectories to distinguish targets.
[0216] The embodiment of the present application also provides a fruit fly multi-target tracking system based on space science experiment videos. Specifically, Figure 5 is a schematic structural diagram of the fruit fly multi-target tracking system provided by the embodiment of the present application. As Figure 5 shown, the fruit fly multi-target tracking system 500 based on space science experiment videos includes: an image extraction module 501, an optical flow information extraction module 502, a feature extraction module 503, a multi-modal feature progressive fusion module 504, a conditional prediction module 505, an offset prediction association module 506, and a trajectory generation module 507.
[0217] Among them, the image extraction module 501 can be used to extract the t-th frame image, the (t + 1)-th frame image, and the (t - 1)-th frame image from the fruit fly space science experiment video, where t is a positive integer greater than 1.
[0218] The optical flow information extraction module 502 can be used to extract optical flow information based on the t-th frame image and the (t + 1)-th frame image to determine the first optical flow image corresponding to the t-th frame image; extract optical flow information based on the t-th frame image and the (t - 1)-th frame image to determine the second optical flow image corresponding to the (t - 1)-th frame image.
[0219] The feature extraction module 503 can be used to extract features from the t-th frame image, the first optical flow image, the (t-1)-th frame image, and the second optical flow image respectively, and determine the first static appearance feature, the first motion optical flow feature, the second static appearance feature, the second motion optical flow feature, and the heatmap feature. The heatmap feature is used to characterize the probability distribution of the target center point in the (t-1)-th frame image, and the target center point is used to characterize the center point of the target Drosophila.
[0220] The multi-modal feature progressive fusion module 504 can be used to perform multi-modal feature progressive fusion on the first static appearance feature, the first motion optical flow feature, the second static appearance feature, the second motion optical flow feature, and the heatmap feature to obtain the target fusion feature.
[0221] The conditional prediction module 505 can be used to perform conditional prediction through the target detection network model according to the target fusion feature, and determine the heatmap prediction result, the offset prediction result, and the target size prediction result of the target center point in the t-th frame image.
[0222] The offset prediction association module 506 can be used to perform offset prediction association according to the heatmap prediction result, the offset prediction result, and the heatmap feature.
[0223] The trajectory generation module 507 can be used to update the motion trajectory information of the target Drosophila corresponding to the t-th frame image according to the offset prediction result and the target size prediction result when the offset prediction association is successful.
[0224] In some embodiments, the above multi-modal feature progressive fusion module 504 includes a multi-modal feature progressive fusion model 600. Specifically, Figure 6 is a schematic structural diagram of the multi-modal feature progressive fusion model provided by the embodiments of the present application. As Figure 6 shown, the multi-modal feature progressive fusion model 600 includes: a first fusion module 601, a second fusion module 602, a first feed-forward neural network 603, a second feed-forward neural network 604, a feature addition module 605, a third fusion module 606, a fourth fusion module 607, a fifth fusion module 608, a sixth fusion module 609, a feature splicing module 610, and a third feed-forward neural network 611.
[0225] Among them, the first fusion module 601 can be used to perform first feature fusion on the first static appearance feature and the second static appearance feature based on the cross-attention mechanism to obtain the static appearance fusion feature.
[0226] The second fusion module 602 can be used to perform second feature fusion on the first motion optical flow feature and the second motion optical flow feature based on the cross-attention mechanism to obtain the motion optical flow fusion feature.
[0227] The first feedforward neural network 603 can be used to determine the static appearance temporal feature according to the static appearance fusion feature and the heatmap feature.
[0228] The second feedforward neural network 604 can be used to determine the motion optical flow temporal feature according to the motion optical flow fusion feature and the heatmap feature.
[0229] The feature addition module 605 can be used to add the static appearance temporal feature and the motion optical flow temporal feature to obtain the initial multimodal feature.
[0230] The third fusion module 606 can be used to perform the third feature fusion based on the cross-attention mechanism, using the initial multimodal feature as K and V, and the static appearance temporal feature as Q, to obtain the first multimodal feature.
[0231] The fourth fusion module 607 can be used to perform the fourth feature fusion based on the cross-attention mechanism, using the initial multimodal feature as Q, and the static appearance temporal feature as K and V, to obtain the second multimodal feature.
[0232] The fifth fusion module 608 can be used to perform the fifth feature fusion based on the cross-attention mechanism, using the initial multimodal feature as K and V, and the motion optical flow temporal feature as Q, to obtain the third multimodal feature.
[0233] The sixth fusion module 609 can be used to perform the sixth feature fusion based on the cross-attention mechanism, using the initial multimodal feature as Q, and the motion optical flow temporal feature as K and V, to obtain the fourth multimodal feature.
[0234] The feature splicing module 610 can be used to splice the first multimodal feature, the second multimodal feature, the third multimodal feature, and the fourth multimodal feature to obtain the multimodal splicing feature.
[0235] The third feedforward neural network 611 can be used to determine the target fusion feature according to the multimodal splicing feature.
[0236] Using the Drosophila multi-object tracking system based on space science experiment videos provided by the embodiments of the present application, first, the microgravity motion characteristics are decoupled by extracting the optical flow information of the frame images in the Drosophila space science experiment videos. Then, the static appearance features, motion optical flow features, and heatmap features are progressively fused in a multi-modal manner to obtain the target fusion features. Next, conditional prediction is performed based on the target fusion features to determine the heatmap prediction result, offset prediction result, and target size prediction result of the target center point. Finally, offset prediction association is performed based on the prediction results. When the offset prediction association is successful, the motion trajectory information of the target Drosophila is updated according to the prediction results, thereby realizing the multi-object tracking of Drosophila. In this way, based on the appearance-motion dual-modal fusion Drosophila multi-object tracking framework, the system can achieve accurate and stable multi-object tracking of Drosophila by extracting optical flow information to decouple motion features and performing multi-modal progressive feature fusion with static appearance features.
[0237] An embodiment of the present invention further provides an electronic device, which may include: a display screen, a memory, and one or more processors. The display screen, memory, and processor are coupled. The memory is used to store computer program code, and the computer program code includes computer instructions. When the processor executes the computer instructions, the electronic device can execute each method or step performed in the above-mentioned embodiments of the Drosophila multi-object tracking method. Of course, the electronic device includes, but is not limited to, the above-mentioned display screen, memory, and one or more processors.
[0238] An embodiment of the present invention further provides a computer-readable storage medium for storing computer instructions for running the above-mentioned Drosophila multi-object tracking method.
[0239] Through the description of the above embodiments, those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above-mentioned division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0240] In the description of the present invention, it should be understood that the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined.
[0241] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0242] For the similar parts between the embodiments provided in this application, reference can be made to each other. The specific embodiments provided above are only several examples under the general concept of this application and do not constitute a limitation on the protection scope of this application. For those skilled in the art, any other embodiments extended based on the solution of this application without creative efforts belong to the protection scope of this application.
Claims
1. A Drosophila multi-object tracking method based on space science experiment videos, characterized in that, Including: Extracting the t-th frame image, the (t + 1)-th frame image, and the (t - 1)-th frame image from the Drosophila spatial science experiment video, where t is a positive integer greater than 1; Performing optical flow information extraction based on the t-th frame image and the (t + 1)-th frame image to determine the first optical flow image corresponding to the t-th frame image; performing optical flow information extraction based on the t-th frame image and the (t - 1)-th frame image to determine the second optical flow image corresponding to the (t - 1)-th frame image; Performing feature extraction on the t-th frame image, the first optical flow image, the (t - 1)-th frame image, and the second optical flow image respectively to determine the first static appearance feature, the first motion optical flow feature, the second static appearance feature, the second motion optical flow feature, and the heatmap feature, where the heatmap feature is used to characterize the probability distribution of the target center point in the (t - 1)-th frame image, and the target center point is used to characterize the center point of the target Drosophila; Performing multi-modal feature progressive fusion on the first static appearance feature, the first motion optical flow feature, the second static appearance feature, the second motion optical flow feature, and the heatmap feature to obtain the target fusion feature; wherein, the multi-modal feature progressive fusion specifically includes: determining the static appearance fusion feature based on the first static appearance feature and the second static appearance feature according to the cross-attention mechanism, determining the motion optical flow fusion feature for the first motion optical flow feature and the second motion optical flow feature based on the cross-attention mechanism; according to the heatmap feature, the static appearance fusion feature, and the motion optical flow fusion feature, determining the static appearance time feature and the motion optical flow time feature through a feed-forward neural network; determining the initial multi-modal feature by adding the static appearance time feature and the motion optical flow time feature; performing feature fusion based on the initial multi-modal feature, the static appearance time feature, and the motion optical flow time feature through the cross-attention mechanism to determine the target fusion feature; Performing conditional prediction through a target detection network model according to the target fusion feature to determine the heatmap prediction result, the offset prediction result, and the target size prediction result of the target center point in the t-th frame image; Performing offset prediction association according to the heatmap prediction result, the offset prediction result, and the heatmap feature; In the case where the offset prediction association is successful, updating the motion trajectory information of the target Drosophila corresponding to the t-th frame image according to the offset prediction result and the target size prediction result.
2. The method according to claim 1, characterized in that, The performing optical flow information extraction based on the t-th frame image and the (t + 1)-th frame image to determine the first optical flow image corresponding to the t-th frame image includes: Performing preprocessing on the t-th frame image and the (t + 1)-th frame image respectively to obtain the corresponding first preprocessed image and second preprocessed image; Performing image subtraction processing based on the first preprocessed image and the second preprocessed image to determine the first difference image; Determining significant feature points according to the first difference image through a corner detection algorithm; Determine a first optical flow vector corresponding to a pixel point in the t-th frame image according to the first preprocessed image, the second preprocessed image, and the significant feature points; Generate the first optical flow image according to the first optical flow vector corresponding to the pixel point in the t-th frame image.
3. The method according to claim 2, wherein The determining of the first optical flow vector corresponding to the pixel point in the t-th frame image according to the first preprocessed image, the second preprocessed image, and the significant feature points includes: Determine the displacement vector of the significant feature points by minimizing the optical flow error function, linearizing by first-order Taylor expansion, and solving by the least squares method according to the first preprocessed image and the second preprocessed image; Determine the displacement vector corresponding to the pixel point in the t-th frame image by bilinear interpolation according to the displacement vector of the significant feature points; Determine the superimposed optical flow vector by superimposing the optical flow fields through constructing an image pyramid with a preset number of layers according to the displacement vector corresponding to the pixel point in the t-th frame image; Perform median filtering on the superimposed optical flow vector to determine the filtered optical flow vector; Perform filtering processing on the filtered optical flow vector according to a preset speed amplitude threshold to obtain the first optical flow vector.
4. The method according to any one of claims 1 to 3, characterized in that, The multi-modal feature progressive fusion of the first static appearance feature, the first motion optical flow feature, the second static appearance feature, the second motion optical flow feature, and the heat map feature to obtain the target fusion feature includes: Perform first feature fusion on the first static appearance feature and the second static appearance feature based on the cross-attention mechanism to obtain the static appearance fusion feature; Perform second feature fusion on the first motion optical flow feature and the second motion optical flow feature based on the cross-attention mechanism to obtain the motion optical flow fusion feature; Determine the static appearance time feature through a first feed-forward neural network according to the static appearance fusion feature and the heat map feature; Determine the motion optical flow time feature through a second feed-forward neural network according to the motion optical flow fusion feature and the heat map feature; Add the static appearance time feature and the motion optical flow time feature to obtain the initial multi-modal feature; Perform third feature fusion with the initial multi-modal feature as the key K and value V and the static appearance time feature as the query Q based on the cross-attention mechanism to obtain the first multi-modal feature; Perform fourth feature fusion with the initial multi-modal feature as Q, the static appearance time feature as K and V based on the cross-attention mechanism to obtain the second multi-modal feature; Perform fifth feature fusion with the initial multi-modal feature as K and V, and the motion optical flow time feature as Q based on the cross-attention mechanism to obtain the third multi-modal feature; Perform sixth feature fusion with the initial multi-modal feature as Q, the motion optical flow time feature as K and V based on the cross-attention mechanism to obtain the fourth multi-modal feature; Perform feature splicing on the first multi-modal feature, the second multi-modal feature, the third multi-modal feature, and the fourth multi-modal feature to obtain the multi-modal splicing feature; The target fusion feature is determined through a third feedforward neural network according to the multi-modal splicing feature.
5. The method according to claim 1, characterized in that, The expression of the heatmap prediction result is: ; Among them, represents the predicted result of the heat map, represents the Sigmoid activation function, represents the heat map output operation of the convolutional network, represents the target fusion feature; The expression of the target size prediction result is: ; Among them, represents the predicted result of the target size, represents the rectified linear unit function, represents the size output operation of the convolutional network; The expression of the offset prediction result is: ; Among them, represents the offset prediction result, represents the offset output operation of the convolutional network.
6. The method according to claim 5, wherein The loss expression of the target detection network model is: ; ; ; ; Among them, represents the loss of the target detection network model; represents the prediction loss of the heatmap, represents the prediction loss of the target size, represents the prediction loss of the offset; represents the number of positive samples in the training set, represents the ground truth of the heatmap, the predicted value of the heatmap, represents the weight coefficient of the positive sample, represents the weight coefficient of the negative sample in the training set; represents the predicted value of the width in the target size, represents the ground truth of the width in the target size, represents the predicted value of the height in the target size, represents the ground truth of the height in the target size; represents the predicted value of the width offset, represents the ground truth of the width offset, represents the predicted value of the height offset, represents the ground truth of the height offset.
7. The method according to claim 1, wherein The offset prediction association according to the heatmap prediction result, the offset prediction result and the heatmap feature includes: Performing local extreme point extraction on the heatmap prediction result to determine the detection point corresponding to the t-th frame image; Determining the backtracking position corresponding to the detection point according to the offset prediction result, and the expression of the backtracking position is: ; Among them, represents the backtracking position, represents the abscissa of the detection point in the t-th frame image, represents the predicted offset value of the abscissa and ordinate of the detection point, represents the ordinate of the detection point in the t-th frame image; Determining the heatmap response corresponding to the backtracking position according to the heatmap feature, and determining whether the heatmap response corresponding to the backtracking position meets a preset response threshold.
8. The method according to claim 1, characterized in that The target fruit fly movement trajectory information includes: the abscissa of the center point, the ordinate of the center point, the width, the height and the trajectory number corresponding to the target fruit fly; In the case where the offset prediction association is successful, updating the target fruit fly movement trajectory information corresponding to the t-th frame image according to the offset prediction result and the target size prediction result includes: Superimposing the offset prediction result, the target size prediction result and the target fruit fly movement trajectory information corresponding to the (t - 1)-th frame image to obtain the target fruit fly movement trajectory information corresponding to the t-th frame image; the expression of the target fruit fly movement trajectory information corresponding to the t-th frame image is: ; = + ; = + ; Among them, represents the th target Drosophila motion trajectory information corresponding to the t-th frame image, represents the th target Drosophila motion trajectory information corresponding to the (t - 1)-th frame image, represents the abscissa of the center point of the th target Drosophila, represents the ordinate of the center point of the th target Drosophila, represents the width prediction value in the target size prediction result of the th target Drosophila, represents the motion trajectory number of the th target Drosophila, represents the abscissa of the center point of the th target Drosophila in the (t - 1)-th frame image, ordinate of the center point of the th target Drosophila in the (t - 1)-th frame image, represents the vertical offset prediction value in the offset prediction result of the th target Drosophila.
9. A Drosophila multi-object tracking system based on space science experiment videos, characterized in that, Including: An image extraction module, an optical flow information extraction module, a feature extraction module, a multi-modal feature progressive fusion module, a conditional prediction module, an offset prediction association module and a trajectory generation module; wherein, The image extraction module is configured to extract the t-th frame image, the (t + 1)-th frame image and the (t - 1)-th frame image from the fruit fly space science experiment video, where t is a positive integer greater than 1; The optical flow information extraction module is configured to extract optical flow information according to the t-th frame image and the (t + 1)-th frame image to determine the first optical flow image corresponding to the t-th frame image; extracting optical flow information according to the t-th frame image and the (t - 1)-th frame image to determine the second optical flow image corresponding to the (t - 1)-th frame image; The feature extraction module is configured to perform feature extraction on the t-th frame image, the first optical flow image, the (t - 1)-th frame image and the second optical flow image respectively to determine a first static appearance feature, a first motion optical flow feature, a second static appearance feature, a second motion optical flow feature and a heatmap feature, where the heatmap feature is used to characterize the probability distribution of the target center point in the (t - 1)-th frame image, and the target center point is used to characterize the center point of the target fruit fly; The multimodal feature progressive fusion module is used to perform multimodal feature progressive fusion on the first static appearance feature, the first motion optical flow feature, the second static appearance feature, the second motion optical flow feature, and the heat map feature to obtain a target fusion feature; wherein, the multimodal feature progressive fusion specifically includes: determining a static appearance fusion feature based on the cross-attention mechanism according to the first static appearance feature and the second static appearance feature, and determining a motion optical flow fusion feature for the first motion optical flow feature and the second motion optical flow feature based on the cross-attention mechanism; according to the heat map feature, the static appearance fusion feature, and the motion optical flow fusion feature, determining a static appearance temporal feature and a motion optical flow temporal feature through a feed-forward neural network; according to the static appearance temporal feature and the motion optical flow temporal feature, determining an initial multimodal feature by feature addition; based on the initial multimodal feature, the static appearance temporal feature, and the motion optical flow temporal feature, performing feature fusion through the cross-attention mechanism to determine the target fusion feature; The conditional prediction module is used to perform conditional prediction through a target detection network model according to the target fusion feature to determine the heat map prediction result, the offset prediction result, and the target size prediction result of the target center point in the t-th frame image; The offset prediction association module is used to perform offset prediction association according to the heat map prediction result, the offset prediction result, and the heat map feature; The trajectory generation module is used to update the target fruit fly motion trajectory information corresponding to the t-th frame image according to the offset prediction result and the target size prediction result when the offset prediction association is successful.
10. The system according to claim 9, wherein The multimodal feature progressive fusion module includes a multimodal feature progressive fusion model; the multimodal feature progressive fusion model includes: a first fusion module, a second fusion module, a first feed-forward neural network, a second feed-forward neural network, a feature addition module, a third fusion module, a fourth fusion module, a fifth fusion module, a sixth fusion module, a feature splicing module, and a third feed-forward neural network; wherein, The first fusion module is used to perform first feature fusion on the first static appearance feature and the second static appearance feature based on the cross-attention mechanism to obtain a static appearance fusion feature; The second fusion module is used to perform second feature fusion on the first motion optical flow feature and the second motion optical flow feature based on the cross-attention mechanism to obtain a motion optical flow fusion feature; The first feed-forward neural network is used to determine a static appearance temporal feature according to the static appearance fusion feature and the heat map feature; The second feed-forward neural network is used to determine a motion optical flow temporal feature according to the motion optical flow fusion feature and the heat map feature; The feature addition module is used to add the static appearance temporal feature and the motion optical flow temporal feature to obtain an initial multimodal feature; The third fusion module is used to perform third feature fusion based on the cross-attention mechanism, using the initial multi-modal features as K and V, and the static appearance temporal features as Q, to obtain the first multi-modal feature; The fourth fusion module is used to perform fourth feature fusion based on the cross-attention mechanism, using the initial multi-modal features as Q, and the static appearance temporal features as K and V, to obtain the second multi-modal feature; The fifth fusion module is used to perform fifth feature fusion based on the cross-attention mechanism, using the initial multi-modal features as K and V, and the motion optical flow temporal features as Q, to obtain the third multi-modal feature; The sixth fusion module is used to perform sixth feature fusion based on the cross-attention mechanism, using the initial multi-modal features as Q, and the motion optical flow temporal features as K and V, to obtain the fourth multi-modal feature; The feature concatenation module is used to perform feature concatenation on the first multi-modal feature, the second multi-modal feature, the third multi-modal feature, and the fourth multi-modal feature to obtain the multi-modal concatenated feature; The third feed-forward neural network is used to determine the target fusion feature according to the multi-modal concatenated feature.
Citation Information
Patent Citations
Multi-target trajectory anomaly processing method and system in micro-manipulation
CN112785630A
Video small target tracking method and device
CN113269808A