Fruit fly multi-target tracking method and system based on space science experiment video

By using optical flow information to decouple microgravity motion in the fruit fly multi-object tracking method, and combining static appearance and moving optical flow characteristics for multi-modal feature fusion, the stability and accuracy problems of fruit fly multi-object tracking in the space station are solved, and accurate tracking of fruit fly is achieved.

CN120070510AActive Publication Date: 2025-05-30TECH & ENG CENT FOR SPACE UTILIZATION CHINESE ACAD OF SCI

Patent Information

Application Number
CN202510541648.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-05-30
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

The fruit fly observation experiments conducted in the space station are unable to achieve multi-target tracking of fruit fly flies stably and accurately due to nonlinear motion in microgravity environments, high density and frequent occlusion of fruit fly targets, and high appearance similarity.

Method used

The multi-objective tracking method of fruit fly based on space science experimental videos is adopted to achieve decoupling of microgravity motion characteristics by extracting optical flow information of frame images in the space science experimental videos of fruit fly. Then, the static appearance features, moving optical flow features and thermal map features are gradually fused to obtain the target fusion features. Next, conditional prediction is performed based on the target fusion characteristics, and the thermal map prediction results and offset prediction results of the target center point are determined. Finally, the offset prediction association is performed based on the prediction results, and when the offset prediction association is successful, the target fruit fly movement trajectory information is updated.

Benefits of technology

Accurate and stable multi-objective tracking of fruit flies is achieved, and can effectively deal with microgravity nonlinear motion and high-density occlusion scenarios, meeting the scientific needs of space station fruit flies behavior analysis for trajectory accuracy and long-term identity consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070510A_ABST
    Figure CN120070510A_ABST
Patent Text Reader

Abstract

The invention provides a fruit fly multi-target tracking method and system based on a space science experiment video, and the method comprises the steps: firstly, extracting the optical flow information of a frame image in the fruit fly space science experiment video, and achieving the microgravity motion decoupling; then, multi-modal feature progressive fusion is carried out on the static appearance features, the moving optical flow features and the thermodynamic diagram features; next, conditional prediction is carried out according to the fused target fusion features, and a thermodynamic diagram prediction result, an offset prediction result and a target size prediction result are determined; and finally, performing offset prediction association according to the prediction result, updating the target drosophila movement track information according to the prediction result under the condition that the offset prediction association is successful, and determining a multi-target tracking result. According to the method, based on an appearance-motion dual-mode fusion multi-target tracking framework, motion feature decoupling is realized by extracting optical flow information, and multi-mode progressive feature fusion with static appearance features is carried out, so that accurate and stable multi-target tracking of fruit flies is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target tracking, and specifically relates to a method and system for multi-target tracking of Drosophila based on space science experiment videos. Background Art

[0002] With the rapid development of space station technology, large-scale multidisciplinary space science research can be carried out in space stations. For example, Drosophila observation experiments can be conducted through the life ecology experiment cabinet carried on the space station laboratory. Specifically, according to the obtained Drosophila space science experiment videos, multi-target tracking technology can be used to locate multiple Drosophila targets, maintain their identity labels, and generate the respective movement trajectories of the Drosophila targets, thereby providing reliable technical support for in-depth analysis of the behavioral characteristics and gene expression laws of organisms under space microgravity and submagnetic environments.

[0003] Generally, multi-target tracking technology can adopt a tracking method of detecting first and then associating, and a method of jointly detecting and tracking. Among them, the tracking method of detecting first and then associating decouples the tracking task into two stages of target detection and data association, and improves the overall tracking performance by independently optimizing the detection accuracy and association strategy. For example, the target center can be located through the detection box, and on this basis, the Kalman filter can be used to predict the linear movement trajectory. The method of jointly detecting and tracking converts the detector into a tracker and combines the two tasks of detection and tracking in the same detection framework. For example, low-dimensional re-identification features can be extracted by sharing the backbone network.

[0004] However, the Drosophila observation experiment carried out in the space station has the following characteristics: the movement under microgravity environment is nonlinear, random, and highly dynamic; the Drosophila targets have high density and frequent occlusion; the appearance similarity of the Drosophila targets is high; and there is a stable scene under a fixed field of view. For the Drosophila space science experiment videos obtained in the space station, the above-mentioned related technologies of multi-target tracking do not model the microgravity nonlinear movement, ignore the significant differences between the characteristics of Drosophila targets and the data distribution of Drosophila space science experiment videos, resulting in the deviation of the multi-target detection center, leading to feature ambiguity, identity switching in frequently occluded scenes, and the failure of nonlinear trajectory prediction. As a result, it is impossible to stably and accurately achieve multi-target tracking of Drosophila, and it is difficult to meet the scientific requirements of trajectory accuracy and long-term identity consistency for Drosophila behavior analysis in the space station. Summary of the Invention

[0005] The technical problem to be solved by the present invention is the problem that it is impossible to stably and accurately achieve multi-target tracking of Drosophila.

[0006] To solve the above technical problem, the present invention provides a method and system for multi-target tracking of Drosophila based on space science experiment videos, and specifically adopts the following technical solutions: First aspect, the present invention provides a Drosophila multi-object tracking method based on space science experiment videos, including: First, extract the t-th frame image, the (t + 1)-th frame image, and the (t - 1)-th frame image from the Drosophila space science experiment video, where t is a positive integer greater than 1. Then, extract optical flow information based on the t-th frame image and the (t + 1)-th frame image to determine the first optical flow image corresponding to the t-th frame image; extract optical flow information based on the t-th frame image and the (t - 1)-th frame image to determine the second optical flow image corresponding to the (t - 1)-th frame image. Next, perform feature extraction on the t-th frame image, the first optical flow image, the (t - 1)-th frame image, and the second optical flow image respectively to determine the first static appearance feature, the first moving optical flow feature, the second static appearance feature, the second moving optical flow feature, and the heatmap feature. The heatmap feature is used to characterize the probability distribution of the target center point in the (t - 1)-th frame image, and the target center point is used to characterize the center point of the target Drosophila. Further, perform multi-modal feature progressive fusion on the first static appearance feature, the first moving optical flow feature, the second static appearance feature, the second moving optical flow feature, and the heatmap feature to obtain the target fusion feature. Secondly, perform conditional prediction through the target detection network model according to the target fusion feature to determine the heatmap prediction result and the offset prediction result of the target center point in the t-th frame image, as well as the target size prediction result. Finally, perform offset prediction association based on the heatmap prediction result, the offset prediction result, and the heatmap feature. In the case where the offset prediction association is successful, update the motion trajectory information of the target Drosophila corresponding to the t-th frame image according to the offset prediction result and the target size prediction result.

[0007] This method first realizes the decoupling of microgravity motion characteristics by extracting the optical flow information of the frame images in the Drosophila space science experiment video. Then, perform multi-modal feature progressive fusion on the static appearance feature, the moving optical flow feature, and the heatmap feature to obtain the target fusion feature. Next, perform conditional prediction according to the target fusion feature to determine the heatmap prediction result and the offset prediction result of the target center point, as well as the target size prediction result. Finally, perform offset prediction association according to the prediction results. In the case where the offset prediction association is successful, update the motion trajectory information of the target Drosophila according to the prediction results, so as to realize accurate multi-object tracking of Drosophila. This method is based on an appearance-motion dual-modal fusion Drosophila multi-object tracking framework. By extracting optical flow information, it realizes the decoupling of motion characteristics and performs multi-modal progressive feature fusion with the static appearance feature, which can achieve accurate and stable multi-object tracking of Drosophila.

[0008] Combined with the first aspect, in an alternative implementation, the above-mentioned extraction of optical flow information based on the t-th frame image and the (t + 1)-th frame image to determine the first optical flow image corresponding to the t-th frame image includes: First, preprocess the t-th frame image and the (t + 1)-th frame image respectively to obtain the corresponding first preprocessed image and second preprocessed image. Then, perform an image subtraction process based on the first preprocessed image and the second preprocessed image to determine the first difference image. Next, according to the first difference image, determine significant feature points through a corner detection algorithm. According to the first preprocessed image, the second preprocessed image, and the significant feature points, determine the first optical flow vector corresponding to the pixel points in the t-th frame image. Finally, generate the first optical flow image based on the first optical flow vector corresponding to the pixel points in the t-th frame image.

[0009] In this implementation, by preprocessing the t-th frame image and the (t + 1)-th frame image, performing an image subtraction process, and determining the first optical flow vector, finally, the first optical flow image corresponding to the t-th frame image can be accurately and effectively determined based on the first optical flow vector.

[0010] Combined with the first aspect, in an alternative implementation, the above-mentioned determination of the first optical flow vector corresponding to the pixel points in the t-th frame image according to the first preprocessed image, the second preprocessed image, and the significant feature points includes: First, according to the first preprocessed image and the second preprocessed image, determine the displacement vector of the significant feature points by minimizing the optical flow error function, linearizing by first-order Taylor expansion, and solving by the least squares method. Then, according to the displacement vector of the significant feature points, determine the displacement vector corresponding to the pixel points in the t-th frame image through bilinear interpolation. Next, according to the displacement vector corresponding to the pixel points in the t-th frame image, perform optical flow field superposition by constructing an image pyramid with a preset number of layers to determine the superimposed optical flow vector. Perform median filtering on the superimposed optical flow vector to determine the filtered optical flow vector. Finally, perform filtering processing on the filtered optical flow vector according to a preset speed amplitude threshold to obtain the first optical flow vector.

[0011] In this implementation, as an implementation of determining the first optical flow vector corresponding to the pixel points in the t-th frame image, the displacement vector of the significant feature points is determined by constructing and solving the optical flow error function. Then, through bilinear interpolation, optical flow field superposition by constructing an image pyramid, median filtering, and filtering processing in sequence, the first optical flow vector can be accurately obtained.

[0012] In combination with the first aspect, in an alternative implementation, the above multi-modal feature progressive fusion of the first static appearance feature, the first motion optical flow feature, the second static appearance feature, the second motion optical flow feature, and the heat map feature to obtain the target fusion feature specifically includes: performing a first feature fusion on the first static appearance feature and the second static appearance feature based on the cross-attention mechanism to obtain a static appearance fusion feature; performing a second feature fusion on the first motion optical flow feature and the second motion optical flow feature based on the cross-attention mechanism to obtain a motion optical flow fusion feature; determining the static appearance time feature through a first feed-forward neural network according to the static appearance fusion feature and the heat map feature; determining the motion optical flow time feature through a second feed-forward neural network according to the motion optical flow fusion feature and the heat map feature; adding the static appearance time feature and the motion optical flow time feature to obtain an initial multi-modal feature; performing a third feature fusion based on the cross-attention mechanism with the initial multi-modal feature as the key K and value V, and the static appearance time feature as the query Q to obtain a first multi-modal feature; performing a fourth feature fusion based on the cross-attention mechanism with the initial multi-modal feature as the Q, and the static appearance time feature as the K and V to obtain a second multi-modal feature; performing a fifth feature fusion based on the cross-attention mechanism with the initial multi-modal feature as the K and V, and the motion optical flow time feature as the query Q to obtain a third multi-modal feature; performing a sixth feature fusion based on the cross-attention mechanism with the initial multi-modal feature as the Q, and the motion optical flow time feature as the K and V to obtain a fourth multi-modal feature; performing feature concatenation on the first multi-modal feature, the second multi-modal feature, the third multi-modal feature, and the fourth multi-modal feature to obtain a multi-modal concatenated feature; and determining the target fusion feature through a third feed-forward neural network according to the multi-modal concatenated feature.

[0013] In combination with the first aspect, in an alternative implementation, the expression of the above heat map prediction result is: 。

[0014] Wherein, represents the heat map prediction result, represents the Sigmoid activation function, represents the heat map output operation of the convolutional network, represents the target fusion feature. The expression of the target size prediction result is: 。

[0015] Wherein, represents the target size prediction result, represents the rectified linear unit function, represents the size output operation of the convolutional network. The expression of the offset prediction result is: 。

[0016] Among them, represents the offset prediction result, and represents the offset output operation of the convolutional network.

[0017] Combined with the first aspect, in an alternative implementation, the loss expression of the above object detection network model is: .

[0018] .

[0019] .

[0020] .

[0021] Among them, represents the loss of the object detection network model; represents the prediction loss of the heatmap, represents the prediction loss of the object size, represents the prediction loss of the offset; represents the number of positive samples in the training set, represents the ground truth of the heatmap, the predicted value of the heatmap, represents the weight coefficient of the positive sample, represents the weight coefficient of the negative sample in the training set; represents the predicted value of the width in the object size, represents the ground truth of the width in the object size, represents the predicted value of the height in the object size, represents the ground truth of the height in the object size; represents the predicted value of the width offset, represents the ground truth of the width offset, represents the predicted value of the height offset, represents the ground truth of the height offset.

[0022] Combined with the first aspect, in an alternative implementation, the above offset prediction association based on the heatmap prediction result, offset prediction result and heatmap feature includes: First, extract the local extreme points of the heatmap prediction result to determine the detection points corresponding to the t-th frame image. Then, determine the backtracking position corresponding to the detection points according to the offset prediction result, and the expression of the backtracking position is: .

[0023] Among them, represents the backtracking position, represents the abscissa of the detection point in the t-th frame image, The predicted offset values of the abscissa and ordinate of the detection point represents the ordinate of the detection point in the t-th frame image. Finally, according to the heatmap features, determine the heatmap response corresponding to the backtracking position, and determine whether the heatmap response corresponding to the backtracking position meets the preset response threshold.

[0024] Combined with the first aspect, in an alternative implementation, the target fruit fly motion trajectory information includes: the abscissa of the center point corresponding to the target fruit fly, the ordinate of the center point, the width, the height, and the trajectory number. In the case where the offset prediction is successfully associated, update the target fruit fly motion trajectory information corresponding to the t-th frame image according to the offset prediction result and the target size prediction result, including: superimposing the offset prediction result, the target size prediction result with the target fruit fly motion trajectory information corresponding to the (t - 1)-th frame image to obtain the target fruit fly motion trajectory information corresponding to the t-th frame image; the expression of the target fruit fly motion trajectory information corresponding to the t-th frame image is: .

[0025] = + .

[0026] = + .

[0027] Among them, represents the th target fruit fly motion trajectory information corresponding to the t-th frame image, represents the th target fruit fly motion trajectory information corresponding to the (t - 1)-th frame image, represents the abscissa of the center point of the th target fruit fly, represents the ordinate of the center point of the th target fruit fly, represents the predicted width value in the target size prediction result of the th target fruit fly, represents the predicted height value in the target size prediction result of the th target fruit fly, represents the trajectory number of the th target fruit fly, represents the abscissa of the center point of the th target fruit fly in the (t - 1)-th frame image, represents the predicted abscissa offset value in the offset prediction result of the th target fruit fly, represents the The vertical coordinate of the center point of a target fruit fly in the (t-1)-th frame image represents the predicted value of the vertical coordinate offset in the offset prediction result of the

[0028] In a second aspect, the present invention provides a multi-target tracking system for fruit flies based on a space science experiment video. The system includes: an image extraction module, an optical flow information extraction module, a feature extraction module, a multi-modal feature progressive fusion module, a conditional prediction module, an offset prediction association module, and a trajectory generation module. Among them, the image extraction module can be used to extract the t-th frame image, the (t + 1)-th frame image, and the (t-1)-th frame image from the fruit fly space science experiment video, where t is a positive integer greater than 1. The optical flow information extraction module can be used to extract optical flow information based on the t-th frame image and the (t + 1)-th frame image to determine the first optical flow image corresponding to the t-th frame image; extract optical flow information based on the t-th frame image and the (t-1)-th frame image to determine the second optical flow image corresponding to the (t-1)-th frame image. The feature extraction module can be used to extract features from the t-th frame image, the first optical flow image, the (t-1)-th frame image, and the second optical flow image respectively to determine the first static appearance feature, the first motion optical flow feature, the second static appearance feature, the second motion optical flow feature, and the heat map feature. The heat map feature is used to characterize the probability distribution of the target center point in the (t-1)-th frame image, and the target center point is used to characterize the center point of the target fruit fly. The multi-modal feature progressive fusion module can be used to perform multi-modal feature progressive fusion on the first static appearance feature, the first motion optical flow feature, the second static appearance feature, the second motion optical flow feature, and the heat map feature to obtain the target fusion feature. The conditional prediction module can be used to perform conditional prediction through a target detection network model based on the target fusion feature to determine the heat map prediction result, the offset prediction result, and the target size prediction result of the target center point in the t-th frame image. The offset prediction association module can be used to perform offset prediction association based on the heat map prediction result, the offset prediction result, and the heat map feature. The trajectory generation module can be used to update the target fruit fly motion trajectory information corresponding to the t-th frame image according to the offset prediction result and the target size prediction result when the offset prediction association is successful.

[0029] Combined with the second aspect, in an alternative implementation, the above-mentioned multi-modal feature progressive fusion module includes a multi-modal feature progressive fusion model. The multi-modal feature progressive fusion model specifically includes: a first fusion module, a second fusion module, a first feed-forward neural network, a second feed-forward neural network, a feature addition module, a third fusion module, a fourth fusion module, a fifth fusion module, a sixth fusion module, a feature splicing module, and a third feed-forward neural network. Among them, the first fusion module is used to perform a first feature fusion on the first static appearance feature and the second static appearance feature based on the cross-attention mechanism to obtain a static appearance fusion feature. The second fusion module is used to perform a second feature fusion on the first motion optical flow feature and the second motion optical flow feature based on the cross-attention mechanism to obtain a motion optical flow fusion feature. The first feed-forward neural network is used to determine the static appearance time feature according to the static appearance fusion feature and the heat map feature. The second feed-forward neural network is used to determine the motion optical flow time feature according to the motion optical flow fusion feature and the heat map feature. The feature addition module is used to add the static appearance time feature and the motion optical flow time feature to obtain an initial multi-modal feature. The third fusion module is used to perform a third feature fusion based on the cross-attention mechanism with the initial multi-modal feature as K and V, and the static appearance time feature as Q to obtain a first multi-modal feature. The fourth fusion module is used to perform a fourth feature fusion based on the cross-attention mechanism with the initial multi-modal feature as Q, and the static appearance time feature as K and V to obtain a second multi-modal feature. The fifth fusion module is used to perform a fifth feature fusion based on the cross-attention mechanism with the initial multi-modal feature as K and V, and the motion optical flow time feature as Q to obtain a third multi-modal feature. The sixth fusion module is used to perform a sixth feature fusion based on the cross-attention mechanism with the initial multi-modal feature as Q, and the motion optical flow time feature as K and V to obtain a fourth multi-modal feature. The feature splicing module is used to splice the first multi-modal feature, the second multi-modal feature, the third multi-modal feature, and the fourth multi-modal feature to obtain a multi-modal splicing feature. The third feed-forward neural network is used to determine the target fusion feature according to the multi-modal splicing feature.

[0030] In a third aspect, the present invention provides an electronic device, including: a memory, one or more processors; the memory is coupled to the processor; wherein, the memory stores computer program code, and the computer program code includes computer instructions, when the computer instructions are executed by the processor, the electronic device is caused to execute the method provided in the first aspect and any of its alternative implementations as described above.

[0031] In a fourth aspect, the present invention provides a computer-readable storage medium, including computer instructions, when the computer instructions are run on an electronic device, the electronic device is caused to execute the method provided in the first aspect and any of its alternative implementations as described above.

[0032] Understandably, for the beneficial effects that can be achieved by the Drosophila multi-object tracking system provided in the second aspect above, the electronic device in the third aspect, and the computer-readable storage medium in the fourth aspect, reference can be made to the beneficial effects in the first aspect and any of its possible design manners, which will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 It is a schematic diagram of the principle of the Drosophila multi-object tracking method based on the space science experiment video provided by an embodiment of the present application; Figure 2 It is a schematic flowchart of the Drosophila multi-object tracking method based on the space science experiment video provided by an embodiment of the present application; Figure 3 It is a schematic flowchart of the method for determining the first optical flow image provided by an embodiment of the present application; Figure 4 It is a schematic diagram of the principle of progressive fusion of multi-modal features provided by an embodiment of the present application; Figure 5 It is a schematic structural diagram of the Drosophila multi-object tracking system based on the space science experiment video provided by an embodiment of the present application; Figure 6 It is a schematic structural diagram of the multi-modal feature progressive fusion model provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0034] The embodiments will be described in detail below, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following examples do not represent all embodiments consistent with the present application. They are only examples of systems and methods consistent with some aspects of the present application detailed in the claims.

[0035] With the rapid development of space station technology, large-scale multi-disciplinary space science research can be carried out in the space station. For example, through the life ecology experiment cabinet carried in the space station laboratory, Drosophila observation experiments can be conducted. Specifically, according to the obtained Drosophila space science experiment video, multi-object tracking technology can be used to locate multiple Drosophila targets, maintain their identity identifiers, and generate the respective movement trajectories of the Drosophila targets, thereby providing reliable technical support for in-depth analysis of the behavioral characteristics and gene expression laws of organisms under space microgravity and sub-magnetic environments.

[0036] Generally, multi-object tracking techniques can adopt the tracking method of detecting first and then associating, and the method of jointly detecting and tracking. Among them, the tracking method of detecting first and then associating decouples the tracking task into two stages: object detection and data association, and improves the overall tracking performance by independently optimizing the detection accuracy and association strategy. For example, the center of the target can be located by the detection box, and on this basis, the Kalman filter is used to predict the linear motion trajectory. The method of jointly detecting and tracking converts the detector into a tracker and combines the two tasks of detection and tracking in the same detection framework. For example, low-dimensional re-identification features can be extracted by sharing the backbone network.

[0037] Exemplarily, related technology one proposed a Multiple Object Tracking and Detection (MOTDT) method based on the framework of detecting first and then associating. By combining the detection and tracking results into candidates and selecting the optimal candidate based on a deep neural network, it can cope with the interference of unreliable detections in tracking. Further design a hierarchical data association strategy to improve the tracking performance. This method specifically includes: real-time object classification, trajectory confidence scoring, appearance feature representation, and hierarchical data association based on the Drosophila space science experiment video. Finally, the position trajectory is obtained through the tracking results.

[0038] Related technology two proposed the FairMOT method based on the framework of jointly detecting and tracking, which transforms multi-object tracking into a pixel-level key point estimation and identity classification problem on a high-resolution feature map through a single-stage end-to-end architecture and feature alignment design. This method specifically includes: extracting high-resolution features based on the Drosophila space science experiment video, and then performing detection and re-identification through the detection branch and the re-identification branch respectively. Next, perform bounding box temporal association. Finally, the position trajectory is obtained through the tracking results.

[0039] However, the Drosophila observation experiment conducted in the space station has the following characteristics: the motion in the microgravity environment is nonlinear, random, and highly dynamic; the Drosophila targets have high density and frequent occlusion; the appearance similarity of Drosophila targets is high; and there is a stable scene under the fixed field of view.

[0040] The related technology 1 relies on the Kalman filter to predict the trajectory candidate boxes. However, the non-linear motion of Drosophila in the microgravity environment (such as high-frequency turning and hovering) will cause a significant increase in the prediction deviation of the motion model, making it difficult for the trajectory confidence calculation to accurately reflect the true tracking state. Especially in the case of long-term occlusion, it is easy to cause trajectory drift or loss. Secondly, although the MOTDT method alleviates the identity switching problem through re-identification features, the highly similar appearance features among Drosophila individuals will greatly weaken the discriminative ability of the feature space, resulting in frequent ID switching. Moreover, this method adopts an online processing mode and cannot use future frame information for trajectory backtracking and correction. It is difficult to recover the interrupted trajectory in the target high-density occlusion scenario, exacerbating trajectory fragmentation. In addition, the candidate box scoring mechanism of the MOTDT method depends on the quality of the detection box, and the fast dynamic motion of Drosophila easily leads to inaccurate detection box positioning or missed detection, which may falsely suppress valid tracking candidate boxes and cause trajectory initialization failure.

[0041] The related technology 2 adopts a linear motion prediction model based on the Kalman filter, which is not optimized for the stable background of the fixed field of view, may introduce irrelevant noise, lacks a motion prior modeling based on physical constraints, and is difficult to adapt to the non-linear and high-dynamic flight trajectories of Drosophila in the microgravity environment, and there are biases in the trajectory prediction deviation. And its re-identification branch only uses 128-dimensional low-dimensional features, which are difficult to capture subtle differences when the appearances of Drosophila are highly similar, and are prone to false associations. At the same time, due to the feature conflict problem between the detection and re-identification tasks, combined with the design that only relies on the alignment of the target center point features, it will lead to feature ambiguity and identity switching in the frequent occlusion scenario.

[0042] It can be seen that for the Drosophila space science experiment videos obtained in the space station, the above-mentioned related technologies for multi-object tracking do not model the non-linear motion in microgravity, ignore the significant differences between the characteristics of Drosophila targets and the data distribution of Drosophila space science experiment videos, resulting in the deviation of the multi-object detection center, leading to feature ambiguity, identity switching in the frequent occlusion scenario, and the failure of non-linear trajectory prediction. As a result, it is impossible to stably and accurately achieve multi-object tracking of Drosophila, and it is difficult to meet the scientific requirements of trajectory accuracy and long-term identity consistency for Drosophila behavior analysis in the space station.

[0043] To solve the above problems, the embodiments of the present application provide a Drosophila multi-object tracking method and system based on space science experiment videos. This method and system can be applied to Drosophila observation experiment videos collected in a stable scenario under microgravity and a fixed field of view. For example, it can be applied to Drosophila space science experiment videos collected during Drosophila observation experiments in the space station. Specifically, Figure 1 is a schematic diagram of the principle of the Drosophila multi-object tracking method based on space science experiment videos provided by the embodiments of the present application, as Figure 1As shown, this method can extract optical flow information from the image frames in the Drosophila space science experiment video to achieve decoupling of microgravity motion. Then, feature extraction is performed based on the image frames and optical flow information in the Drosophila space science experiment video to determine static appearance features and motion optical flow features. Next, multi-modal feature progressive fusion is performed based on the static appearance features and motion optical flow features to obtain the fused target fusion features. Further, conditional prediction is performed based on the target fusion features to obtain the heatmap prediction result, offset prediction result, and target size prediction result. Finally, offset prediction association is performed based on the prediction results. In the case where the offset prediction association is successful, the motion trajectory information of the target Drosophila is updated based on the prediction results to obtain the multi-target tracking result, thereby achieving accurate multi-target tracking of Drosophila. This method is based on a Drosophila multi-target tracking framework that fuses appearance and motion dual modalities. By extracting optical flow information, it realizes the decoupling of motion features and performs multi-modal progressive feature fusion with static appearance features, enabling precise and stable multi-target tracking of Drosophila.

[0044] Next, the solution provided in the embodiments of the present application will be introduced with reference to the accompanying drawings.

[0045] Specifically, Figure 2 is a schematic flowchart of the Drosophila multi-target tracking method based on the space science experiment video provided in the embodiments of the present application. As Figure 2 shown, the Drosophila multi-target tracking method based on the space science experiment video provided in the embodiments of the present application includes the following steps S101-S107: S101. Extract the t-th frame image, the (t + 1)-th frame image, and the (t - 1)-th frame image from the Drosophila space science experiment video.

[0046] In the embodiments of the present application, the Drosophila space science experiment video is pre-acquired video data. Exemplarily, the Drosophila space science experiment video can be an experimental observation video collected from a Drosophila observation experiment conducted in the Life Ecological Experiment Cabinet carried on the space station. This Drosophila space science experiment video can be used to characterize the motion process of multiple target Drosophila in the Life Ecological Experiment Cabinet.

[0047] Specifically, the Drosophila space science experiment video consists of multiple consecutive image frames. In the embodiments of the present application, taking the multi-target tracking of the target Drosophila in the t-th frame image corresponding to the t-th moment as an example, first, the t-th frame image, the previous frame image adjacent to the t-th frame image, and the subsequent frame image, that is, the (t - 1)-th frame image and the (t + 1)-th frame image, need to be extracted from the Drosophila space science experiment video. Among them, t is a positive integer greater than 1. The t-th frame image, the (t + 1)-th frame image, and the (t - 1)-th frame image can all be RGB frame images.

[0048] S102. Extract optical flow information based on the t-th frame image and the (t + 1)-th frame image to determine the first optical flow image corresponding to the t-th frame image; extract optical flow information based on the t-th frame image and the (t - 1)-th frame image to determine the second optical flow image corresponding to the (t - 1)-th frame image.

[0049] Specifically, the first optical flow image corresponding to the t-th frame image can be determined by extracting optical flow information from the t-th frame image and the (t + 1)-th frame image. The first optical flow image includes the optical flow information (i.e., optical flow vectors, including: direction and magnitude of motion) corresponding to the pixel points in the t-th frame image.

[0050] The second optical flow image corresponding to the (t - 1)-th frame image can be determined by extracting optical flow information from the t-th frame image and the (t - 1)-th frame image. The second optical flow image includes the optical flow information corresponding to the pixel points in the (t - 1)-th frame image.

[0051] In the embodiments of the present application, in view of the characteristics of the video background stability and no dynamic interference in the Drosophila space science experiment video, optical flow information can be extracted from the frame images of the Drosophila space science experiment video, that is, the microgravity motion decoupling of the Drosophila space science experiment video is performed through optical flow estimation by frame difference, and the non-linear motion characteristics in the microgravity environment are separated, such as: hovering, acceleration mutation, etc. In this way, the dynamic characteristics of the motion trajectory of the target Drosophila can be directly modeled through the optical flow information, so as to effectively solve the problem of motion trajectory prediction deviation caused by the microgravity drift of the target Drosophila.

[0052] In some embodiments, Figure 3 is a schematic flowchart of the method for determining the first optical flow image provided by the embodiments of the present application. As Figure 3 shown, taking the extraction of optical flow information from the t-th frame image and the (t + 1)-th frame image to determine the first optical flow image corresponding to the t-th frame image as an example, S102 may specifically include the following steps S1021 - S1025: S1021. Perform preprocessing on the t-th frame image and the (t + 1)-th frame image respectively to obtain the corresponding first preprocessed image and second preprocessed image.

[0053] Specifically, the preprocessing may include: image grayscale processing, Gaussian blur processing, etc. In this way, the computational complexity of subsequent image processing can be reduced, and the influence of noise in the t-th frame image and the (t + 1)-th frame image can be smoothed, and then the accuracy of subsequent image processing can be improved.

[0054] S1022. Perform image subtraction processing based on the first preprocessed image and the second preprocessed image to determine the first difference image.

[0055] Then, the first preprocessed image and the second preprocessed image can be subjected to image subtraction processing. For example, frame difference method can be used for image subtraction processing to determine the first difference image. The first difference image includes the gray value difference corresponding to each pixel point.

[0056] S1023. Determine the significant feature points according to the first difference image through a corner detection algorithm.

[0057] Furthermore, whether there is a moving target (i.e., the target Drosophila) can be determined according to the first difference image, that is, the gray value difference. Specifically, the feature points with obvious structural changes (i.e., significant feature points) in the first difference image can be determined through a corner detection algorithm, and the significant feature points can be used for the target Drosophila.

[0058] Exemplarily, first, the Shi-Tomasi corner detection algorithm can be used to detect the corner response value corresponding to the pixel points in the first difference image. The expression of the corner response value is: .

[0059] Where, represents the corner response value, and represent the eigenvalues of the image gradient matrix in the first difference image, and the smaller of the two eigenvalues is taken as the corner response value.

[0060] Then, the pixel points corresponding to the corner response values in the first difference image that satisfy the preset corner response threshold can be determined as significant feature points.

[0061] S1024. Determine the first optical flow vector corresponding to the pixel points in the t-th frame image according to the first preprocessed image, the second preprocessed image and the significant feature points.

[0062] Furthermore, according to the first preprocessed image and the second preprocessed image, the displacement vector of the significant feature points can be determined, and it can be extended to the pixel-level displacement of the entire image of the t-th frame image through bilinear interpolation. Then, through constructing an image pyramid, filtering and filtering processing, the first optical flow vector can be obtained.

[0063] Specifically, in some embodiments, the above S1024 may specifically include the following steps S10241 - S10245: S10241. Determine the displacement vector of the significant feature points according to the first preprocessed image and the second preprocessed image by minimizing the optical flow error function, first-order Taylor expansion linearization and least squares method.

[0064] Specifically, for the significant feature points determined in S1023, with the precondition that the pixel points within the first preset neighborhood window of the significant feature points have the same motion vector, the displacement vector of the significant feature points can be determined by minimizing the optical flow error function. The expression of the displacement vector of the significant feature points is as follows: .

[0065] Among them, represents the displacement vector of the significant feature point, represents the first preset neighborhood window. For example: can be a 3×3 or 5×5 square area. represents the gray value of the pixel point ( ) in the first preprocessed image, represents the gray value of the pixel point ( ) in the second preprocessed image.

[0066] Then, based on the above expression of the displacement vector of the significant feature points, through first-order Taylor expansion for linearization, an approximate error function is obtained: .

[0067] .

[0068] .

[0069] .

[0070] Among them, and represent the image space gradient, represents the time gradient.

[0071] Next, by solving with the least squares method, a closed-form solution of the displacement vector of the significant feature points is obtained: .

[0072] S10242. Determine the displacement vector corresponding to the pixel points in the t-th frame image according to the displacement vector of the significant feature points through bilinear interpolation.

[0073] Specifically, through bilinear interpolation, the displacement vector of the significant feature points can be extended to the pixel-level displacement of the entire image of the t-th frame, and the displacement vector corresponding to the pixel points in the t-th frame image is obtained.

[0074] .

[0075] .

[0076] Among them, Denote the pixel point in the t-th frame image The corresponding displacement vector Denote the one that is The nearest significant feature point Denote the significant feature point The displacement vector of Is the interpolation weight

[0077] S10243. According to the displacement vector corresponding to the pixel point in the t-th frame image, perform optical flow field superposition by constructing an image pyramid with a preset number of layers to determine the superimposed optical flow vector

[0078] Furthermore, according to the displacement vector corresponding to the pixel point in the t-th frame image determined in S10242, an image pyramid can be constructed, and the extracted optical flow field can be optimized from coarse to fine using multi-scale representation. Calculate the optical flow field at each layer and use the result as the initial value of the optical flow field at the next layer

[0079] Exemplarily, the expression of the optical flow field at the layer can be: .

[0080] Wherein Is the optical flow field at the layer, that is, the displacement vector at the pixel point , which can be determined according to to obtain Denote the optical flow correction amount of the optical flow field at the layer Denote the bilinear upsampling operation

[0081] The higher the layer number ( the larger), the lower the resolution of the image, and the processing result is from coarse to fine. Therefore, first calculate the optical flow field , is the preset number of layers) of the top layer, and then upsample it to the layer, add the optical flow correction amount of this layer , and so on, until the bottom layer . The final expression of the superimposed optical flow vector after optical flow field superposition can be recursively expanded as: .

[0082] Wherein Denote the superimposed optical flow vector Denote the preset number of layers. For example, in the embodiments of the present application Can be set to 3 Denote the cumulative upsampling operation from the layer to the 0-th layer is the optical flow correction amount of the layer, and its coordinates need to be mapped to the original resolution space.

[0083] S10244. Perform median filtering on the superimposed optical flow vectors to determine the filtered optical flow vectors.

[0084] Furthermore, median filtering can be performed on the horizontal and vertical components of the superimposed optical flow vectors to remove noise and obtain the filtered optical flow vectors.

[0085] .

[0086] Among them, represents the filtered optical flow vector, represents the second preset neighborhood window, which can be the same as the first preset neighborhood window. For example: it can be a square area of 3×3 or 5×5. represents the median operation, and the middle value is taken after sorting the horizontal and vertical components of the superimposed optical flow vectors within .

[0087] S10245. Filter the filtered optical flow vectors according to a preset speed amplitude threshold to obtain the first optical flow vectors.

[0088] Finally, the filtered optical flow vectors can be further filtered through a preset speed amplitude threshold to filter out abnormal movements and obtain the filtered optical flow vectors, that is, the first optical flow vectors. Exemplarily, the expression of the first optical flow vector is: .

[0089] Among them, represents the first optical flow vector, represents the speed amplitude of the optical flow, represents the preset speed amplitude threshold. Exemplarily, can be the 95th percentile of the optical flow amplitude distribution in the image.

[0090] S1025. Generate a first optical flow image based on the first optical flow vectors corresponding to the pixel points in the t-th frame image.

[0091] Finally, a first optical flow image can be generated according to the first optical flow vectors corresponding to the pixel points in the t-th frame image determined in S1024.

[0092] It should be noted that the method principle for determining the second optical flow image corresponding to the (t - 1)-th frame image in S102 is the same as the method principles shown in the above S1021 - S1025, and will not be elaborated here.

[0093] S103. Extract features from the t-th frame image, the first optical flow image, the (t - 1)-th frame image, and the second optical flow image respectively to determine the first static appearance feature, the first motion optical flow feature, the second static appearance feature, the second motion optical flow feature, and the heat map feature.

[0094] In the embodiments of the present application, in order to perform multi-modal feature progressive fusion, it is first necessary to extract features from the t-th frame image, the first optical flow image, the (t - 1)-th frame image, and the second optical flow image. Specifically, the t-th frame image, the first optical flow image, the (t - 1)-th frame image, and the second optical flow image can be respectively subjected to feature extraction through a feature extraction network (for example: ResNet network) to obtain the first static appearance feature corresponding to the t-th frame image, the first motion optical flow feature corresponding to the first optical flow image, the second static appearance feature corresponding to the (t - 1)-th frame image, and the second motion optical flow feature corresponding to the second optical flow image. Among them, the static appearance features (the first static appearance feature, the second static appearance feature) can be used to characterize the static features of the pixel points in the frame images (the t-th frame image, the (t - 1)-th frame image), such as color values, grayscale values, etc. The motion optical flow features (the first motion optical flow feature, the second motion optical flow feature) can be used to characterize the dynamic features of the pixel points in the optical flow images (the first optical flow image, the second optical flow image), such as motion directions, motion speeds, etc.

[0095] Moreover, the heat map feature can also be extracted from the (t - 1)-th frame image through a target detection network for feature fusion. Specifically, this heat map feature can be used to characterize the probability distribution of the target center point in the (t - 1)-th frame image, and this target center point can be used to characterize the center point of the target Drosophila.

[0096] S104. Perform multi-modal feature progressive fusion on the first static appearance feature, the first motion optical flow feature, the second static appearance feature, the second motion optical flow feature, and the heat map feature to obtain the target fusion feature.

[0097] In the embodiments of the present application, the first static appearance feature, the first motion optical flow feature, the second static appearance feature, the second motion optical flow feature, and the heat map feature extracted in S103 are subjected to multi-level and progressive fusion. The static appearance feature extraction is guided by the optical flow field to enhance the robustness of the motion blur area detection. The subsequent generation of the conditional prediction heat map is constrained by the temporal consistency of the motion field to suppress false detections and missed detections in high-density occlusion scenarios, so as to provide an appearance-motion joint representation. Specifically, the multi-modal feature progressive fusion is divided into two fusion stages, respectively fusing the temporal features and the multi-modal features to make full use of the multi-modal information of visual appearance and motion optical flow as well as the temporal information.

[0098] Exemplarily, the expression of the target fusion feature can be: 。

[0099] Among them, represents the target fusion feature, represents the time feature fusion operation, represents the multi-modal feature fusion operation, represents the first static appearance feature, represents the first motion optical flow feature, represents the second static appearance feature, represents the second motion optical flow feature, represents the heat map feature.

[0100] For multi-object tracking, extracting time information is crucial for improving multi-object tracking performance. Since the t-th frame image and the (t - 1)-th frame image are not strictly spatially aligned, it is difficult to effectively integrate the feature information of the (t - 1)-th frame image. Therefore, in the time feature fusion stage, the embodiments of the present application adopt an attention mechanism that does not rely on strict spatial alignment of adjacent frames to effectively integrate time information.

[0101] Specifically, in some embodiments, S104 may specifically include the following steps S1041 - S1046: S1041. Perform a first feature fusion on the first static appearance feature and the second static appearance feature based on the cross-attention mechanism to obtain a static appearance fusion feature; perform a second feature fusion on the first motion optical flow feature and the second motion optical flow feature based on the cross-attention mechanism to obtain a motion optical flow fusion feature.

[0102] Specifically, Figure 4 is a schematic diagram of the principle of progressive multi-modal feature fusion provided by the embodiments of the present application. As Figure 4 shown, a first feature fusion can be performed based on the cross-attention mechanism using the first static appearance feature as the query Q (i.e., Figure 4 shown in ), the second static appearance feature as the key K (i.e., Figure 4 shown in ) and the value V (i.e., Figure 4 shown in ) to capture the spatio-temporal context information in the static appearance feature and obtain a static appearance fusion feature.

[0103] Moreover, a second feature fusion can also be performed based on the cross-attention mechanism using the first motion optical flow feature as the query Q (i.e., Figure 4 shown in ), the second motion optical flow feature as the key K (i.e., Figure 4 shown in ) and the value V (i.e., Figure 4 shown in )(2) Perform the second feature fusion to capture the spatio-temporal context information in the motion optical flow features, and obtain the motion optical flow fusion features.

[0104] Exemplarily, the expression of the static appearance fusion feature can be: .

[0105] Wherein, represents the static appearance fusion feature, represents the position encoding, represents the cross-attention operation.

[0106] The expression of the motion optical flow fusion feature can be: .

[0107] Wherein, represents the motion optical flow fusion feature.

[0108] S1042. Determine the static appearance temporal feature according to the static appearance fusion feature and the heat map feature; determine the motion optical flow temporal feature according to the motion optical flow fusion feature and the heat map feature through a second feed-forward neural network.

[0109] Then, in order to enhance the localization ability of target tracking, the heat map feature of the (t - 1)-th frame image is used as the position condition and integrated into the static appearance fusion feature and the motion optical flow fusion feature, and finally the corresponding temporal feature is obtained through a feed-forward neural network to enhance the target feature representation.

[0110] Exemplarily, the expression of the static appearance temporal feature is: .

[0111] .

[0112] Wherein, represents the static appearance temporal feature, represents the layer normalization operation, represents the feed-forward neural network, represents the intermediate result of feature processing.

[0113] The expression of the motion optical flow temporal feature is: .

[0114] .

[0115] Wherein, represents the motion optical flow temporal feature.

[0116] S1043. Add the static appearance temporal feature and the motion optical flow temporal feature to obtain an initial multi-modal feature.

[0117] In the progressive multi-modal feature fusion adopted in the embodiments of the present application, the interaction between modalities is crucial for obtaining effective multi-modal features. The cross-attention mechanism tends to enhance the similarity information (homogeneous information) between modalities, while potentially ignoring modality-specific information (heterogeneous information). Therefore, in the multi-modal fusion stage, an addition operation is used to obtain the initial multi-modal feature, and the initial multi-modal feature is used as a bridging feature to interact with the single-modal features. In this way, the problems caused by directly interacting with single-modal features can be effectively avoided.

[0118] Specifically, the static appearance temporal feature and the motion optical flow temporal feature can be integrated through a feature addition operation to obtain an initial multi-modal feature representation, that is, an initial multi-modal feature is obtained.

[0119] Exemplarily, the expression of the initial multi-modal feature is: .

[0120] Where represents the initial multi-modal feature.

[0121] S1044. Based on the cross-attention mechanism, use the initial multi-modal feature as the key K and value V, and the static appearance temporal feature as the query Q for the third feature fusion to obtain the first multi-modal feature; based on the cross-attention mechanism, use the initial multi-modal feature as Q, and the static appearance temporal feature as K and V for the fourth feature fusion to obtain the second multi-modal feature; based on the cross-attention mechanism, use the initial multi-modal feature as K and V, and the motion optical flow temporal feature as Q for the fifth feature fusion to obtain the third multi-modal feature; based on the cross-attention mechanism, use the initial multi-modal feature as Q, and the motion optical flow temporal feature as K and V for the sixth feature fusion to obtain the fourth multi-modal feature.

[0122] Furthermore, the multi-modal feature (i.e., the initial multi-modal feature) can be used as K and V, while the two single-modal features (the static appearance temporal feature and the motion optical flow temporal feature) are used as Q respectively. Based on the cross-attention mechanism, interact the single-modal features and the multi-modal features to enhance the representation of modality-specific features. At the same time, the multi-modal feature can be used as Q, and the two single-modal features are used as K and V respectively. Then, based on the cross-attention mechanism, further enhance the multi-modal feature.

[0123] S1045. Concatenate the first multi-modal feature, the second multi-modal feature, the third multi-modal feature, and the fourth multi-modal feature to obtain a multi-modal concatenated feature.

[0124] Exemplarily, the expression of the multi-modal concatenated feature is: 。

[0125] Among them, represents the multi-modal splicing feature, represents the cross-attention operation, represents the first multi-modal feature, represents the second multi-modal feature, represents the third multi-modal feature, represents the fourth multi-modal feature, represents the feature splicing operation.

[0126] S1046. Determine the target fusion feature through the third feed-forward neural network according to the multi-modal splicing feature.

[0127] Finally, input the multi-modal splicing feature into the third feed-forward neural network to obtain the refined multi-modal feature, that is, the target fusion feature. Exemplarily, the target fusion feature can be specifically expressed as: 。

[0128] Among them, represents the target fusion feature.

[0129] S105. Perform conditional prediction through the target detection network model according to the target fusion feature, and determine the heatmap prediction result, offset prediction result, and target size prediction result of the target center point in the t-th frame image.

[0130] In the embodiments of the present application, the conditional prediction can introduce the multi-object tracking results, image features, and motion features of the previous frame (i.e., the (t - 1)-th frame image) according to the target fusion feature, enhance the detection robustness of the current frame (i.e., the t-th frame image), and achieve the joint optimization of detection and tracking. It not only explicitly uses the temporal information but also fuses the motion information of the optical flow field for joint inference.

[0131] In one implementation, the target detection network model can adopt the CenterNet model. Specifically, based on the target fusion feature determined in S104, through three parallel branches of the CenterNet model, the heatmap prediction result and offset prediction result of the target center point in the t-th frame image, and the target size prediction result can be obtained respectively.

[0132] In some embodiments, the expression of the heatmap prediction result is: 。

[0133] Among them, represents the heatmap prediction result, represents the Sigmoid activation function, represents the heatmap output operation of the convolutional network, represents the target fusion feature.

[0134] The expression for the target size prediction result is: .

[0135] Among them, represents the target size prediction result, represents the rectified linear unit function, represents the size output operation of the convolutional network.

[0136] The expression for the offset prediction result is: .

[0137] Among them, represents the offset prediction result, represents the offset output operation of the convolutional network.

[0138] In some embodiments, the target detection network model is a network model trained based on a sample training set. Specifically, the loss expression of the target detection network model is: .

[0139] .

[0140] .

[0141] .

[0142] Among them, represents the loss of the target detection network model; represents the prediction loss of the heatmap, represents the prediction loss of the target size, represents the prediction loss of the offset; represents the number of positive samples in the training set, represents the ground truth of the heatmap, the predicted value of the heatmap, represents the weight coefficient of the positive sample, represents the weight coefficient of the negative sample in the training set; represents the predicted value of the width in the target size, represents the ground truth of the width in the target size, represents the predicted value of the height in the target size, represents the ground truth of the height in the target size; represents the predicted value of the width offset, Represents the true value of the width offset, Represents the predicted value of the height offset, Represents the true value of the height offset.

[0143] S106. Perform offset prediction association based on the heatmap prediction result, the offset prediction result, and the heatmap feature.

[0144] In the embodiment of the present application, in order to solve the problem of cross-frame target matching ambiguity caused by the fact that distance matching is vulnerable to occlusion interference and the appearance of target fruit flies is similar and re-identification features cannot be relied on. The target displacement is predicted through the offset, improving the adaptability to non-linear patterns. Using the previous frame heatmap (i.e., the heatmap feature) as prior knowledge to constrain the search space for the detection of the t-th frame image, reducing false detections and missed detections caused by occlusion, etc.

[0145] Specifically, in some embodiments, S106 may specifically include the following steps S1061 - S1063: S1061. Extract local extreme points from the heatmap prediction result to determine the detection points corresponding to the t-th frame image.

[0146] S1062. Determine the backtracking position corresponding to the detection point according to the offset prediction result.

[0147] Among them, the expression of the backtracking position is: .

[0148] Among them, Represents the backtracking position, Represents the abscissa of the detection point in the t-th frame image, Represents the predicted value of the offset of the abscissa and ordinate of the detection point, Represents the ordinate of the detection point in the t-th frame image.

[0149] S1063. Determine the heatmap response corresponding to the backtracking position according to the heatmap feature, and determine whether the heatmap response corresponding to the backtracking position meets the preset response threshold.

[0150] Specifically, the heatmap response corresponding to the backtracking position can be determined according to the heatmap feature. Further, it can be determined whether the heatmap response meets the preset response threshold, that is, it can be determined whether there is a significant heatmap response near the backtracking position to determine whether the offset prediction association is successful. Among them, the preset response threshold can be preset according to prior knowledge and the requirements of actual applications.

[0151] S107. In the case where the offset prediction association is successful, update the motion trajectory information of the target fruit fly corresponding to the t-th frame image according to the offset prediction result and the target size prediction result.

[0152] Finally, when the heatmap response corresponding to the backtracking position satisfies a preset response threshold, it can be determined that the offset prediction is successfully associated. In this case, the target fruit fly motion trajectory information corresponding to the t-th frame image can be updated according to the offset prediction result and the target size prediction result to obtain the multi-target tracking result of the target fruit fly.

[0153] In some embodiments, the target fruit fly motion trajectory information includes: the abscissa of the center point, the ordinate of the center point, the width, the height, and the trajectory number corresponding to the target fruit fly.

[0154] Then S107 specifically includes: superimposing the offset prediction result, the target size prediction result, and the target fruit fly motion trajectory information corresponding to the (t - 1)-th frame image to obtain the target fruit fly motion trajectory information corresponding to the t-th frame image.

[0155] Among them, the expression of the target fruit fly motion trajectory information corresponding to the t-th frame image is: .

[0156] = + .

[0157] = + .

[0158] Among them, represents the -th target fruit fly motion trajectory information corresponding to the t-th frame image, represents the -th target fruit fly motion trajectory information corresponding to the (t - 1)-th frame image, represents the abscissa of the center point of the -th target fruit fly, represents the ordinate of the center point of the -th target fruit fly, represents the predicted width value in the target size prediction result of the -th target fruit fly, represents the predicted height value in the target size prediction result of the -th target fruit fly, represents the trajectory number of the -th target fruit fly, represents the abscissa of the center point of the -th target fruit fly in the (t - 1)-th frame image, represents the predicted abscissa offset value in the offset prediction result of the -th target fruit fly, represents the The vertical coordinate of the center point of a target fruit fly in the (t - 1)-th frame image, denotes the predicted value of the vertical coordinate offset in the offset prediction result of the

[0159] Using the fruit fly multi-target tracking method based on space science experiment videos provided in the embodiments of the present application, first, the microgravity motion characteristics are decoupled by extracting the optical flow information of the frame images in the fruit fly space science experiment video. Then, the static appearance features, motion optical flow features, and heatmap features are progressively fused in a multi-modal manner to obtain the target fusion features. Next, conditional prediction is performed based on the target fusion features to determine the heatmap prediction result, offset prediction result, and target size prediction result of the target center point. Finally, offset prediction association is performed based on the prediction results. When the offset prediction association is successful, the motion trajectory information of the target fruit fly is updated according to the prediction results, thereby realizing the multi-target tracking of fruit flies.

[0160] Aiming at the characteristics of the nonlinear and random motion of fruit flies in the microgravity environment, this method adopts an optical flow feature extraction and microgravity motion decoupling scheme. Through multi-scale feature point detection and sub-pixel level positioning optimization, a high-precision optical flow field is generated to capture instantaneous motion vectors such as hovering and sudden acceleration, providing complementary semantic enhancement for the appearance features. Aiming at the problems that fruit flies have high appearance similarity and high-density occlusion, a single modality (such as RGB) is prone to identity confusion (IDSW), and direct feature addition causes modal noise superposition, this method uses multi-modal feature progressive fusion to enhance the robustness of motion blur area detection. By the temporal consistency constraint of the motion field, a conditional detection heatmap can be generated subsequently, which can suppress false detections and missed detections in high-density occlusion scenarios and provide an appearance-motion joint representation. In this way, based on the fruit fly multi-target tracking framework of appearance-motion bimodal fusion, this method can realize accurate and stable multi-target tracking of fruit flies by extracting optical flow information to decouple motion features and performing multi-modal progressive feature fusion with static appearance features.

[0161] In some embodiments, to verify the performance and effectiveness of the fruit fly multi-target tracking method based on space science experiment videos provided in the above embodiments of the present application, verification and analysis are carried out.

[0162] Experimental dataset: Representative time periods in the fruit fly life cycle are selected, and a fruit fly multi-target tracking dataset is refined and manually annotated. The fruit fly space science experiment video dataset: It contains 20 videos with a frame rate of 25 frames per second, a total of 2500 frame images, and a total of 27500 fruit fly instances.

[0163] Evaluation method: The evaluation metrics for the performance of the Drosophila multi-object tracking method in this technical solution include: HOTA (Higher Order Tracking Accuracy), MOTA (Multi-Object Tracking Accuracy), IDF1 (Identity F1 Score), and IDs (Number of identity switches). The higher the values of HOTA, MOTA, and IDF1, and the lower the value of the IDs metric, the better the tracking effect. The specific calculation methods are as follows: 。

[0164] 。

[0165] 。

[0166] Among them, represents the detection accuracy rate, represents the association accuracy rate, represents the number of missed detection targets, represents the number of false detection targets, represents the number of times of incorrect identity switches, represents the number of ground truth targets, represents the proportion of detection boxes with correctly associated target identities, represents the proportion of correctly associated true targets with the same identity.

[0167] The Drosophila multi-object tracking accuracy experiments were respectively carried out on the Drosophila space science experiment video dataset using the above-mentioned related technology one, related technology two, and the Drosophila multi-object tracking method provided by the embodiment of this application (hereinafter referred to as: this solution). Table 1 shows the experimental results of the Drosophila multi-object tracking accuracy.

[0168] Table 1 Experimental results of Drosophila multi-object tracking accuracy As shown in Table 1, this solution is superior to related technology one and related technology two in terms of indicators such as HOTA (83.15), MOTA (89.045), IDF1 (85.174), and IDs (65). The HOTA indicator combines the dual advantages of detection accuracy (DetA) and association robustness (AssA). The score of 83.15 in this solution verifies the synergistic effect of optical flow motion feature decoupling and cross-modal feature fusion. The extremely high value of MOTA (89.045) is based on the targeted design of the experimental scene background stability of the Drosophila space science experiment video in this solution. The significant advantages of IDF1 (85.174) and IDs (65) reflect the key role of motion trajectory decoupling in the identity consistency of target Drosophila. Through the dynamic weight fusion of motion optical flow features and static appearance features, the limitation of relying on a single visual feature in related technologies is broken. Even if the appearances of target Drosophila are highly similar, the uniqueness of the motion trajectory can still ensure the stability of identity association.

[0169] Furthermore, in some embodiments, ablation experiments were conducted to verify the effects of the steps in this solution. By gradually adding steps such as introducing the motion optical flow feature step and the multi-modal feature progressive fusion step, the influence of each step on the accuracy of multi-object tracking of target fruit flies was analyzed. Table 2 shows the results of the ablation experiment for fruit fly multi-object tracking. Among them, "O" in Table 2 indicates the use of the corresponding step method.

[0170] Table 2 Experimental results of fruit fly multi-object tracking accuracy As shown in Table 2, after introducing the motion optical flow feature step, HOTA increased from 78.8 to 80.889 (+2.64%), MOTA increased from 82.583 to 85.528 (+3.55%), IDF1 increased from 77.998 to 82.675 (+5.99%), and IDs decreased from 92 to 69 (-25%). This shows that introducing the motion optical flow feature can effectively model the non-linear motion in the microgravity environment (such as the "floating" and "somersaulting" behaviors of fruit flies), enhance the robustness of trajectory prediction through the motion vector field, and reduce trajectory breaks (FN) and false detections (FP) caused by sudden accelerations. In addition, the spatio-temporal continuity of the motion optical flow feature supports short-term trajectory extrapolation, maintains the consistency of target identity during occlusion, and significantly reduces the number of IDSW.

[0171] After the superposition of the multi-modal feature progressive fusion step, HOTA increased from 80.889 to 83.15 (+2.8%), MOTA increased from 85.528 to 89.045 (+4.12%), IDF1 increased from 82.675 to 85.174 (+3.03%), and IDs decreased from 69 to 51 (-26%). This shows that multi-modal feature progressive fusion can fuse the motion optical flow feature and the static appearance feature in stages through the cross-modal attention mechanism, enhance the positioning ability for blurred targets in the detection stage (heat map generation), suppress the interference of appearance similarity in the association stage (feature matching), automatically adjust the fusion weights of the motion optical flow feature and the static appearance feature for different scenarios (such as high-density occlusion and fast movement), and preferentially use the differences in motion trajectories to distinguish targets.

[0172] The embodiment of the present application also provides a fruit fly multi-object tracking system based on space science experiment videos. Specifically, Figure 5 is a schematic structural diagram of the fruit fly multi-object tracking system provided by the embodiment of the present application. As Figure 5 shown, the fruit fly multi-object tracking system 500 based on space science experiment videos includes: an image extraction module 501, an optical flow information extraction module 502, a feature extraction module 503, a multi-modal feature progressive fusion module 504, a conditional prediction module 505, an offset prediction association module 506, and a trajectory generation module 507.

[0173] Among them, the image extraction module 501 can be used to extract the t-th frame image, the (t + 1)-th frame image, and the (t - 1)-th frame image from the Drosophila space science experiment video, where t is a positive integer greater than 1.

[0174] The optical flow information extraction module 502 can be used to extract optical flow information based on the t-th frame image and the (t + 1)-th frame image to determine the first optical flow image corresponding to the t-th frame image; extract optical flow information based on the t-th frame image and the (t - 1)-th frame image to determine the second optical flow image corresponding to the (t - 1)-th frame image.

[0175] The feature extraction module 503 can be used to extract features from the t-th frame image, the first optical flow image, the (t - 1)-th frame image, and the second optical flow image respectively to determine the first static appearance feature, the first motion optical flow feature, the second static appearance feature, the second motion optical flow feature, and the heatmap feature. The heatmap feature is used to characterize the probability distribution of the target center point in the (t - 1)-th frame image, and the target center point is used to characterize the center point of the target Drosophila.

[0176] The multi-modal feature progressive fusion module 504 can be used to perform multi-modal feature progressive fusion on the first static appearance feature, the first motion optical flow feature, the second static appearance feature, the second motion optical flow feature, and the heatmap feature to obtain the target fusion feature.

[0177] The conditional prediction module 505 can be used to perform conditional prediction through the target detection network model according to the target fusion feature to determine the heatmap prediction result, the offset prediction result, and the target size prediction result of the target center point in the t-th frame image.

[0178] The offset prediction association module 506 can be used to perform offset prediction association according to the heatmap prediction result, the offset prediction result, and the heatmap feature.

[0179] The trajectory generation module 507 can be used to update the target Drosophila motion trajectory information corresponding to the t-th frame image according to the offset prediction result and the target size prediction result when the offset prediction association is successful.

[0180] In some embodiments, the above multi-modal feature progressive fusion module 504 includes a multi-modal feature progressive fusion model 600. Specifically, Figure 6 is a schematic structural diagram of the multi-modal feature progressive fusion model provided by the embodiments of the present application, as Figure 6As shown in the figure, the multi-modal feature progressive fusion model 600 includes: a first fusion module 601, a second fusion module 602, a first feed-forward neural network 603, a second feed-forward neural network 604, a feature addition module 605, a third fusion module 606, a fourth fusion module 607, a fifth fusion module 608, a sixth fusion module 609, a feature splicing module 610, and a third feed-forward neural network 611.

[0181] Among them, the first fusion module 601 can be used to perform a first feature fusion on the first static appearance feature and the second static appearance feature based on the cross-attention mechanism to obtain a static appearance fusion feature.

[0182] The second fusion module 602 can be used to perform a second feature fusion on the first motion optical flow feature and the second motion optical flow feature based on the cross-attention mechanism to obtain a motion optical flow fusion feature.

[0183] The first feed-forward neural network 603 can be used to determine the static appearance time feature according to the static appearance fusion feature and the heat map feature.

[0184] The second feed-forward neural network 604 can be used to determine the motion optical flow time feature according to the motion optical flow fusion feature and the heat map feature.

[0185] The feature addition module 605 can be used to add the static appearance time feature and the motion optical flow time feature to obtain an initial multi-modal feature.

[0186] The third fusion module 606 can be used to perform a third feature fusion based on the cross-attention mechanism, using the initial multi-modal feature as K and V, and the static appearance time feature as Q, to obtain a first multi-modal feature.

[0187] The fourth fusion module 607 can be used to perform a fourth feature fusion based on the cross-attention mechanism, using the initial multi-modal feature as Q, and the static appearance time feature as K and V, to obtain a second multi-modal feature.

[0188] The fifth fusion module 608 can be used to perform a fifth feature fusion based on the cross-attention mechanism, using the initial multi-modal feature as K and V, and the motion optical flow time feature as Q, to obtain a third multi-modal feature.

[0189] The sixth fusion module 609 can be used to perform a sixth feature fusion based on the cross-attention mechanism, using the initial multi-modal feature as Q, and the motion optical flow time feature as K and V, to obtain a fourth multi-modal feature.

[0190] The feature splicing module 610 can be used to splice the first multi-modal feature, the second multi-modal feature, the third multi-modal feature, and the fourth multi-modal feature to obtain a multi-modal splicing feature.

[0191] The third feedforward neural network 611 can be used to determine the target fusion feature according to the multimodal splicing feature.

[0192] Using the Drosophila multi-target tracking system based on the space science experiment video provided by the embodiments of the present application, first, the microgravity motion feature decoupling is realized by extracting the optical flow information of the frame images in the Drosophila space science experiment video. Then, the static appearance feature, the motion optical flow feature, and the heatmap feature are gradually fused in a multimodal manner to obtain the target fusion feature. Next, conditional prediction is performed according to the target fusion feature to determine the heatmap prediction result, the offset prediction result, and the target size prediction result of the target center point. Finally, offset prediction association is performed according to the prediction results. When the offset prediction association is successful, the motion trajectory information of the target Drosophila is updated according to the prediction results, thereby realizing the multi-target tracking of Drosophila. In this way, based on the appearance-motion bimodal fusion Drosophila multi-target tracking framework, the system can realize accurate and stable multi-target tracking of Drosophila by extracting optical flow information to decouple motion features and performing multimodal progressive feature fusion with static appearance features.

[0193] The embodiments of the present invention further provide an electronic device, which may include: a display screen, a memory, and one or more processors. The display screen, the memory, and the processor are coupled. The memory is used to store computer program code, and the computer program code includes computer instructions. When the processor executes the computer instructions, the electronic device can execute each method or step executed in the above embodiments of the Drosophila multi-target tracking method. Of course, the electronic device includes, but is not limited to, the above display screen, memory, and one or more processors.

[0194] The embodiments of the present invention further provide a computer-readable storage medium for storing the computer instructions for running the above Drosophila multi-target tracking method.

[0195] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and conciseness of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0196] In the description of the present invention, it should be understood that the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined.

[0197] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0198] For the similar parts between the embodiments provided in this application, reference can be made to each other. The specific embodiments provided above are only several examples under the general concept of this application and do not constitute a limitation on the protection scope of this application. For those skilled in the art, any other embodiments extended based on the solution of this application without creative efforts belong to the protection scope of this application.

Claims

1. A method for tracking multiple fruit flies based on space science experimental videos, characterized in that: include: Extract the t-th frame image, the t+1-th frame image and the t-1-th frame image from the fruit fly space science experiment video, where t is a positive integer greater than 1; Extracting optical flow information based on the t-th frame image and the t+1-th frame image to determine a first optical flow image corresponding to the t-th frame image; extracting optical flow information based on the t-th frame image and the t-1-th frame image to determine a second optical flow image corresponding to the t-1-th frame image; Performing feature extraction on the t-th frame image, the first optical flow image, the t-1-th frame image, and the second optical flow image, respectively, to determine a first static appearance feature, a first motion optical flow feature, a second static appearance feature, a second motion optical flow feature, and a heat map feature, wherein the heat map feature is used to characterize the probability distribution of the target center point in the t-1-th frame image, and the target center point is used to characterize the center point of the target fruit fly; Performing multimodal feature progressive fusion on the first static appearance feature, the first motion optical flow feature, the second static appearance feature, the second motion optical flow feature, and the heat map feature to obtain the target fusion feature; Performing conditional prediction through a target detection network model according to the target fusion feature, determining a heat map prediction result and an offset prediction result of the target center point in the t-th frame image, and a target size prediction result; Performing offset prediction association according to the heat map prediction result, the offset prediction result and the heat map feature; When the offset prediction association is successful, the target fruit fly motion trajectory information corresponding to the t-th frame image is updated according to the offset prediction result and the target size prediction result.

2. The method according to claim 1, characterized in that The extracting optical flow information according to the t-th frame image and the t+1-th frame image to determine a first optical flow image corresponding to the t-th frame image includes: Preprocessing the t-th frame image and the t+1-th frame image respectively to obtain corresponding first preprocessed images and second preprocessed images; Performing image subtraction processing based on the first preprocessed image and the second preprocessed image to determine a first difference image; Determining significant feature points using a corner detection algorithm according to the first difference image; Determine a first optical flow vector corresponding to a pixel point in the t-th frame image according to the first preprocessed image, the second preprocessed image and the significant feature point; The first optical flow image is generated according to the first optical flow vector corresponding to the pixel point in the t-th frame image.

3. The method according to claim 2, characterized in that The step of determining a first optical flow vector corresponding to a pixel point in the t-th frame image according to the first preprocessed image, the second preprocessed image and the salient feature point comprises: According to the first preprocessed image and the second preprocessed image, determining the displacement vector of the significant feature point by minimizing the optical flow error function, first-order Taylor expansion linearization and least squares solution; Determine the displacement vector corresponding to the pixel point in the t-th frame image by bilinear interpolation method according to the displacement vector of the significant feature point; According to the displacement vector corresponding to the pixel point in the t-th frame image, optical flow field superposition is performed by constructing an image pyramid with a preset number of layers to determine the superimposed optical flow vector; Performing median filtering on the superimposed optical flow vector to determine a filtered optical flow vector; The filtered optical flow vector is filtered according to a preset speed amplitude threshold to obtain the first optical flow vector.

4. The method according to any one of claims 1 to 3, characterized in that: The step of performing multi-modal feature progressive fusion on the first static appearance feature, the first motion optical flow feature, the second static appearance feature, the second motion optical flow feature, and the heat map feature to obtain the target fusion feature includes: Performing a first feature fusion on the first static appearance feature and the second static appearance feature based on a cross attention mechanism to obtain a static appearance fusion feature; Based on the cross attention mechanism, a second feature fusion is performed on the first motion optical flow feature and the second motion optical flow feature to obtain a motion optical flow fusion feature; Determining static appearance temporal features through a first feedforward neural network according to the static appearance fusion features and the heat map features; Determining the motion optical flow temporal features through a second feedforward neural network according to the motion optical flow fusion features and the heat map features; Adding the static appearance time feature and the motion optical flow time feature to obtain an initial multimodal feature; Based on the cross attention mechanism, the initial multimodal feature is used as the key K and the value V, and the static appearance time feature is used as the query Q for third feature fusion to obtain the first multimodal feature; Based on the cross attention mechanism, the initial multimodal feature is used as Q, the static appearance time feature is used as K and V to perform a fourth feature fusion to obtain a second multimodal feature; Based on the cross attention mechanism, the initial multimodal features are used as K and V, and the motion optical flow time features are used as Q to perform fifth feature fusion to obtain a third multimodal feature; Based on the cross attention mechanism, the initial multimodal feature is used as Q, the motion optical flow time feature is used as K and V to perform the sixth feature fusion to obtain the fourth multimodal feature; Performing feature splicing on the first multimodal feature, the second multimodal feature, the third multimodal feature and the fourth multimodal feature to obtain a multimodal splicing feature; The target fusion feature is determined by a third feedforward neural network based on the multimodal splicing feature.

5. The method according to claim 1, characterized in that The expression of the heat map prediction result is: ; in, represents the prediction result of the heat map, represents the Sigmoid activation function, represents the heatmap output operation of the convolutional network, represents the target fusion feature; The expression of the target size prediction result is: ; in, represents the target size prediction result, represents the linear rectification function, Represents the size output operation of the convolutional network; The expression of the offset prediction result is: ; in, represents the offset prediction result, Represents the offset output operation of a convolutional network.

6. The method according to claim 5, characterized in that The loss expression of the target detection network model is: ; ; ; ; in, Represents the loss of the target detection network model; represents the prediction loss of the heat map, represents the prediction loss of the target size, represents the prediction loss of the offset; represents the number of positive samples in the training set, represents the true value of the heat map, The predicted value of the heat map, represents the weight coefficient of the positive sample, Represents the weight coefficient of negative samples in the training set; represents the predicted value of the width in the target size, represents the true value of the width in the target size, represents the predicted value of the height in the target size, represents the pre-truth value of the height in the target size; represents the predicted value of the width offset, represents the true value of the width offset, represents the predicted value of the height offset, True value representing the height offset.

7. The method according to claim 1, characterized in that The performing offset prediction association according to the heat map prediction result, the offset prediction result and the heat map feature includes: Extracting local extreme points from the prediction results of the heat map to determine the detection points corresponding to the t-th frame image; The backtracking position corresponding to the detection point is determined according to the offset prediction result, and the expression of the backtracking position is: ; in, represents the backtracking position, represents the horizontal coordinate of the detection point in the t-th frame image, Indicates the predicted value of the horizontal and vertical coordinates of the detection point. represents the vertical coordinate of the detection point in the t-th frame image; A heat map response corresponding to the backtracking position is determined according to the heat map feature, and it is determined whether the heat map response corresponding to the backtracking position meets a preset response threshold.

8. The method according to claim 1, characterized in that The target fruit fly movement trajectory information includes: the horizontal coordinate of the center point, the vertical coordinate of the center point, the width, the height and the trajectory number corresponding to the target fruit fly; When the offset prediction association is successful, updating the target fruit fly motion trajectory information corresponding to the t-th frame image according to the offset prediction result and the target size prediction result includes: The target fruit fly motion track information corresponding to the t-1 frame image is superimposed according to the offset prediction result, the target size prediction result and the target fruit fly motion track information corresponding to the t-1 frame image to obtain the target fruit fly motion track information corresponding to the t frame image; the expression of the target fruit fly motion track information corresponding to the t frame image is: ; = + ; = + ; in, Indicates the tth frame image corresponding to the The target fruit fly movement trajectory information, Indicates the t-1th frame image corresponding to the The target fruit fly movement trajectory information, Indicates The horizontal coordinate of the center point of the target fruit fly, Indicates The vertical coordinate of the center point of the target fruit fly, Indicates The width prediction value in the target size prediction results of the target fruit fly, Indicates The height prediction value in the target size prediction results of the target fruit fly, Indicates The movement trajectory number of the target fruit fly, Indicates The horizontal coordinate of the center point of the target fruit fly in the t-1 frame image, Indicates The predicted value of the horizontal axis offset in the offset prediction result of the target fruit fly, Indicates The vertical coordinate of the center point of the target fruit fly in the t-1 frame image, Indicates The predicted value of the ordinate offset in the offset prediction results of the target fruit flies.

9. A fruit fly multi-target tracking system based on space science experiment video, characterized in that: include: Image extraction module, optical flow information extraction module, feature extraction module, multimodal feature progressive fusion module, conditional prediction module, offset prediction association module and trajectory generation module; wherein, The image extraction module is used to extract the t-th frame image, the t+1-th frame image and the t-1-th frame image from the fruit fly space science experiment video, where t is a positive integer greater than 1; The optical flow information extraction module is used to extract optical flow information according to the t-th frame image and the t+1-th frame image, and determine a first optical flow image corresponding to the t-th frame image; extract optical flow information according to the t-th frame image and the t-1-th frame image, and determine a second optical flow image corresponding to the t-1-th frame image; The feature extraction module is used to extract features from the t-th frame image, the first optical flow image, the t-1-th frame image and the second optical flow image respectively, and determine a first static appearance feature, a first motion optical flow feature, a second static appearance feature, a second motion optical flow feature and a heat map feature, wherein the heat map feature is used to characterize the probability distribution of the target center point in the t-1-th frame image, and the target center point is used to characterize the center point of the target fruit fly; The multimodal feature progressive fusion module is used to perform multimodal feature progressive fusion on the first static appearance feature, the first motion optical flow feature, the second static appearance feature, the second motion optical flow feature and the heat map feature to obtain the target fusion feature; The conditional prediction module is used to perform conditional prediction through the target detection network model according to the target fusion feature, and determine the heat map prediction result and offset prediction result of the target center point in the t-th frame image, as well as the target size prediction result; The offset prediction association module is used to perform offset prediction association according to the thermal map prediction result, the offset prediction result and the thermal map feature; The trajectory generation module is used to update the target fruit fly motion trajectory information corresponding to the t-th frame image according to the offset prediction result and the target size prediction result when the offset prediction association is successful.

10. The system according to claim 9, characterized in that The multimodal feature progressive fusion module includes a multimodal feature progressive fusion model; the multimodal feature progressive fusion model includes: a first fusion module, a second fusion module, a first feedforward neural network, a second feedforward neural network, a feature addition module, a third fusion module, a fourth fusion module, a fifth fusion module, a sixth fusion module, a feature splicing module and a third feedforward neural network; wherein, The first fusion module is used to perform a first feature fusion on the first static appearance feature and the second static appearance feature based on a cross attention mechanism to obtain a static appearance fusion feature; The second fusion module is used to perform a second feature fusion on the first motion optical flow feature and the second motion optical flow feature based on a cross attention mechanism to obtain a motion optical flow fusion feature; The first feedforward neural network is used to determine a static appearance temporal feature according to the static appearance fusion feature and the heat map feature; The second feedforward neural network is used to determine the motion optical flow time feature according to the motion optical flow fusion feature and the heat map feature; The feature addition module is used to perform feature addition on the static appearance time feature and the motion optical flow time feature to obtain an initial multimodal feature; The third fusion module is used to perform a third feature fusion based on a cross attention mechanism using the initial multimodal features as K and V and the static appearance time features as Q to obtain the first multimodal features; The fourth fusion module is used to perform a fourth feature fusion based on a cross attention mechanism using the initial multimodal feature as Q and the static appearance time feature as K and V to obtain a second multimodal feature; The fifth fusion module is used to perform a fifth feature fusion based on a cross attention mechanism using the initial multimodal features as K and V and the motion optical flow time features as Q to obtain a third multimodal feature; The sixth fusion module is used to perform a sixth feature fusion based on a cross attention mechanism using the initial multimodal feature as Q and the motion optical flow time feature as K and V to obtain a fourth multimodal feature; The feature splicing module is used to perform feature splicing on the first multimodal feature, the second multimodal feature, the third multimodal feature and the fourth multimodal feature to obtain a multimodal splicing feature; The third feedforward neural network is used to determine the target fusion feature based on the multimodal splicing feature.

Citation Information

Patent Citations

  • Multi-target trajectory anomaly processing method and system in micro-manipulation

    CN112785630A

  • Video small target tracking method and device

    CN113269808A

  • Image processing method and device and computer readable storage medium

    CN113706577A

  • Multi-target tracking method based on spatial correlation and optical flow registration

    CN115100565A

  • Method for detecting cattle in complex cattle farm environment

    CN117935314A

Cited By

  • Zebra fish multi-target tracking method and system based on space science experiment video

    CN120894540A

  • Zebrafish multi-target tracking method and system based on spatial science experiment video

    CN120894540B