Multi-modal multi-target tracking method guided by adaptive key frame mining and space-time diagram learning

Through adaptive keyframe mining and space-time graph learning-guided multi-objective tracking methods, combined with multi-modal information fusion and reinforced learning video segmentation, the identification difficulties caused by occlusion and target similarity in multi-objective tracking are solved, and the tracking accuracy and robustness are significantly improved.

CN120070506APending Publication Date: 2025-05-30ANHUI UNIV

Patent Information

Application Number
CN202510241766.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing multi-target tracking technology is prone to ID switching and target error problems caused by occlusion in crowded people, especially when the targets look similar and have similar positions, it is difficult to effectively solve them.

Method used

Adaptive keyframe mining and space-time graph learning-guided multi-objective tracking method is adopted, and visible light and thermal infrared image information are fused through feature fusion modules, video adaptive segmentation is used to use reinforcement learning, and inter-frame target feature extraction is performed through intra-frame feature fusion modules and SUSHI blocks to solve the problems of occlusion and similar appearance.

Benefits of technology

It significantly improves the accuracy and robustness of multi-target tracking, effectively alleviates the identification difficulties caused by occlusion and target similarity, and improves the tracking effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070506A_ABST
    Figure CN120070506A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal multi-target tracking method guided by adaptive key frame mining and space-time diagram learning, and the method comprises the steps: obtaining all frame images of a video segment, inputting a visible light image and a thermal infrared image corresponding to the same frame image into a feature fusion module, and generating embedding; performing information fusion among multiple modes by using cross attention to obtain a multi-mode fusion feature; carrying out adaptive video segmentation on the video through a key frame extraction module; the key frame extraction module continuously iterates an optimal segmentation strategy and an optimal reward score in a learning process based on a reinforcement learning method; and repeatedly inputting the adaptively divided video sequence into the intra-frame feature fusion module and the SUSHI block to obtain a final tracking result. According to the method, the thermal infrared image is used for making up for the deficiency of single-mode information, and video segmentation is carried out through reinforcement learning self-adaption to solve the IDS problem; the time relation of inter-frame targets is mined by using an SUSHI module, and the spatial relation of intra-frame targets is mined by using an IFF module, so that the problems of shielding and similar appearance are further solved, and the tracking effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to target tracking and deep learning technologies, and specifically to a multi-modal multi-object tracking method guided by adaptive key frame mining and spatio-temporal graph learning. Background Art

[0002] Multi-object tracking (MOT) aims to simultaneously track multiple objects in a video, maintain the identities of the objects, and generate their motion trajectories. It has a wide range of applications in different fields such as video surveillance, autonomous driving, and video analysis.

[0003] Although many methods have been proposed for multi-object tracking, ID switches (IDS) caused by frequent occlusions in crowded crowds and incorrect object tracking due to similar appearances and close positions of two objects remain major challenges, as Figure 6 and Figure 7 shown. Summary of the Invention

[0004] Object of the Invention: The object of the present invention is to solve the deficiencies existing in the prior art and provide a multi-modal multi-object tracking method guided by adaptive key frame mining and spatio-temporal graph learning.

[0005] Technical Solution: A multi-modal multi-object tracking method guided by adaptive key frame mining and spatio-temporal graph learning according to the present invention includes the following steps:

[0006] Step 1: Obtain all frame images of a video segment, input the visible light image and the thermal infrared image corresponding to the same frame image into a feature fusion module FFM, extract features from the visible light image and the thermal infrared image respectively to generate embeddings; then use cross-attention to fuse the generated embeddings of the two modalities for inter-modal information to obtain a multi-modal fusion feature f (using multi-modal information to make the features of the object more comprehensive); input the multi-modal fusion features f of all frames into a key frame extraction module KFE;

[0007] Step 2: Based on the multi-modal fusion features f of all frames, adaptively segment the video through a key frame extraction module KFE;

[0008] The key frame extraction module KFE is based on the Q-learning method of reinforcement learning and includes an action selection, a reward mechanism, an active exploration intensity, and an update strategy of a Q-table, and continuously iterates the optimal segmentation strategy FS best and the corresponding optimal reward score K best ;

[0009] Select the first frame action Γ of the video segment through the action selection strategy aAnd the tail frame action Γ b , and determine future actions; the reward mechanism uses the feature difference between the first frame and the tail frame of the obtained video segment as the reward value to determine the direction of model optimization:

[0010] Until the video is completely segmented in the m-th iteration, the final optimal reward score K is obtained best , and select the optimal segmentation sequence;

[0011] Step 3: Repeatedly input the video sequence adaptively segmented in Step 2 into the intra-frame feature fusion module IFF module and the SUSHI block N times to obtain the final tracking result;

[0012] First, the intra-frame feature fusion module IFF module extracts the intra-frame target features based on the graph convolutional network, and then the SUSHI block extracts the inter-frame target features.

[0013] Furthermore, the implementation process of the feature fusion module FFM in Step 1 is as follows:

[0014]

[0015] In the above formula, d k is the feature dimension of the key Q, value V, and query K; the calculation formulas of Q, V, and K are as follows:

[0016] In the above formula, P RGB refers to the visible light image, P T refers to the thermal infrared image, and the implementation formula of the ECM module is as follows:

[0017] ECM(P) = HardSwish(BN(Conv1(HardSwish(BN(Conv2(P))))))

[0018] In the above formula, HardSwish is the activation function, BN is the normalization function, Conv1 is Conv 1x1, and Conv2 is Conv3x3;

[0019] Finally, the expression of the multi-modal fusion feature f is as follows:

[0020] f = FFM(Q, V, K).

[0021] Furthermore, the key frame extraction module KFE determines the first frame action Γ of the video segment through the action selection strategy a and the process of determining the tail frame action Γ of the video segment b is the same. The selection process of the first frame action Γ a is as follows:

[0022] First, set Γ as the action selection, record the future choices through Γ, set QT as the Q-value, record the total expected return obtained by taking a specific action in a certain state through QT, set F as the state, and QT(F,Γ) as the expected return of the future action selection Γ in state F, where Γ ∈ Γ range ; Then, implement the future action selection through the following formula:

[0023]

[0024] Among them, represents the state the next selection in the expected return, the superscript i refers to the current state, and the superscript i + 1 refers to the next selection, represents Γ a the value range, η ∈ (0,1), i ∈ (0,N), and ε is the active exploration rate (i.e., the active exploration intensity);

[0025] The key frame extraction module KFE uses the feature difference between the first and last frames of video segments as the reward value, and the calculation formula is as follows:

[0026]

[0027] In the above formula, k i+1 is the reward for the (i + 1)-th segmentation result of this round, is the multi-modal fusion feature of the first frame of the i-th video segment, is the multi-modal fusion feature of the last frame of the i-th video segment, is the multi-modal fusion feature of the first frame of the (i + 1)-th video segment, is the multi-modal fusion feature of the last frame of the (i + 1)-th video segment; σ, ξ are constants, and φ is the cosine similarity;

[0028] In the key frame extraction module KFE, the Q-value is related to F i , F i+1 , the learning rate q, the reward Γ, and the discount factor α. The update strategy of the Q-table is as follows:

[0029]

[0030] QT[F i ,Γ] = λ × q + QT[F i ,Γ]

[0031] is the optimal action selection for the next time, λ = κ i + α × QT[F i+1 ,Γ best -QT[F i ,Γ];

[0032] Until the video is completely segmented in the m-th iteration, the optimal reward score κ of this round of iteration sum The formula is as follows:

[0033]

[0034] Finally, select the optimal segmentation sequence through the reward value:

[0035]

[0036] Repeat the above video adaptive segmentation steps M times, where M is a constant.

[0037] Furthermore, before using SUSHI blocks for different-level fusion in step 3, first construct a graph G=(V, E) for the current level, where the target is represented by a node v i ∈V, and the edge e i represents the hypothesis for judging the relationship between nodes;

[0038] In the case where the target is occluded, the graph convolutional GCN of IFE supplements the feature representation of the occluded target by fusing the information of the surrounding unoccluded targets, which helps to alleviate the impact of occlusion on the recognition performance. When the targets are too similar and close in position, GCN can utilize the small differences between them (such as spatial position, relative size, etc.) and the context information of the surrounding environment to distinguish these targets. Thus, the recognition accuracy is improved.

[0039] Therefore, the graph method GCN is selected for feature fusion. As Figure 5 shown, the feature is obtained through formula (10)

[0040]

[0041] After the video is adaptively segmented, repeating the divided video sequence through the IFF module and SUSHI blocks N times can obtain the final tracking result.

[0042] Beneficial effects: The present invention uses the thermal infrared map to make up for the deficiency of single-modal information, uses reinforcement learning to adaptively segment the video to solve the IDS problem; uses the SUSHI module to mine the temporal relationship between targets in frames, and uses the IFF module to mine the spatial relationship between targets in frames to further solve the problems of occlusion and similar appearance, so as to improve the tracking effect. Description of the Drawings

[0043] Figure 1 is the overall framework diagram of the present invention;

[0044] Figure 2 is the structural schematic diagram of the FFM module of the present invention;

[0045] Figure 3 It is a schematic diagram of the ECM structure in the FFM module of the present invention;

[0046] Figure 4 It is a schematic diagram of the KFE module of the present invention;

[0047] Figure 5 It is a schematic diagram of the IFF module of the present invention;

[0048] Figure 6 It is a diagram showing the problems handled by the prior art in the embodiment;

[0049] Figure 7 It is a diagram showing the problems handled by the prior art in the embodiment. Detailed implementation manners

[0050] The technical solution of the present invention will be described in detail below, but the protection scope of the present invention is not limited to the described embodiments.

[0051] The present invention significantly improves the accuracy of tracking association by making full use of multi-modal information, short-term and long-term associations, the structured spatial relationship between different targets, and the temporal relationship of targets in different frames.

[0052] As Figure 1 shown, the network model of the present invention includes three core modules: a feature fusion module FFM, a key frame extraction module KFE, and an intra-frame feature fusion module IFF; taking multi-modal information as input, the KFE module based on reinforcement learning adaptively segments the video, guiding the tracker to deeply explore the internal logic in the video, thereby significantly enhancing the association effect. In addition, the present invention also uses GCN to capture the spatial relationship between targets within a frame, effectively solving the problems of target occlusion and difficult distinction between targets with similar appearances or similar targets.

[0053] The present invention specifically includes the following steps:

[0054] Step 1: Obtain all frame images of the video segment, input the visible light image and the thermal infrared image corresponding to the same frame image into the feature fusion module FFM, extract features from the visible light image and the thermal infrared image respectively to generate embeddings; then use cross-attention to fuse the generated embeddings of the two modalities to obtain multi-modal fusion feature f (using multi-modal information makes the features of the target more comprehensive); input the multi-modal fusion feature f of all frames into the key frame extraction module KFE together;

[0055] Step 2: Based on the multi-modal fusion feature f of all frames, the key frame extraction module KFE adaptively segments the video;

[0056] The key frame extraction module KFE is based on the reinforcement learning Q-learning method, including action selection, reward mechanism, active exploration intensity, and the update strategy of the Q-table. During the learning process, it continuously iterates the optimal segmentation strategy FS best and the corresponding optimal reward score K best ;

[0057] Select the first-frame action Γ of the video segment and the last-frame action Γ of the video segment through the action selection strategy, and determine future actions; the reward mechanism uses the feature difference between the first frame and the last frame of the obtained video segment as the reward value to determine the direction of model optimization: a and the last-frame action Γ b , and determine future actions; the reward mechanism uses the feature difference between the first frame and the last frame of the obtained video segment as the reward value to determine the direction of model optimization:

[0058] Until the video is completely segmented in the m-th iteration, the final optimal reward score K is obtained best , and select the optimal segmentation sequence;

[0059] Step 3: Repeatedly input the video sequence adaptively segmented in Step 2 into the intra-frame feature fusion module IFF module and the SUSHI block N times to obtain the final tracking result;

[0060] First, the intra-frame feature fusion module IFF module extracts intra-frame target features based on the graph convolutional network, and then the SUSHI block extracts inter-frame target features.

[0061] As Figure 2 and Figure 3 shown, the implementation process of the feature fusion module FFM in Step 1 of this embodiment is as follows:

[0062]

[0063] In the above formula, d k is the feature dimension of the key Q, value V, and query K; the calculation formulas of Q, V, and K are as follows:

[0064] In the above formula, P RGB refers to the visible light image, P T refers to the thermal infrared image, and the implementation formula of the ECM module is as follows:

[0065] ECM(P) = HardSwish(BN(Conv1(HardSwish(BN(Conv2(P))))))

[0066] In the above formula, HardSwish is the activation function, BN is the normalization function, Conv1 is Conv 1x1, and Conv2 is Conv3x3;

[0067] Finally, the expression of the multi-modal fusion feature f is as follows:

[0068] f = FFM(Q, V, K).

[0069] From Figure 3 It can be seen that the same set of visible light images and thermal infrared images are both operated on by the ECM and then processed using the attention mechanism; after processing the visible light, we get V and K, and for the infrared, it is Q.

[0070] As Figure 4 shown, the key frame extraction module KFE determines the first-frame action Γ of the video segment through the action selection strategy a and determines the last-frame action Γ of the video segment b in the same process. The selection process of the first-frame action Γ a is as follows:[[]]

[0071] First, set Γ as the action selection. Record the future choices through Γ. Let QT be the Q value, and record the total expected return obtained by taking a specific action in a certain state through QT. Let F represent the state, and QT(F, Γ) be the expected return of the future action selection Γ in state F, where Γ ∈ Γ range ; then, the future action selection is achieved through the following formula:[[]]

[0072]

[0073] where,[[]] represents the expected return of the next selection in the state . The superscript i refers to the current state, and the superscript i + 1 refers to the next selection.[[]] represents the value range of Γ a . η ∈ (0, 1), i ∈ (0, N), and ε is the active exploration rate.[[]]

[0074] The key frame extraction module KFE uses the feature difference between the first and last frames of the video segment as the reward value, and the calculation formula is as follows:[[]]

[0075]

[0076] In the above formula, κ i+1 is the reward for the (i + 1)-th segmentation result in this round,[[]] is the multi-modal fusion feature of the first frame of the i-th video segment,[[]] is the multi-modal fusion feature of the last frame of the i-th segment,[[]] is the multi-modal fusion feature of the first frame of the i-th video segment,[[]] is the multi-modal fusion feature of the last frame of the i-th segment;[[]]

[0077] σ and ξ are constants,[[]] is the cosine similarity;[[]]

[0078] In the key frame extraction module KFE, the Q value and Fi , F i+1 , learning rate q, reward Γ, and discount factor α. The update strategy of the Q - table is as follows:

[0079]

[0080] QT[F i , Γ] = λ×q + QT[F i , Γ]

[0081] is the optimal action selection for the next time, λ = k i + α×QT[F i+1 , Γ best -QT[F i , Γ];

[0082] Until the video is completely segmented in the m - th iteration, the final optimal reward score k sum of this round of iteration is given by the following formula:

[0083]

[0084] Finally, the optimal segmented sequence is selected through the reward value:

[0085]

[0086] Repeat the above video adaptive segmentation steps M times, where M is a constant.

[0087] Before using SUSHI blocks for different - level fusion in step 3 of this embodiment, first construct a graph G=(V, E) for the current level. The target is represented by a node v i ∈V, and the edge e i represents the hypothesis for judging the relationship between nodes;

[0088] As Figure 5 shown, in the case where the target is occluded, GCN supplements the feature representation of the occluded target by fusing the information of the surrounding un - occluded targets, which helps to alleviate the impact of occlusion on the recognition performance. When the targets are too similar and close in position, GCN can use the subtle differences between them (such as spatial position, relative size, etc.) and the context information of the surrounding environment to distinguish these targets. Thus, the recognition accuracy is improved.

[0089] Therefore, the graph - based method GCN is selected for feature fusion. As Figure 5 shown, the feature

[0090]

[0091] After the video is adaptively segmented, repeating the divided video sequence through the IFF module and the SUSHI block N times can obtain the final tracking result.

[0092] To further verify the technical effect of the present invention, all frames of a certain video segment are input into the network model of the present invention. The time intervals of the marked target states from appearance, occlusion to reappearance are not fixed. The key frame extraction module KFE is used for adaptive segmentation, which can better explore the internal logic of the video to obtain a more appropriate short sequence. In this embodiment, the two marked targets have similar appearances and close positions. The intra-frame feature fusion module IFF uses GCN to supplement the information of surrounding targets to the target, which can increase the difference between the two and thus distinguish them.

Claims

1. A multi-modal multi-target tracking method guided by adaptive keyframe mining and spatiotemporal graph learning, characterized in that: The following steps are involved: Step 1: Get all the frame images of the video segment, input the visible light image and thermal infrared image corresponding to the same frame image into the feature fusion module FFM, extract features of the visible light image and thermal infrared image, and generate embeddings respectively; then use cross attention to fuse the generated embeddings of the two modalities to obtain multimodal information, and obtain multimodal fusion features f; input the multimodal fusion features f of all frames into the key frame extraction module KFE; Step 2: Based on the multimodal fusion features f of all frames, the video is adaptively segmented through the key frame extraction module KFE; The key frame extraction module KFE is based on the reinforcement learning Q-learning method, including action selection, reward mechanism, active exploration intensity and Q table update strategy, and continuously iterates the optimal segmentation strategy FS during the learning process best And the corresponding optimal reward score K best ; The first frame action Γ of the video segment is selected by the action selection strategy a and the last frame action Γ b , and determine future actions; the reward mechanism uses the feature difference between the first frame and the last frame of the obtained video segment as the reward value to determine the direction of model optimization: Until the video is completely segmented in the mth iteration, the final optimal reward score K is obtained best , and select the optimal sequence of sub-bids; Step 3, repeatedly input the video sequence adaptively divided in step 2 into the intra-frame feature fusion module IFF module and SUSHI block N times to obtain the final tracking result; First, the intra-frame feature fusion module (IFF) extracts intra-frame target features based on the graph convolutional network, and then the SUSHI block extracts inter-frame target features.

2. The multimodal multi-target tracking method guided by adaptive keyframe mining and spatiotemporal graph learning according to claim 1 is characterized in that: The implementation process of the feature fusion module FFM in step 1 is as follows: In the above formula, d k is the feature dimension of key Q, value V and query K; the calculation formulas of Q, V and K are as follows: In the above formula, P RGB refers to the visible light image, P T Refers to thermal infrared images. The implementation formula of the ECM module is as follows: ECM(P)=HardSwish(BN(Conv1(HardSwish(BN(Conv2(P)))))) In the above formula, HardSwish is the activation function, BN is the normalization function, Conv1 is Conv 1x1, and Conv2 is Conv3x3; Finally, the expression of the multimodal fusion feature f is as follows: f=FFM(Q,V,K).

3. The multimodal multi-target tracking method guided by adaptive keyframe mining and spatiotemporal graph learning according to claim 1 is characterized in that: The key frame extraction module KFE determines the first frame action Γ of the video segment through the action selection strategy a and determine the video segment end frame action v b The process is the same as that of the first frame action Γ a The selection process is as follows: First, let Γ be the action selection, and use Γ to record the choices to be made in the future. QT is the Q value, and QT is used to record the total return expected to be obtained by taking a specific action in a certain state. F represents the state, and QT(F,Γ) is the expected return of the future action selection Γ in state F, where Γ∈Γ range ; Then, the future action selection is realized by the following formula: in, Display status Next option The expected return, superscript i refers to the current state, superscript i+1 refers to the next choice, Represents Γ a The value range of , η∈(0,1), i∈(0,N), ε is the active exploration rate (i.e., the intensity of active exploration); The key frame extraction module KFE uses the feature difference between the first and last frames of the video segment as the reward value. The calculation formula is as follows: In the above formula, κ i+1 The reward for the i+1th segment result in this round. is the multimodal fusion feature of the first frame of the i-th video, is the multimodal fusion feature of the last frame of the i-th segment, is the multimodal fusion feature of the first frame of the i+1th video, is the multimodal fusion feature of the last frame of the i+1th segment; σ, ξ are constants, and φ is the cosine similarity; In the key frame extraction module KFE, the Q value and F i 、F i+1 , learning rate q, reward Γ and discount factor α, the update strategy of Q table is as follows: QT[F i ,Γ]=λ×q+QT[F i ,Γ] For the next optimal action selection, λ=κ i +α×QT[F i+1 ,Γ best ]-QT[F i ,Γ]; Until the video is completely segmented in the mth iteration, the final optimal reward score κ of this iteration sum The formula is as follows: Finally, the optimal segment sequence is selected based on the reward value: Repeat the above video adaptive segmentation M times, where M is a constant.

4. The multimodal multi-target tracking method guided by adaptive keyframe mining and spatiotemporal graph learning according to claim 1, characterized in that: Before using the SUSHI block to fuse different levels in step 3, a graph G = (V, E) is constructed for the current level, with the target node v i ∈V means that edge e i Represents the assumptions about the relationship between nodes; When the target is occluded, IFE's graph convolutional GCN supplements the feature representation of the occluded target by fusing the information of the surrounding unoccluded targets, which helps alleviate the impact of occlusion on recognition performance. When the targets are too similar and close to each other, GCN can use the slight differences between them (such as spatial position, relative size, etc.) and the contextual information of the surrounding environment to distinguish these targets, thereby improving recognition accuracy. Therefore, the graph method GCN is used to perform feature fusion, as shown in Figure 5. The feature is obtained by formula (10) After the video is adaptively segmented, the segmented video sequence is repeated with the IFF module and the SUSHI block N times to obtain the final tracking result.

Citation Information

Patent Citations

  • Fusion method of visible light and infrared images based on focusing loss function constraint

    CN117115065A

  • Target tracking system and method fusing multi-modal features

    CN117522923A

Cited By

  • Self-learning multi-target tracking method and system based on cross-modal perception

    CN120997253A

  • Self-learning multi-target tracking method and system based on cross-modal perception

    CN120997253B