Dynamic feature space construction method for time dimension coupling multi-modal data

The dynamic feature space construction method coupled with the time dimension solves the problems of time synchronization and dynamic interaction in traditional multimodal data processing, improves the adaptability and robustness of the multimodal system, and is suitable for video description generation, cross-modal retrieval and autonomous driving scenario analysis.

CN120804659AInactive Publication Date: 2025-10-17CHENGDU YUNZHONGLE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510912739.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-10-17
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional multimodal data processing methods rely on a single time scale or static feature extraction, which makes it difficult to fully capture the temporal synchronization and dynamic interaction between modalities, limiting the expressive power and decision-making effect of multimodal systems.

Method used

A dynamic feature space construction method coupled with time dimension is adopted. Through time synchronization modeling, time granularity alignment and dynamic feature space construction, combined with local and global feature fusion, a dynamic shared feature space is generated to realize time synchronization and dynamic interaction between modalities.

Benefits of technology

It improves the adaptability and robustness of multimodal systems in complex scenarios, making them suitable for tasks such as video description generation, cross-modal retrieval, and autonomous driving scene analysis, while reducing processing complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120804659A_ABST
    Figure CN120804659A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic feature space construction method for time dimension coupling multi-modal data, which comprises the following steps: time synchronization modeling is realized through time dynamic correlation calculation and time step relation weight generation; time granularity alignment is realized through dynamic time warping optimization and time scale unification; and dynamic feature space construction is realized through global time feature extraction and local-global feature fusion. The invention discloses a dynamic feature space construction method for time dimension coupling multi-modal data, and the method comprises the steps: capturing a time synchronization relation between modals through dynamic correlation calculation and time step weight generation; unifying a multi-modal time resolution and a time step relationship through time granularity alignment; and finally, through fusion of local and global time features, a dynamic shared feature space is generated, and the multi-modal time dynamic state is comprehensively expressed.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and particularly relates to a dynamic feature space construction method for time-dimension coupled multi-modal data. BACKGROUND

[0002] In modern artificial intelligence applications, the processing of multi-modal data is the key to solving complex tasks, especially multi-modal data with time coupling relationship (such as video, audio, and text). The time dynamic alignment and feature fusion between modalities directly affect the system performance. Traditional multi-modal methods usually rely on a single time scale or static feature extraction strategy, which is difficult to fully capture the time synchronization and dynamic interaction between modalities, limiting the expression ability and decision effect of the multi-modal system. SUMMARY

[0003] In view of the above problems in the prior art, the present application provides a dynamic feature space construction method for time-dimension coupled multi-modal data, which aims to solve the problem that traditional multi-modal methods usually rely on a single time scale or static feature extraction strategy, which is difficult to fully capture the time synchronization and dynamic interaction between modalities, limiting the expression ability and decision effect of the multi-modal system.

[0004] In order to achieve the above application purpose, the technical scheme adopted by the present application is as follows:

[0005] A dynamic feature space construction method for time-dimension coupled multi-modal data, comprising the following steps:

[0006] S1: Time synchronization modeling, realized by time dynamic correlation calculation and time step relationship weight generation;

[0007] S2: Time granularity alignment: realized by dynamic time warping optimization and time scale unification;

[0008] S3: Dynamic feature space construction, realized by global time feature extraction and local-global feature fusion.

[0009] Further, the time dynamic correlation calculation is:

[0010] Let the time step features of modal A and modal B be and Then the time dynamic correlation degree R ij is defined as:

[0011] Wherein:

[0012] represents the similarity between modal features (such as dot product similarity):

[0013]

[0014] W(|t i -t j |) represents the influence of time difference on correlation:

[0015]

[0016] Further, time step relationship weight generation:

[0017] Based on the time dynamic correlation R ij , generate time step alignment weight α ij :

[0018]

[0019] Weight α ij represents the alignment degree of time step i of modality A and time step j of modality B.

[0020] Further, dynamic time warping optimization:

[0021] For the time step features of modality A and modality B, define the time step alignment path The alignment cost of the path is:

[0022]

[0023] Where:

[0024] Represents the matching cost between time step features.

[0025] Δ(i k , j k ) represents the continuity constraint of the path:

[0026] Δ(i k , j k ) = |(i k -i k-1 ) - (j k -j k-1 )|

[0027] λ controls the weight of matching cost and path smoothness.

[0028] Solve the optimal path P * by dynamic programming: Realize the optimal alignment of time steps between modalities.

[0029] Further, time scale unification:

[0030] Based on the optimal path P * of dynamic time warping, unify the time resolution of modality A and modality B, and align the time steps to the unified scale: ​

[0031] where Align denotes a time step alignment operation based on path P * .

[0032] Further, global temporal feature extraction:

[0033] Combining local temporal feature and global temporal feature is generated by attention weighted pooling:

[0034]

[0035] where the weight β t is generated by attention mechanism:

[0036]

[0037] is an importance score function of time step feature, which is usually defined as:

[0038]

[0039] where w, W, b are learnable parameters.

[0040] Further, local-global feature fusion:

[0041] Combining local temporal feature and global temporal feature generates dynamic shared feature

[0042]

[0043] where W1, W2 are learnable parameters, used to dynamically adjust the contribution ratio of local and global features.

[0044] The beneficial effects of the present application are:

[0045] The present invention provides a method for constructing a dynamic feature space for time-dimension coupled multimodal data. The method captures the temporal synchronization relationship between modalities through dynamic correlation calculation and time-step weight generation; unifies the multimodal temporal resolution and time-step relationship through time granularity alignment; and finally generates a dynamic shared feature space through the fusion of local and global temporal features to fully express the multimodal temporal dynamics. This mechanism does not require complex modality-specific processing, and integrates temporal dynamic relationships, multi-time scale information, and inter-modal feature interactions into a unified representation through mathematical models. When modal time step differences or synchronization anomalies occur, they are discovered and processed in a timely manner; when the modal relationship is stable, the dynamic feature expression is optimized to improve task performance. This mechanism effectively avoids the limitations of single-time-scale modeling and static feature extraction, takes into account both local modal dynamics and global trends, reduces the complexity of multimodal processing, and improves the adaptability and robustness of multimodal systems in complex scenarios. It is suitable for multi-task scenarios such as video description generation, cross-modal retrieval, and autonomous driving scene analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 This is a flow chart of the method for constructing a dynamic feature space of time-dimension coupled multimodal data according to the present invention. DETAILED DESCRIPTION

[0047] The technical solutions in the embodiments of the present invention are described clearly and completely below with reference to the accompanying drawings.

[0048] like Figure 1 As shown, a method for constructing a dynamic feature space of time-dimensional coupled multimodal data includes the following steps:

[0049] Time Synchronous Modeling:

[0050] Temporal synchronization modeling aims to quantify the interactions between multimodal data at time steps, capturing temporal dynamics and intermodal synchronization. This is achieved through temporal dynamic correlation calculations and time step relationship weight generation.

[0051] Time dynamic correlation calculation:

[0052] Assume that the time step characteristics of mode A and mode B are and Then the temporal dynamic correlation R ij Defined as:

[0053] in:

[0054] Represents the similarity between modal features (such as dot product similarity):

[0055]

[0056] W(|ti -t j |) represents the influence of time difference on correlation:

[0057]

[0058] Time step relationship weight generation:

[0059] Based on time dynamic correlation R ij , generate time step alignment weight α ij :

[0060]

[0061] Weight α ij represents the alignment degree of time step i of modality A and time step j of modality B.

[0062] Time granularity alignment:

[0063] The goal of time granularity alignment is to unify the time resolution of different modalities and ensure the alignment relationship across modalities time steps. It is mainly achieved through dynamic time warping optimization and time scale unification.

[0064] Dynamic time warping optimization:

[0065] For the time step features of modality A and modality B, define the time step alignment path The alignment cost of the path is:

[0066]

[0067] Where:

[0068] represents the matching cost between time step features.

[0069] Δ(i k , j k ) represents the continuity constraint of the path:

[0070] Δ(i k , j k ) = |(i k -i k-1 ) - (j k -j k-1 )|

[0071] λ controls the weight of matching cost and path smoothness.

[0072] Solve the optimal path P * by dynamic programming: Achieve the optimal alignment of time steps between modalities.

[0073] Time scale unification:

[0074] Optimal path P based on dynamic time warping * , unify the time resolution of modal A and modal B, align the time steps to a unified scale:

[0075] where Align denotes the time step alignment operation based on path P * .

[0076] Dynamic feature space construction:

[0077] The goal of dynamic feature space is to combine local temporal features and global temporal features to generate a unified shared representation. This is mainly achieved through global temporal feature extraction and local-global feature fusion.

[0078] Global temporal feature extraction:

[0079] Combine local temporal features and global temporal features to generate:

[0080]

[0081] where the weight β t is generated by the attention mechanism:

[0082]

[0083] is the importance score function of the time step feature, which is usually defined as:

[0084]

[0085] where w, W, b are learnable parameters.

[0086] Local-global feature fusion:

[0087] Combine local temporal features and global temporal features to generate dynamic shared features

[0088]

[0089] where W1, W2 are learnable parameters, used to dynamically adjust the contribution ratio of local and global features.

[0090] The application is a multi-modal time dynamic processing mechanism based on time synchronization modeling, time granularity alignment and dynamic feature space construction. Through dynamic correlation calculation and time step weight generation, the time synchronization relationship between modalities is captured. Through time granularity alignment, the multi-modal time resolution and time step relationship are unified. Finally, through the fusion of local and global time features, a dynamic shared feature space is generated to fully express the multi-modal time dynamics.

[0091] The mechanism does not require complex modality-specific processing, and integrates time dynamic relationship, multi-time scale information and inter-modal feature interaction into a unified representation through a mathematical model.

[0092] When the modal time step difference or synchronization is abnormal, it is timely discovered and processed. When the modality relationship is stable, the dynamic feature expression is optimized and the task performance is improved.

[0093] The mechanism effectively avoids the limitations of single time scale modeling and static feature extraction, takes into account the local dynamics and global trend of modalities, reduces the complexity of multi-modal processing, and improves the adaptability and robustness of multi-modal systems in complex scenarios. It is suitable for multi-task scenarios such as video description generation, cross-modal retrieval, automatic driving scene analysis, etc.

[0094] The above only describes the preferred embodiments of the application patent and does not limit the application patent. Any modification, equivalent replacement and improvement made within the spirit and principle of the application patent should be included in the protection scope of the application patent.

Claims

1. A method for constructing a dynamic feature space of time-dimensional coupled multimodal data, characterized in that: The following steps are involved: S1: Time synchronization modeling, achieved through time dynamic association calculation and time step relationship weight generation; S2: Time granularity alignment: achieved through dynamic time warping optimization and time scale unification; S3: Dynamic feature space construction: achieved through global temporal feature extraction and local-global feature fusion.

2. The method for constructing a dynamic feature space of time-dimension coupled multimodal data according to claim 1, characterized in that: Time dynamic correlation calculation: Assume that the time step characteristics of mode A and mode B are and Then the temporal dynamic correlation R ij Defined as: in: Represents the similarity between modal features (such as dot product similarity): W(|t i -t j |) represents the effect of time difference on correlation:

3. The method for constructing a dynamic feature space of time-dimension coupled multimodal data according to claim 1, characterized in that: Time step relationship weight generation: Based on the time dynamic correlation R ij , generate time step alignment weight α ij : Weight α ij Indicates how closely time step i of mode A is aligned with time step j of mode B.

4. The method for constructing a dynamic feature space of time-dimension coupled multimodal data according to claim 1, characterized in that: Dynamic Time Warping Optimization: For the time step characteristics of mode A and mode B, define the time step alignment path The alignment cost of the path is: in: Represents the matching cost between time step features. Δ(i k ,j k ) represents the continuity constraint of the path: Δ(i k ,j k )=|(i k -i k-1 )-(j k -j k-1 )| λ controls the weight between matching cost and path smoothness. Solve the optimal path P through dynamic programming * : Achieve optimal alignment of time steps between modalities.

5. The method for constructing a dynamic feature space of time-dimension coupled multimodal data according to claim 1, characterized in that: Time scale uniformity: Optimal path P based on dynamic time warping * , unify the time resolution of mode A and mode B and align the time steps to a unified scale: Where Align represents the path P * The time step alignment operation.

6. The method for constructing a dynamic feature space of time-dimension coupled multimodal data according to claim 1, characterized in that: Global temporal feature extraction: Combined with local temporal features and global temporal characteristics Generated by attention-weighted pooling: Among them, the weight β t Generated by the attention mechanism: is the importance scoring function of the time step feature, usually defined as: Where w, W, b are learnable parameters.

7. The method for constructing a dynamic feature space of time-dimension coupled multimodal data according to claim 1, characterized in that: Local-global feature fusion: Combined with local temporal features and global temporal characteristics Generate dynamic shared features Among them, W1 and W2 are learnable parameters used to dynamically adjust the contribution ratio of local and global features.