Method and System for Labeling Dynamic and Static Detection Data Based on Time Key Association

By aligning and processing multimodal data based on time key correlation, using neural network and Transformer model to extract features, combined with attention mechanism and dynamic weight adjustment, the problem that traditional annotation methods are difficult to capture modal association is solved, and efficient and accurate annotation results are achieved.

CN119830229BActive Publication Date: 2025-06-13SICHUAN INST OF PROD QUALITY SUPERVISION INSPECTION & TESTING (SICHUAN QUALITY & TECH REVIEW & EVALUATION CENT) +2
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510330292.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-13
Estimated Expiration
2045-03-20

AI Technical Summary

Technical Problem

Traditional annotation methods are difficult to effectively capture the correlation characteristics between dynamic and static modal data, resulting in low labeling efficiency, high cost and unstable quality.

Method used

Through a time-key correlation method, all modal data are aligned to a unified timeline, modal data across modal time series are generated, and multimodal features are extracted using neural network and Transformer model, combining attention mechanism and dynamic modal weight adjustment, and high-quality annotation results are output.

Benefits of technology

It realizes the rapid generation of high-precision and low-cost annotation results, significantly improving the labeling efficiency and consistency, and is suitable for scenarios such as industrial testing and medical testing that require efficient multimodal data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119830229B_ABST
    Figure CN119830229B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of data processing and relates to a method and system for annotating dynamic and static detection data based on time key association. The method includes aligning all modal data to a unified time axis; performing standardized alignment through linear mapping processing to generate a unified feature dimension; extracting multi-modal feature data through feature extraction; extracting global time correlation features of the multi-modal feature data; extracting interaction features between modalities based on an attention mechanism; and outputting an annotation result based on the interaction features, and updating the annotation result by dynamically adjusting the weights and biases of the output layer according to the input interaction features. The present invention solves the problems of time synchronization and insufficient feature expression in multi-modal data annotation through time series feature modeling, inter-modal association alignment, and feature dynamic optimization; and can quickly generate high-precision and low-cost annotation results by dynamicizing static modalities and unifying time series features, combined with cross-modal feature fusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data processing, and more specifically, relates to a method and system for annotating dynamic and static detection data based on time key association. Background Art

[0002] In information-based detection tasks, automatic annotation of multi-modal data is a key link to improve data processing efficiency and model performance. Automatic annotation mainly targets information such as the category, location, state, and behavior characteristics of target objects in multi-modal data, aiming to provide high-quality supervised data for subsequent data analysis and model training, while reducing the cost of manual intervention and improving the efficiency and accuracy of tasks. In the face of complex scenarios where dynamic modalities (such as videos, audio, sensor data) and static modalities (such as images) coexist, traditional annotation methods usually rely on manual processing or single-modal strategies, making it difficult to effectively capture the correlation characteristics between modalities, resulting in low annotation efficiency, high cost, and unstable quality. Summary of the Invention

[0003] To solve the above technical problems, the present invention provides a method and system for annotating dynamic and static detection data based on time key association.

[0004] In a first aspect, the present invention provides a method for annotating dynamic and static detection data based on time key association, including:

[0005] Align all modal data to a unified time axis to generate modal data of cross-modal time series;

[0006] Encode all modal data into high-dimensional features and perform standardized alignment through linear mapping processing to generate a unified feature dimension;

[0007] Use a neural network to extract features from each modal data to obtain multi-modal feature data;

[0008] Use a time series model based on Transformer to extract global time correlation features of multi-modal feature data;

[0009] Extract interaction features between modalities based on the attention mechanism;

[0010] Output an annotation result based on the interaction features, set dynamic modal weights, assign learned weights to the features of each modality, and update the annotation result by dynamically adjusting the weights and biases of the output layer according to the input interaction features.

[0011] In a second aspect, the present invention provides a system for annotating dynamic and static detection data based on time key association, including an alignment unit, an encoding and alignment unit, a first feature extraction unit, a second feature extraction unit, an interaction feature extraction unit, and an output and update unit;

[0012] An alignment unit for aligning all modal data to a unified timeline to generate modal data of cross-modal time series;

[0013] An encoding and alignment unit for encoding all modal data into high-dimensional features and performing standardized alignment through linear mapping processing to generate a unified feature dimension;

[0014] A first feature extraction unit for extracting features from each modal data using a neural network to obtain multi-modal feature data;

[0015] A second feature extraction unit for extracting global time correlation features of multi-modal feature data using a time series model based on Transformer;

[0016] An interaction feature extraction unit for extracting interaction features between modalities based on an attention mechanism;

[0017] An output and update unit for outputting an annotation result based on the interaction feature, setting dynamic modal weights, assigning learned weights to the features of each modality, and dynamically adjusting the weights and biases of the output layer according to the input interaction feature to update the annotation result.

[0018] Based on the above technical solutions, the present invention can also be improved as follows.

[0019] Further, before aligning all modal data to a unified timeline, it further includes: performing time serialization processing on static modal data, and converting discrete static data into continuous time series features through time alignment, interpolation filling, and feature generation; for dynamic modal data, distributing features to a unified timeline according to time windows.

[0020] Further, the time serialization processing of static data includes:

[0021] Using an interpolation method to fill a preset time window lacking picture data;

[0022] When the picture acquisition interval is greater than a set value, generating transition features through a generative adversarial network or generating transition features through time-aware linear interpolation, so that discrete static data is converted into continuous time series features.

[0023] Further, considering the influence of time context on interpolation, selecting a nearest neighbor interpolation algorithm combined with time weights, setting the time as , in the picture features obtained from the time difference , the time The nearest known picture feature before is , is The moment before, At the moment after the nearest known image feature after , is the weight of is the weight of, then:

[0024] ;

[0025] ;

[0026] ;

[0027] When the image acquisition interval is greater than the set value, generate transitional features through a generative adversarial network. Let the adversarial network be , then:

[0028] ;

[0029] Or generate transitional features through time-aware linear interpolation, then:

[0030] .

[0031] Furthermore, align all modal data to a unified time axis to generate modal data of cross-modal time series, including: for dynamic modal data, allocate the features of dynamic modal data to the unified time axis according to time windows; for static modal data, map the static modal data into the time series through a time serialization method.

[0032] Furthermore, encode all modal data into high-dimensional features and perform standardized alignment through linear mapping processing to generate a unified feature dimension, including:

[0033] For image modal data, introduce a time-aware self-attention mechanism through a convolutional neural network or a vision transformer to extract static features of the image modality;

[0034] For video modal data, adopt a separable spatio-temporal convolution and a multi-head spatio-temporal attention mechanism to jointly extract the time information and spatial information of the video;

[0035] For audio modal data, introduce a time-aware self-attention mechanism through a convolutional neural network or a transformer-based network to extract the frequency-domain features and time-domain features of the audio modality;

[0036] For sensor modal data, extract the time series features of the sensor modality through a one-dimensional convolutional network or a recurrent neural network.

[0037] Furthermore, a time series model based on Transformer is adopted to extract the global time correlation features of multi-modal feature data, including: Let the unified aligned multi-modal feature matrix be , be the output of the time series modeling module, represent real numbers, represent the total number of time steps, represent the feature dimension, represent the time series model, then:

[0038] .

[0039] Furthermore, the interaction features between modalities are extracted based on the attention mechanism, including: Let represent the query vector of the features of the target modality at the current time step, represent the key vector of the features of the historical time step, represent the eigenvalue of the candidate modality, be the output of the time series modeling module, and the interaction features between modalities are , be the attention module, then:

[0040] .

[0041] Furthermore, based on the interaction features, the annotation results are output, and dynamic modality weights are set to assign learned weights to the features of each modality, and the weights and biases of the output layer are dynamically adjusted according to the input interaction features to update the annotation results, including: Let represent the number of modalities, represent the total number of modalities, represent the time step, represent the weight of each modality, represent the annotation result, represent the th modality's fused features at the time step, represent the weight matrix of the output layer, represent the bias term of the output layer, represent the weight matrix of the output layer at time step , represent the bias term of the output layer at time step , represent the dynamic generation function of the weight matrix, represent the dynamic generation function of the bias term. For different time steps and different input features, a dynamic weight adjustment mechanism is introduced to enable the output layer to adaptively optimize the weights and biases, then:

[0042] ;

[0043] ;

[0044] 。

[0045] The beneficial effects of the present invention are as follows: The present invention proposes an automated annotation method. By means of temporal feature modeling, inter-modal correlation alignment, and feature dynamic optimization, it solves the problems of time synchronization and insufficient feature expression in multi-modal data annotation. By dynamicizing static modalities and unifying time series features, combined with cross-modal feature fusion, it can quickly generate high-precision and low-cost annotation results, realizing the intelligence and automation of the annotation process. This method requires little manual intervention, can adapt to the time step differences and complex dynamic relationships of different modal data, significantly improves the annotation efficiency and consistency, and is applicable to scenarios such as industrial inspection and medical inspection that require efficient multi-modal data processing, providing high-quality automated annotation support for information-based inspection systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 It is a schematic diagram of the method for annotating dynamic and static detection data based on time key association provided in Embodiment 1 of the present invention;

[0047] Figure 2 It is a system block diagram of the system for annotating dynamic and static detection data based on time key association provided in Embodiment 2 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations.

[0049] Embodiment 1

[0050] As an embodiment, as shown in the appendix Figure 1 To solve the above technical problems, this embodiment provides a method for annotating dynamic and static detection data based on time key association, including:

[0051] Align all modal data to a unified time axis to generate modal data of cross-modal time series;

[0052] Encode all modal data into high-dimensional features and perform standardized alignment through linear mapping processing to generate a unified feature dimension;

[0053] Use a neural network to extract features from each modal data to obtain multi-modal feature data;

[0054] Adopt a Transformer-based time series model to extract the global time correlation features of multimodal feature data;

[0055] Extract the interaction features between modalities based on the attention mechanism;

[0056] Output the annotation result based on the interaction features, set dynamic modality weights, assign learned weights to the features of each modality, and update the annotation result by dynamically adjusting the weights and biases of the output layer according to the input interaction features.

[0057] Optionally, before aligning all modality data to a unified time axis, it further includes: performing time serialization processing on static modality data, and converting discrete static data into continuous time series features through time alignment, interpolation filling, and feature generation; for dynamic modality data, allocate features to the unified time axis according to time windows.

[0058] Optionally, the time serialization processing of static data includes:

[0059] Adopt an interpolation method to fill the preset time window lacking picture data;

[0060] When the picture acquisition interval is greater than the set value, generate transition features through a generative adversarial network or generate transition features through time-aware linear interpolation, so that discrete static data is converted into continuous time series features.

[0061] Since the timestamps of static modality data such as pictures are often sparse and discontinuous and are difficult to directly participate in time series modeling, the present invention proposes a time serialization strategy, and converts discrete static data into continuous time series features through time alignment, interpolation filling, and feature generation.

[0062] Optionally, considering the influence of time context on interpolation, select the nearest neighbor interpolation algorithm combined with time weights. Let the time be , and the picture feature obtained from the time difference is , the time The nearest known picture feature before is , is The moment before, is The moment after, the time The nearest known picture feature after is , is The weight of is The weight of

[0063] ;

[0064] ;

[0065] ;

[0066] When the image acquisition interval is greater than the set value, transitional features are generated through a generative adversarial network. Let the adversarial network be , then:

[0067] ;

[0068] Or transitional features are generated through time-aware linear interpolation, then:

[0069] .

[0070] Compared with traditional linear interpolation, the time-aware mechanism combined with time keys makes the interpolation result more robust.

[0071] Optionally, all modal data is aligned to a unified time axis to generate cross-modal time series modal data, including: for dynamic modal data, the features of the dynamic modal data are allocated to the unified time axis according to time windows; for static modal data, the static modal data is mapped into the time series through a time serialization method.

[0072] Realize the seamless alignment of static data and dynamic modalities in the time dimension, which not only solves the problem of lack of time continuity of static modal data, but also ensures the smoothness and authenticity of the time series through the combination of interpolation and feature generation, providing high-quality input for multi-modal fusion modeling.

[0073] Multi-modal detection data includes modal data such as videos, pictures, audios, and sensors. There are differences in the sampling timestamps and frequencies of different modalities, and it is a great challenge to directly align these data. The present invention proposes an association mechanism based on time keys to align all modal data to a unified time axis and generate cross-modal time series, providing a unified time series input for subsequent annotation tasks.

[0074] Optionally, all modal data is encoded into high-dimensional features and standardized and aligned through linear mapping processing to generate a unified feature dimension, including:

[0075] For picture modal data, a time-aware self-attention mechanism is introduced through a convolutional neural network or a vision transformer to extract static features of the picture modality;

[0076] For video modal data, a separate spatio-temporal convolution and a multi-head spatio-temporal attention mechanism are used to jointly extract the time information and spatial information of the video;

[0077] For audio modal data, a temporal self-attention mechanism is introduced through a convolutional neural network or a Transformer-based network to extract the frequency-domain and time-domain features of the audio modality;

[0078] For sensor modal data, a one-dimensional convolutional network or a recurrent neural network is used to extract the time-series features of the sensor modality.

[0079] In multi-modal annotation tasks, the feature information contained in different modal data (such as images, videos, audio, and sensors) varies greatly. Therefore, a dedicated feature extraction network is required to efficiently encode each modality. The present invention encodes all modal data into high-dimensional feature representations and normalizes and aligns them through linear mapping to generate a unified feature dimension, providing a consistent input for subsequent modeling.

[0080] As an alternative implementation, for image modal data, a self-attention mechanism is introduced to enhance the global receptive field of feature extraction, thereby better capturing the global context information in the image. In addition, a multi-scale feature fusion strategy is combined to extract the fine-grained information and high-level semantic information in the image. For example, let represent the th image of the image modal data, represent the image features extracted by the image encoder, represent the total number of time steps, represent the feature dimension, represent the feature matrix of the extracted image modal data, represent the image encoder, then:

[0081] .

[0082] As an alternative implementation, for video modal data, a spatio-temporal joint feature of the video modality is extracted through a convolutional network or a spatio-temporal feature extractor based on self-attention, solving the problem of insufficient modeling of the time dimension in traditional methods. Optionally, separable spatio-temporal convolution and a multi-head spatio-temporal attention mechanism are used to jointly extract the time information and space information of the video. For example, let represent the video modal data at the th time step, represent the features of the video modal data extracted by the spatio-temporal feature extractor, represent the video encoder, represent the total number of time steps, represent the feature dimension, represent the feature matrix of the extracted video modal data, then:

[0083] .

[0084] As an alternative implementation, for audio modality data, frequency-domain and time-domain features of the audio modality are extracted through a convolutional neural network or a Transformer-based network. On the basis of traditional Mel spectrogram feature extraction, a multi-scale feature extraction module is introduced to capture cross-band dependencies in the audio modality. In addition, a method based on a frequency-domain attention mechanism is combined to enhance the ability to capture features of key frequency bands. For example, let represent the audio modality data at the -th time step, represent the features of the audio modality data extracted by the audio encoder, represent the audio encoder, represent the total number of time steps, represent the feature dimension, represent the feature matrix of the extracted video modality data, then:

[0085] .

[0086] As an alternative implementation, for sensor modality data, time series features of the sensor modality are extracted through a one-dimensional convolutional network or a recurrent neural network. On the basis of traditional time series modeling methods, a time-aware self-attention mechanism is introduced to enhance the ability to model long-term dependencies. At the same time, a residual connection structure is adopted to alleviate the problem of information loss in sensor data. For example, let represent the sensor modality data at the -th time step, represent the features of the sensor modality data extracted by the sensor encoder, represent the sensor encoder, represent the total number of time steps, represent the feature dimension, represent the feature matrix of the extracted video modality data, then:

[0087] .

[0088] Finally, the modality features are stacked into a unified feature matrix. Let the feature matrix be , be the number of modalities, then:

[0089] .

[0090] Through the standardization and alignment of modality features, the present invention eliminates the differences in feature dimensions between modalities and significantly improves the adaptability of the annotation model to multi-modal data.

[0091] Optionally, a Transformer-based time series model is adopted to extract the global time correlation features of the multimodal feature data, including: Let the unified aligned multimodal feature matrix be , be the output of the time series modeling module, represent real numbers, represent the total number of time steps, represent the feature dimension, represent the time series model, then:

[0092] .

[0093] Optionally, the interaction features between modalities are extracted based on the attention mechanism, including: Let represent the query vector of the features of the target modality at the current time step, represent the key vector of the features of the historical time step, represent the eigenvalue of the candidate modality, be the output of the time series modeling module, and the interaction features between modalities are , be the attention module, then:

[0094] .

[0095] Optionally, based on the interaction features, the annotation results are output, and dynamic modality weights are set to assign learned weights to the features of each modality, and the weights and biases of the output layer are dynamically adjusted according to the input interaction features to update the annotation results, including: Let represent the number of modalities, represent the total number of modalities, represent the time step, represent the weight of each modality, represent the annotation result, represent the th modality's fused feature at the time step, represent the weight matrix of the output layer, represent the bias term of the output layer, represent the weight matrix of the output layer at the time step , represent the bias term of the output layer at the time step , represent the dynamic generation function of the weight matrix, represent the dynamic generation function of the bias term. For different time steps and different input features, a dynamic weight adjustment mechanism is introduced to enable the output layer to adaptively optimize the weights and biases, then:

[0096] ;

[0097] ;

[0098] 。

[0099] Introduce a dynamic weight adjustment mechanism for different time steps or different input features, enabling the output layer to adaptively optimize weights and biases.

[0100] Through global time modeling and cross-modal interaction modeling, the present invention captures complex time dynamics and modal dependency relationships and can generate high-quality annotation results.

[0101] The present invention proposes an automated annotation method. By means of temporal feature modeling, inter-modal correlation alignment, and feature dynamic optimization, it solves the problems of time synchronization and insufficient feature expression in multi-modal data annotation. By dynamicizing static modalities and unifying time series features, combined with cross-modal feature fusion, it can quickly generate high-precision and low-cost annotation results, realizing the intelligence and automation of the annotation process. This method requires little manual intervention, can adapt to the time step differences and complex dynamic relationships of different modal data, significantly improves the annotation efficiency and consistency, and is applicable to scenarios such as industrial inspection and medical inspection that require efficient multi-modal data processing, providing high-quality automated annotation support for information-based inspection systems.

[0102] Embodiment 2

[0103] Based on the same principle as the method shown in Embodiment 1 of the present invention, as shown in the appendix Figure 2 shown, the embodiment of the present invention also provides a dynamic and static detection data annotation system based on time key association, including an alignment unit, an encoding and alignment unit, a first feature extraction unit, a second feature extraction unit, an interaction feature extraction unit, and an output and update unit;

[0104] The alignment unit is used to align all modal data to a unified time axis to generate modal data of cross-modal time series;

[0105] The encoding and alignment unit is used to encode all modal data into high-dimensional features and perform standardized alignment through linear mapping processing to generate a unified feature dimension;

[0106] The first feature extraction unit is used to extract features from each modal data by using a neural network to obtain multi-modal feature data;

[0107] The second feature extraction unit is used to extract the global time correlation features of multi-modal feature data by using a Transformer-based time series model;

[0108] An interaction feature extraction unit for extracting interaction features between modalities based on an attention mechanism;

[0109] An output and update unit for outputting an annotation result based on the interaction features, setting dynamic modality weights, assigning learned weights to the features of each modality, and updating the weights and biases of the output layer according to the input interaction features to update the annotation result.

[0110] Optionally, before aligning all modality data to a unified time axis, it further includes: performing time serialization processing on static modality data, and converting discrete static data into continuous time series features through time alignment, interpolation filling, and feature generation; for dynamic modality data, distributing the features to the unified time axis according to time windows.

[0111] Optionally, the time serialization processing of static data includes:

[0112] Using an interpolation method to fill a preset time window lacking picture data;

[0113] When the picture acquisition interval is greater than a set value, generating transition features through a generative adversarial network or generating transition features through time-aware linear interpolation, so that discrete static data is converted into continuous time series features.

[0114] Optionally, considering the influence of time context on interpolation, selecting a nearest neighbor interpolation algorithm combined with time weights. Let the time be , and the picture feature obtained from the time difference is , the time The nearest known picture feature before is , is The moment before, is The moment after, the time The nearest known picture feature after is , is The weight of, is The weight of, then:

[0115] ;

[0116] ;

[0117] ;

[0118] When the picture acquisition interval is greater than a set value, generating transition features through a generative adversarial network. Let the generative adversarial network be , then:

[0119] ;

[0120] Or generate transitional features through time-aware linear interpolation, then:

[0121] .

[0122] Optionally, align all modal data to a unified timeline to generate modal data of cross-modal time series, including: for dynamic modal data, allocate the features of dynamic modal data to the unified timeline according to time windows; for static modal data, map static modal data into a time series through a time serialization method.

[0123] Optionally, encode all modal data into high-dimensional features and perform standardized alignment through linear mapping processing to generate a unified feature dimension, including:

[0124] For picture modal data, introduce a time-aware self-attention mechanism through a convolutional neural network or a vision transformer to extract static features of picture modal data;

[0125] For video modal data, adopt a separable spatio-temporal convolution and a multi-head spatio-temporal attention mechanism to jointly extract the time information and space information of the video;

[0126] For audio modal data, introduce a time-aware self-attention mechanism through a convolutional neural network or a transformer-based network to extract frequency-domain features and time-domain features of audio modal data;

[0127] For sensor modal data, extract time series features of sensor modal data through a one-dimensional convolutional network or a recurrent neural network.

[0128] Optionally, adopt a time series model based on Transformer to extract global time correlation features of multi-modal feature data, including: Let the unified aligned multi-modal feature matrix be , be the output of the time series modeling module, denote real numbers, denote the total number of time steps, denote the feature dimension, denote the time series model, then:

[0129] .

[0130] Optionally, extract interaction features between modalities based on the attention mechanism, including: Let denote the query vector of the feature of the target modality at the current time step, denote the key vector of the feature of the historical time step, Represents the eigenvalue of the candidate modality, is the output of the time series modeling module, and the interaction feature between modalities is , is the attention module, then:

[0131] .

[0132] Optionally, based on the interaction feature, output the annotation result, set the dynamic modality weight, assign the learned weight to the feature of each modality, and update the weight and bias of the output layer according to the input interaction feature dynamically, including: Let represent the number of modalities, represent the total number of modalities, represent the time step, represent the weight of each modality, represent the annotation result, represent the th modality's fused feature at the time step, represent the weight matrix of the output layer, represent the bias term of the output layer, represent the weight matrix of the output layer at time step , represent the bias term of the output layer at time step , represent the dynamic generation function of the weight matrix, represent the dynamic generation function of the bias term. Introduce a dynamic weight adjustment mechanism for different time steps and different input features, so that the output layer can adaptively optimize the weights and biases, then:

[0133] ;

[0134] ;

[0135] .

[0136] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A dynamic and static detection data labeling method based on time key association, characterized in that: include: Perform time series processing on static modal data, and convert discrete static data into continuous time series features through time alignment, interpolation filling and feature generation; For dynamic modal data, the features are assigned to a unified time axis by time window; all modal data are aligned to a unified time axis to generate modal data of cross-modal time series; Time series processing of static data, including: considering the impact of time context on interpolation, selecting the nearest neighbor interpolation algorithm combined with time weight to fill the preset time window with missing image data; when the image acquisition interval is greater than the set value, generating transition features through generative adversarial networks or through time-aware linear interpolation, so that discrete static data is converted into continuous time series features; All modal data are encoded into high-dimensional features and standardized and aligned through linear mapping to generate a unified feature dimension; A neural network is used to extract features from each modal data to obtain multimodal feature data; A Transformer-based time series model is used to extract global time correlation features of multimodal feature data; Extract interaction features between modalities based on attention mechanism; The annotation results are output based on the interactive features, dynamic modal weights are set, and learning weights are assigned to the features of each modality. The weights and biases of the output layer are dynamically adjusted according to the input interactive features to update the annotation results.

2. According to claim 1, the method for labeling dynamic and static detection data based on time key association is characterized in that: Considering the influence of time context on interpolation, we choose the nearest neighbor interpolation algorithm combined with time weight. Let time be , the image features obtained in the time difference ,time The most recent known image feature is , for The moment before, for After that moment, time The nearest known image feature is , for The weight of for The weight of , then: ; ; ; When the image collection interval is greater than the set value, the transition features are generated by generating adversarial networks. Let the adversarial network be ,but: ; Or generate transition features through time-aware linear interpolation, then: 。 3. The dynamic and static detection data labeling method based on time key association according to claim 1 is characterized in that: All modal data are aligned to a unified time axis to generate modal data of cross-modal time series, including: for dynamic modal data, the features of dynamic modal data are assigned to a unified time axis according to time windows; for static modal data, the static modal data are mapped to the time series through a time serialization method.

4. According to claim 1, the method for labeling dynamic and static detection data based on time key association is characterized in that: All modal data are encoded into high-dimensional features and standardized and aligned through linear mapping to generate a unified feature dimension, including: For image modality data, a time-aware self-attention mechanism is introduced through convolutional neural networks or visual transformers to extract static features of the image modality; For video modality data, a separated spatiotemporal convolution and a multi-head spatiotemporal attention mechanism are used to jointly extract the temporal and spatial information of the video. For audio modality data, a time-aware self-attention mechanism is introduced through a convolutional neural network or a transformer-based network to extract the frequency domain features and time domain features of the audio modality; For sensor modal data, the time series features of the sensor modality are extracted through a one-dimensional convolutional network or a recurrent neural network.

5. According to claim 1, the method for labeling dynamic and static detection data based on time key association is characterized in that: The Transformer-based time series model is used to extract the global time correlation features of multimodal feature data, including: assuming that the uniformly aligned multimodal feature matrix is , is the output of the time series modeling module, represents a real number, represents the total number of time steps, represents the feature dimension, represents a time series model, then: 。 6. The dynamic and static detection data labeling method based on time key association according to claim 1 is characterized in that: Extract the interactive features between each modality based on the attention mechanism, including: The query vector representing the features of the target modality at the current time step, The key vector representing the features of the historical time step, represents the eigenvalue of the candidate mode, is the output of the time series modeling module, and the interaction characteristics between modes are , is the attention module, then: 。 7. The dynamic and static detection data labeling method based on time key association according to claim 1 is characterized in that: Output the annotation results based on the interactive features, set dynamic modal weights, assign learning weights to the features of each modality, and dynamically adjust the weights and biases of the output layer according to the input interactive features to update the annotation results, including: Indicates the number of modes, Represents the total number of modes, represents the time step, represents the weight of each mode, Indicates the labeling result. Indicates The fusion features of the modalities at the time step, represents the weight matrix of the output layer, represents the bias term of the output layer, Represents the time step The weight matrix of the output layer is Represents the time step The bias term of the output layer of represents the dynamic generation function of the weight matrix, The dynamic generation function of the bias term is represented. For different time steps and different input features, a dynamic weight adjustment mechanism is introduced to enable the output layer to adaptively optimize the weights and biases. Then: ; ; 。 8. A dynamic and static detection data labeling system based on the dynamic and static detection data labeling method based on time key association according to claim 1, characterized in that: It includes an alignment unit, an encoding and alignment unit, a first feature extraction unit, a second feature extraction unit, an interactive feature extraction unit, and an output and update unit; The alignment unit is used to: perform time series processing on static modal data, and convert discrete static data into continuous time series features through time alignment, interpolation filling and feature generation; For dynamic modal data, the features are assigned to a unified time axis by time window; all modal data are aligned to a unified time axis to generate modal data of cross-modal time series; Time series processing of static data, including: considering the impact of time context on interpolation, selecting the nearest neighbor interpolation algorithm combined with time weight to fill the preset time window with missing image data; when the image acquisition interval is greater than the set value, generating transition features through generative adversarial networks or through time-aware linear interpolation, so that discrete static data is converted into continuous time series features; The encoding and alignment unit is used to encode all modal data into high-dimensional features and perform standardized alignment through linear mapping to generate a unified feature dimension; A first feature extraction unit, used to extract features from each modal data using a neural network to obtain multimodal feature data; A second feature extraction unit is used to extract global time correlation features of multimodal feature data using a Transformer-based time series model; Interaction feature extraction unit, used to extract interaction features between modalities based on the attention mechanism; The output and update unit is used to output the annotation results based on the interactive features, set dynamic modal weights, assign learning weights to the features of each modality, and dynamically adjust the weights and biases of the output layer according to the input interactive features to update the annotation results.

Citation Information

Patent Citations

  • Multi-modal model and method for fusing characters, images and audios

    CN118861988A