Multi-modal sequence data processing method and device, equipment and medium

By generating multi-scale feature pyramid sets, cross-modal feature alignment, local and global attention processing, and cross-layer information interaction, the problem of insufficient modal correlation in multimodal long sequence data is solved, realizing refined modeling and dynamic fusion of multimodal tasks, and improving processing effect and accuracy.

CN120951247APending Publication Date: 2025-11-14PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 8 Cited by

Patent Information

Application Number
CN202511060298.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing technologies lack collaborative extraction, alignment, and dynamic fusion mechanisms for multi-scale features when processing multimodal long sequence data. This results in weak correlations between different modalities at different scales and insufficient utilization of features, making it difficult to meet the needs of unified modeling of refined and global information in complex tasks.

Method used

By acquiring raw data from visual, language, and motion sensor sequences, an initial feature sequence is generated. Multi-scale feature levels are extracted and combined into a multi-scale feature pyramid set. Cross-modal feature alignment is performed for each feature level, local and global attention processing is executed, long-distance dependent features are generated, and cross-layer information interaction is performed. Finally, multi-modal information is dynamically fused to generate the final fused feature.

Benefits of technology

It effectively solves the problem of insufficient information correlation between different modalities at different scales in multimodal long sequence data, realizes refined modeling and dynamic fusion of multi-scale features of different modalities, and improves the processing effect and accuracy of multimodal tasks in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120951247A_ABST
    Figure CN120951247A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, can be applied to business scenes such as agent autonomous decision making, financial science and technology and medical health, and discloses a multi-modal sequence data processing method, device and equipment and a medium. Extracting multi-scale feature hierarchies, combining the multi-scale feature hierarchies into a multi-scale feature pyramid set, performing cross-modal feature alignment to generate a multi-scale alignment feature sequence, executing local and global attention processing to generate long-distance dependency features, performing cross-layer information interaction to generate comprehensive multi-scale features, and performing multi-scale feature extraction; and dynamically fusing the multi-modal information and inputting the multi-modal information into a task decision network to obtain a target task result. According to the method, through the multi-scale feature pyramid, cross-modal alignment, attention processing and cross-layer information interaction, the problem of insufficient relevance between different modals and different scales in multi-modal long sequence data is solved, and fine modeling and dynamic fusion of multi-modal and multi-scale features are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device, and storage medium for processing multimodal sequence data. Background Technology

[0002] In the fintech business, current intelligent analysis and decision-making applications for multimodal sequence data still have significant shortcomings in processing multimodal information such as visual, linguistic, and behavioral data. Traditional methods often employ single-scale feature extraction and fusion approaches, making it difficult to simultaneously capture image details and global layout information in invoice recognition, or to simultaneously model lexical-level semantic details and overall communication logic of customer dialogue in large-value transaction risk analysis. Existing methods often fuse features by simply splicing or stacking them, lacking cross-modal and cross-scale collaborative processing mechanisms, resulting in weak correlations between features and insufficient information utilization. When dealing with long multimodal sequence data such as visual documents, text information, and user behavior sequences in financial scenarios, existing models struggle to process them efficiently, often exhibiting inconsistent feature representations and excessive computational resource consumption, thereby affecting the accuracy of risk assessment and customer service response.

[0003] In the healthcare field, clinical auxiliary diagnostic and health monitoring systems increasingly rely on the comprehensive analysis of multimodal data, including visual (e.g., medical images), linguistic (e.g., electronic medical records), and motor (e.g., patient motion data). However, existing technologies face multiple challenges when processing these long-term multimodal data sequences. Traditional sequence modeling methods, such as recurrent neural networks and their variants, often suffer from gradient vanishing or exploding when handling ultra-long-term medical monitoring sequences, making it difficult for the model to retain long-term dependencies. This manifests as a difficulty in simultaneously understanding the correlation between past patient behavior records and recent monitoring results. Furthermore, while Transformer-based methods theoretically possess the capability for long-term sequence modeling, computational and memory overhead increases dramatically when dealing with large amounts of high-resolution medical image frames, lengthy medical record texts, and long-term motion data, severely limiting the real-time performance and practicality of clinical applications. In particular, existing technologies lack fine-grained dynamic interaction mechanisms between multimodal data, failing to effectively utilize complementary information between different modalities, thus affecting the accurate diagnosis of complex conditions.

[0004] In the fields of robotics and human-computer interaction, existing models also have significant shortcomings in multimodal long-sequence tasks. Traditional methods struggle to simultaneously extract detailed features of small objects and large-scale global scene information in the visual modality; in the language modality, it is difficult to simultaneously capture lexical-level local semantic details and the overall logic of sentence discourse; and in the action modality, it is difficult to uniformly represent motion details and trend features at different time scales. Existing multimodal fusion methods lack effective multi-scale feature processing and alignment mechanisms, making it easy for robots to forget or execute incorrect instructions when performing multi-step tasks. Furthermore, current Transformer-type methods suffer from high computational complexity and memory consumption in long-sequence multimodal tasks, lacking efficient feature alignment and dynamic fusion strategies, making it difficult to guarantee performance and efficiency in dynamic scenes and failing to meet the demands of refined, real-time processing for multimodal long-sequence tasks. Summary of the Invention

[0005] The main objective of this invention is to provide a method, apparatus, device, and storage medium for processing multimodal sequence data. This invention aims to address the technical problem that existing technologies lack collaborative extraction, alignment, and dynamic fusion mechanisms for multi-scale features when processing long multimodal sequence data. This results in weak correlations between different modalities at different scales, insufficient utilization of features, and difficulty in meeting the technical requirements for refined and global information unified modeling in complex tasks.

[0006] To achieve the above objectives, the present invention provides a multimodal sequence data processing method, comprising:

[0007] Acquire raw visual sequence data, raw language text sequence data, and raw motion sensor sequence data, and generate an initial feature sequence based on the raw visual sequence data, raw language text sequence data, and raw motion sensor sequence data;

[0008] Multi-scale feature levels are extracted from the initial feature sequence, and the multi-scale feature levels are combined into a multi-scale feature pyramid set.

[0009] Cross-modal feature alignment is performed for each feature level of the multi-scale feature pyramid set to generate a multi-scale aligned feature sequence;

[0010] Local attention and global attention processing are performed on the multi-scale aligned feature sequence to generate long-distance dependent features;

[0011] Cross-layer information interaction is performed on the long-distance dependent features to generate comprehensive multi-scale features;

[0012] Based on the comprehensive multi-scale features, multi-modal information is dynamically fused to generate the final fused features;

[0013] The final fused features are input into the task decision network to obtain the target task result.

[0014] Furthermore, to achieve the above objectives, the present invention provides a multimodal sequence data processing apparatus, comprising:

[0015] A multimodal input encoding module is used to acquire raw visual sequence data, raw language text sequence data, and raw motion sensor sequence data, and generate an initial feature sequence based on the raw visual sequence data, raw language text sequence data, and raw motion sensor sequence data.

[0016] A multi-scale feature extraction module is used to extract multi-scale feature levels from the initial feature sequence and combine the multi-scale feature levels into a multi-scale feature pyramid set.

[0017] The cross-modal alignment module is used to perform cross-modal feature alignment for each feature level of the multi-scale feature pyramid set, generating a multi-scale aligned feature sequence;

[0018] The local and global attention module is used to perform local and global attention processing on the multi-scale aligned feature sequence to generate long-distance dependent features.

[0019] The cross-layer interaction fusion module is used to perform cross-layer information interaction on the long-distance dependent features to generate comprehensive multi-scale features;

[0020] The dynamic multimodal fusion module is used to dynamically fuse multimodal information based on the comprehensive multi-scale features to generate the final fused features;

[0021] The task decision module is used to input the final fused features into the task decision network to obtain the target task result.

[0022] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and a multimodal sequence data processing program stored in the memory and executable on the processor, wherein when the multimodal sequence data processing program is executed by the processor, it implements the steps of the multimodal sequence data processing method as described above.

[0023] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing a multimodal sequence data processing program, which, when executed by a processor, implements the steps of the multimodal sequence data processing method described above.

[0024] Beneficial Effects: This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as intelligent agent autonomous decision-making, fintech, and healthcare. It discloses a method, apparatus, device, and medium for processing multimodal sequence data, including: acquiring raw visual sequence data, raw language text sequence data, and raw motion sensor sequence data and generating an initial feature sequence; extracting multi-scale feature levels from the initial feature sequence and combining them into a multi-scale feature pyramid set; performing cross-modal feature alignment on each feature level of the multi-scale feature pyramid set to generate a multi-scale aligned feature sequence; performing local attention processing and global attention processing on the multi-scale aligned feature sequence to generate long-distance dependent features; performing cross-layer information interaction on the long-distance dependent features to generate comprehensive multi-scale features; dynamically fusing multimodal information based on the comprehensive multi-scale features to generate a final fused feature; and finally, inputting the fused feature into a task decision network to obtain the target task result. This invention effectively addresses the problem of insufficient information correlation between different modalities at different scales in multimodal long sequence data by using multi-scale feature pyramid sets, cross-modal feature alignment, combination of local and global attention, and cross-layer information interaction. It achieves refined modeling and dynamic fusion of multi-scale features of different modalities, thereby improving the processing effect and accuracy of multimodal tasks in complex scenarios. Attached Figure Description

[0025] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings:

[0026] Figure 1 This is a schematic diagram of an application environment for a multimodal sequence data processing method according to an embodiment of the present invention;

[0027] Figure 2 This is a flowchart illustrating an embodiment of the multimodal sequence data processing method of the present invention;

[0028] Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the multimodal sequence data processing device of the present invention;

[0029] Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention;

[0030] Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0031] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0032] The multimodal sequence data processing method provided in this invention can be applied to, for example... Figure 1In this application environment, the user terminal communicates with the server via a network. The server can obtain raw visual sequence data, raw language text sequence data, and raw motion sensor sequence data from the user terminal and generate an initial feature sequence. Multi-scale feature levels are extracted from the initial feature sequence and combined into a multi-scale feature pyramid set. Cross-modal feature alignment is performed on each feature level of the multi-scale feature pyramid set to generate a multi-scale aligned feature sequence. Local and global attention processing is performed on the multi-scale aligned feature sequence to generate long-distance dependent features. Cross-layer information interaction is performed on the long-distance dependent features to generate comprehensive multi-scale features. Based on the comprehensive multi-scale features, multi-modal information is dynamically fused to generate the final fused feature. The final fused feature is input into the task decision network to obtain the target task result. This invention effectively solves the problem of insufficient information correlation between different modalities at different scales in long multi-modal sequence data by using a multi-scale feature pyramid set, cross-modal feature alignment, combination of local and global attention, and cross-layer information interaction. It achieves refined modeling and dynamic fusion of multi-modal multi-scale features, improving the processing effect and accuracy of multi-modal tasks in complex scenarios. The user terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will now be described in detail through specific embodiments.

[0033] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the multimodal sequence data processing method provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0034] like Figure 2 As shown, the multimodal sequence data processing method proposed in this invention includes the following steps:

[0035] S10, acquire the original visual sequence data, the original language text sequence data, and the original motion sensor sequence data, and generate an initial feature sequence based on the original visual sequence data, the original language text sequence data, and the original motion sensor sequence data;

[0036] In this embodiment, the acquisition of raw visual sequence data, raw language text sequence data, and raw motion sensor sequence data aims to provide comprehensive and multi-dimensional input data sources for subsequent information processing. Raw visual sequence data is a sequence of time-continuous image frames acquired by imaging devices, including but not limited to cameras, depth cameras, and stereo vision systems, capturing visual scene information at different points in time. Raw language text sequence data is derived from text data converted by speech recognition devices or directly received, including text transcribed from user voice commands, structured or unstructured text paragraphs, etc., used to record semantic information expressing intent or describing tasks. Raw motion sensor sequence data is time-series motion parameter data acquired by inertial measurement units, force sensors, position encoders, etc., recording the motion state of the physical world.

[0037] Generating the initial feature sequence requires aligning and standardizing the aforementioned data sources by time. For visual sequence raw data, frame rate normalization can be used to reduce or insert image frames, synchronizing it with other modalities in time. Then, size normalization unifies the image resolution, ensuring consistency in subsequent processing. For language text sequence raw data, a unified character encoding format can be used for conversion, combined with linguistic model segmentation and part-of-speech tagging, converting the text into semantic units suitable for vectorization. For motion sensor sequence raw data, the data acquisition rate can be adjusted by synchronizing the sampling frequency, followed by noise filtering and amplitude normalization for each sensor channel to remove high-frequency noise and zero-point offset during acquisition.

[0038] Standardized data is input into different processing modules. Visual data can be fed into a convolutional neural network, which uses multi-layer convolutional operators to extract low- to high-order features such as spatial texture, edges, and color distribution, forming a sequence of visual feature vectors. Language data can be fed into a word embedding model, which converts each word segmentation unit into a high-dimensional dense vector and introduces positional encoding vectors to preserve the order relationship, forming a sequence of language feature vectors. Action data can be fed into a fully connected network, which processes it through multiple layers of weight transformation and activation functions to extract features such as action amplitude, acceleration, and velocity, forming a sequence of action feature vectors.

[0039] In the temporal dimension, the visual feature vector sequence, language feature vector sequence, and action feature vector sequence need to be concatenated in timestamp order to form a unified initial feature sequence. This operation not only requires strict alignment between data points but also requires the ability to comprehensively include feature representations of vision, language, and action at each time point, laying the foundation for subsequent multimodal joint analysis. This concatenation process uses dynamic buffer management to synchronize data blocks from the three modalities according to timestamps, and then concatenates the vectors from their respective time points into a higher-dimensional joint vector.

[0040] Raw visual sequence data acquisition can be achieved through an RGB camera mounted on the robot's front end. Real-time frame rate resampling and pixel size adjustment are performed using an image preprocessing module to ensure the input image sequence has a uniform time interval and spatial resolution. Raw language text sequence data processing involves a built-in speech recognition module that converts on-site speech commands into text. Simultaneously, character encoding is performed on the text stream, and a self-trained lexical analyzer is used for Chinese or English word segmentation. Part-of-speech tagging results are then incorporated to enrich the semantic structure of the language data. Raw motion sensor sequence data acquisition is accomplished by an inertial measurement unit mounted on the robot's joints. The raw signals are input to a multi-channel digital signal processing module, where low-pass filtering removes high-frequency noise, and mean normalization eliminates dimensional differences between sensor channels.

[0041] The initial feature sequence generation can employ a parallel dataflow architecture, with each modality's data fed into a GPU-parallel feature extraction network. Visual data is input into a deep convolutional network, using convolutional blocks with batch normalization and activation functions to extract local texture and overall shape features. Language data is input into a multi-dimensional embedding vector computation module, which, combined with a positional encoding mechanism, assigns a spatial embedding vector representation to each word, enhancing the ability to model word order relationships. Action data is input into a multi-layer fully connected network, with each layer incorporating activation functions for non-linear feature transformation to extract action change patterns across different time periods.

[0042] Finally, the visual, linguistic, and action feature vectors synchronized according to a unified time base are aligned and synthesized through multi-dimensional vector concatenation operations to form an initial feature sequence in time series form, which can be directly used for subsequent multimodal information fusion or feature analysis tasks. At this point, the joint vector at each time point not only retains the original information features of each modality, but also strengthens the semantic consistency of multimodality through the contextual relationships between them.

[0043] Example description: In the field of healthcare, it can be used in robot-assisted surgery to record the surgical scene in real time by acquiring visual image streams, recording doctors' verbal instructions by acquiring voice input, and recording the dynamic operation of the robotic arm by acquiring motion sensor data. Through unified processing, an initial feature sequence is formed, realizing a multimodal comprehensive understanding of vision, language and motion, and providing data support for the next step of analysis and decision-making in surgical operations.

[0044] In the fintech business, it can be applied to remote customer service. By combining visual data of customer behavior captured by cameras, customer inquiry content obtained by voice recognition, and interactive behavior data monitored by desktop motion sensors, an initial feature sequence is formed through unified timeline alignment and standardized processing, which supports the dynamic perception and interactive decision-making of financial service robots in the service process.

[0045] In robotic operation scenarios, it can be applied to automated handling robots. It collects visual data of the handling environment via cameras, gathers operator voice task commands, and collects motion data of the robotic arm to generate an initial feature sequence, which then drives the robot to perform safe and precise operations. In human-computer interaction scenarios, it can be applied to intelligent assistants. Visual data acquires user expressions and postures, language data records dialogue content, and motion data collects user touch and operation feedback. The initial feature sequence ensures the temporal consistency of multimodal information, making the interaction process more natural and the feedback more accurate.

[0046] This embodiment synchronously collects and standardizes three types of data—visual, verbal text, and motion sensor data—to form an initial feature sequence with temporal consistency and unified multimodal expression. This effectively solves the problems of heterogeneous multimodal data sources, asynchronous sampling, and inconsistent data expression in existing technologies, improves the basic consistency level before data fusion, provides high-quality input for multi-scale feature extraction and subsequent alignment processing, and ensures the reliability and accuracy of subsequent multimodal processing.

[0047] S20, extract multi-scale feature levels from the initial feature sequence, and combine the multi-scale feature levels into a multi-scale feature pyramid set;

[0048] In this embodiment, starting from the initial feature sequence of multimodal sequence data, the first step is to achieve internal deconstruction and scale hierarchy partitioning of different modalities. The core technology of this process lies in the systematic hierarchical expression of the inherent multidimensional (spatial, temporal, semantic) heterogeneous information in the initial feature sequence. The multi-scale representation of visual information stems from the human perception's need for "coexistence of detail and global perspective." Shallow visual levels focus on local information such as edges, corners, and textures, while higher visual levels gradually extract object contours, semantic regions, and global layouts. The machine implementation of this idea relies on the increasing receptive field structure of convolutional neural networks. In linguistic information, words themselves carry the smallest semantic units, phrases embody local contextual combinations, and sentences and paragraphs gradually capture cross-sentence logic and long-distance dependencies. This hierarchical partitioning technique can be traced back to the hierarchical representation theory of natural language processing. Action information, as a time series, reflects short-term action change trends at the detail scale and the integrity of the behavioral sequence at the global scale. The technical implementation originates from hierarchical convolutional processing with multiple time windows.

[0049] Therefore, multi-scale feature extraction of the initial feature sequence is essentially guided by the information complexity and abstraction level within each modality, constructing adaptive scale-layered frameworks for visual, linguistic, and action modalities respectively. This operation requires visual information to be downsampled layer by layer in terms of spatial resolution, linguistic information to be aggregated step by step in terms of syntactic and semantic blocks, and action information to be parsed in parallel along the time axis according to local windows and global trends. The formation of multi-scale feature layers within each modality also requires temporal synchronization and spatial alignment to ensure that expressions at different scales have consistent reference timestamps and spatial semantic contexts.

[0050] After independently extracting multi-scale features at visual, linguistic, and behavioral levels, these levels must be organized into a multi-scale feature pyramid set. The meaning of this pyramid set lies not only in the stacking of levels but also in the structured alignment across modal levels, ensuring a one-to-one correspondence between visual detail levels and linguistic word levels, short-term behavioral change levels, and consistent representation of the visual global level, linguistic paragraph level, and long-term behavioral trend level. This process involves nested mapping of multimodal spaces, scale index construction, and structured storage to ensure the addressability and consistency of subsequent cross-modal alignment and fusion operations.

[0051] Visual sequences can be input into a depth-adjustable convolutional network. The first to third layers use smaller receptive fields (e.g., 3×3 convolutional kernels) to generate edge, texture, and local graphic detail feature levels, respectively. The fourth to sixth layers use larger receptive fields (e.g., 5×5 and 7×7 convolutional kernels) to generate medium-scale graphic contours, large-scale region distributions, and overall scene semantic levels, respectively. Each layer is progressively downsampled in the spatial dimension (e.g., 2x pooling) to form a multi-resolution pyramid structure.

[0052] After word segmentation, part-of-speech tagging, and word vectorization, the language sequence is processed using a semantic aggregation network (e.g., the Transformer Encoder layer). The first layer captures phrase-level contextual relationships through window aggregation; the second layer forms sentence representations based on semantic blocks; and the third layer integrates multiple sentence relationships to form paragraph-level representations. The length of the output sequence at each level decreases hierarchically, reflecting the characteristic of a progressively increasing semantic abstraction scale.

[0053] Action sequences are processed using multi-scale one-dimensional convolutions. Short-window (e.g., 3 or 5 frames) convolutional layers focus on extracting local dynamic features, medium-window (e.g., 15 or 30 frames) convolutional layers focus on mid-range motion trends, and long-window (e.g., 60 frames or more) convolutional layers capture global behavioral patterns. The convolutional kernel stride and window size are adjusted in tandem to ensure consistent alignment of action features across all scales on the timeline.

[0054] The multi-scale feature levels of each modality extracted independently by the three modules mentioned above need to be organized in modality, level, and time order using a unified indexing mechanism. This organizational structure is stored as a multi-scale feature pyramid set, with spatial, temporal, and semantic multi-dimensional index support, which can support efficient and rapid retrieval and localization in the subsequent cross-modal fusion stage.

[0055] In different scenarios, visual networks can employ different depths (e.g., shallow networks are suitable for high frame rate video streams, while deep networks are suitable for high-definition still images); language networks can adjust the aggregation window length (short windows are suitable for instructional statements, while long windows are suitable for document-level text); and action networks can adjust the convolution kernel stride (short strides are suitable for high-sampling-rate action data, while long strides are suitable for low-sampling-rate data streams). This versatility allows for flexible adaptation to different data sampling conditions and application scenarios.

[0056] Example description: In the field of healthcare, surgical robots utilize multi-scale feature pyramid sets to simultaneously extract microvascular details and analyze the global surgical scene layout of endoscopic visual images. They also simultaneously parse terminology-level details and complete operational process logic for voice commands, and simultaneously capture instantaneous operational details and full-process action intentions for robotic arm movements, thereby achieving precise surgical assistance.

[0057] In the fintech business, interactive robots utilize multi-scale feature pyramid sets to simultaneously analyze customer facial expression details and overall emotional state, parse keyword commands and long business statements, understand customer micro-movements and overall interaction intentions, and improve customer service experience.

[0058] In robot operation scenarios, automated assembly robots utilize this set to simultaneously analyze the detailed features of screw holes and the overall assembly layout in the visual modality, analyze phrase commands and the entire operation intent in the language modality, and analyze local assembly details and overall path planning in the action modality.

[0059] In human-computer interaction scenarios, service robots use this set of visual input to capture the user's gaze and full-body posture, to understand word-level commands and paragraph-level requests in language, and to combine click gestures and body language in actions to provide a detailed, accurate, and natural human-computer interaction experience.

[0060] This embodiment, by independently extracting and combining multimodal, multi-scale feature levels, can fully cover the rich expression of local details and global semantics within different modalities, overcoming the limitation of traditional processing methods that cannot simultaneously capture fine-grained and coarse-grained information. The organization of the multi-scale feature pyramid set ensures the orderly combination of features of different modalities and scales, providing efficient retrieval and alignment support, and laying a comprehensive, systematic, and hierarchical foundation for subsequent multimodal collaborative alignment and deep fusion.

[0061] S30, perform cross-modal feature alignment for each feature level of the multi-scale feature pyramid set to generate a multi-scale aligned feature sequence;

[0062] In this embodiment, the goal of cross-modal feature alignment is to ensure consistent alignment of each feature level across different modalities within a unified representation space in the multi-scale feature pyramid set. The multi-scale feature pyramid set, originating from the multi-modal separation and multi-level feature extraction of the initial feature sequence, is an ordered set of hierarchical features formed in the visual, linguistic, and action modalities, respectively. Each feature level refers to a scale of the visual pyramid, a semantic level of the linguistic pyramid, or a temporal window granularity of the action pyramid. The concept of feature levels originates from the theory of "multi-granularity representation" in machine learning, which involves the division of intra-modal details into global multi-layered information.

[0063] Cross-modal feature alignment refers to mapping features at different scales—visual, linguistic, and action—to the same semantic space, enabling direct comparison, combination, and computation of these features in space. The construction of the cross-modal mapping function involves embedding each modal feature into a representation space with a unified dimension and semantic reference system using specialized linear or nonlinear projections. This operation can be traced back to the theory of shared embedding spaces. Within this unified space, cosine similarity is calculated to measure the feature similarity relationships between visual and linguistic, visual and action, and linguistic and action modes. Here, cosine similarity represents angular similarity and is used to eliminate differences in the original scale and numerical distribution of different modalities. The first, second, and third cosine similarities reflect the local semantic consistency between each pair of modalities and serve as important feedback for dynamically adjusting the parameters of the cross-modal mapping function.

[0064] The significance of adjusting the parameters of the cross-modal mapping function lies in making the spatial distribution of features from different modalities more convergent in a unified space, reducing inconsistencies and distribution drift between modalities. Through this dynamic adjustment mechanism, the final projected visual, linguistic, and action multi-scale feature layers can be more tightly and consistently distributed in the shared space, forming a multi-scale aligned feature sequence. The multi-scale aligned feature sequence not only includes the original modal information but also embeds the inter-modal consistency correction results, ensuring the efficiency of subsequent attention processing and fusion operations.

[0065] Three independent but structurally symmetrical mapping modules can be constructed to handle visual multi-scale feature layers, linguistic multi-scale feature layers, and action multi-scale feature layers, respectively. Each mapping module can employ a single-layer or multi-layer fully connected network structure, with inputs being feature tensors for each modality and outputs being feature representations with consistent embedding dimensions. The projected visual, linguistic, and action features are then fed into a cosine similarity calculation module, which sequentially calculates the similarity scores between visual and linguistic features, visual and action features, and linguistic and action features. Similarity feedback serves as an adaptive adjustment signal, updating the weights of the cross-modal mapping module through backpropagation gradient descent to improve the alignment of the mapping results.

[0066] The alignment process can employ batch normalization techniques to ensure that the mean and variance of different samples in the batch data are consistent in the embedding space, further mitigating differences in statistical distribution between modalities. During the training phase, an alignment loss function can be introduced as an optimization objective to directly minimize the similarity differences among the three modalities. The alignment operation at multiple feature levels is performed synchronously layer by layer and scale by scale; that is, the i-th layer visual features are aligned with the i-th layer language features and the i-th layer action features in the same space, ensuring consistency and traceability of cross-modal level alignment.

[0067] When a high-resolution feature level exists in the visual modality but the corresponding scale is missing in the language or action modality, interpolation padding or sampling compression can be used to make them correspond one-to-one in the scale index. This strategy works effectively not only in static multimodal input scenarios but also in dynamic multimodal inputs where the visual frame rate, speech rate, and action sampling frequency are asynchronous.

[0068] Example: In the field of healthcare, in robot-assisted surgery scenarios, cross-modal feature alignment is performed on the endoscopic details of the visual modality, the vocabulary and sentence structure of the physician's instructions in the language modality, and the micro-motion data of the robotic arm in the action modality. This ensures that the anatomical details of the tissue in the visual image correspond consistently with the physician's voice description and the current action of the robotic arm in a unified space, thus ensuring that the robot responds accurately to complex surgical steps.

[0069] In the fintech business, intelligent customer service robots use cross-modal feature alignment to map details of customer facial expressions, speech, and body language into a unified expression space. This enables precise perception of subtle emotions, text content, and micro-gesture intentions during customer consultations, thereby enhancing the intelligent experience and risk control capabilities of financial services.

[0070] In robot operation scenarios, assembly robots use cross-modal feature alignment to ensure that visual details, voice commands, and robotic arm movements at the assembly station are aligned in a consistent manner, thereby ensuring precise alignment of parts and complete execution of assembly instructions.

[0071] In human-computer interaction scenarios, service robots enhance contextual understanding and the naturalness of interaction by expressing users' visual expressions, voice commands, and interactive actions in the same space through cross-modal alignment.

[0072] This embodiment overcomes the problems of inconsistent spatial distribution of information across different modalities and scales, inconsistent numerical scales, and large differences in statistical properties by performing cross-modal feature alignment on each feature level of the multi-scale feature pyramid set. It addresses the semantic inconsistencies and alignment errors in multi-scale representations of visual, linguistic, and action multimodal data in existing methods. This process ensures that subsequent attention calculations, dependency modeling, and fusion operations can be performed in a consistent semantic space, significantly improving the model's multimodal understanding accuracy and fusion efficiency. Particularly in scenarios involving large-scale, multi-dimensional, and multimodal data, it significantly reduces semantic mismatches and redundant computations, improving real-time performance and robustness in long sequences and dynamic scenarios.

[0073] S40, perform local attention processing and global attention processing on the multi-scale aligned feature sequence to generate long-distance dependent features;

[0074] In this embodiment, performing local attention processing on the multi-scale aligned feature sequence refers to dividing the sequence into local windows according to the temporal or spatial dimensions. Within each local window, attention weights are calculated using local query vectors, local key vectors, and local value vectors, enabling fine-grained modeling of the interaction relationships between features within the window. The multi-scale aligned feature sequence originates from multi-level representations formed by unifying the alignment of visual, linguistic, and action modalities, exhibiting intermodal consistency. Local window partitioning methods include fixed-length partitioning, sliding window partitioning, or adaptive partitioning, capable of covering local continuous patterns. The local query vectors, key vectors, and value vectors can be generated using single-layer or multi-layer linear transformations, mapping the feature sequence within the local window to a low-dimensional or high-dimensional attention space representation for local correlation analysis.

[0075] The calculation of the local scaling dot product involves taking the inner product of each pair of local query vectors and local key vectors, using the square root of the feature dimension as a scaling factor to ensure numerical stability. The application of the normalized exponential function means that the scaling dot product result is transformed into a probability distribution of local attention weights using the softmax function or its variants, ensuring the sum of the weights is 1, thus enhancing the model's attribution ability within the local window. Weighted aggregation of local value vectors involves summing the local value vectors using the local attention weights to generate a local attention output sequence, preserving short-range dependency features within the window.

[0076] The local attention output sequences are concatenated along the time dimension to form a local window output sequence, building the foundation for the global context. Global attention processing, based on the local window output sequences, models long-distance interaction relationships throughout the sequence using global query vectors, global key vectors, and global value vectors. The calculation of the global scaling dot product is the same as the local method, but its scope covers the entire sequence. A normalized exponential function ensures that the global attention weights also possess probabilistic normalization properties. Finally, the global value vector is weighted by the global attention weights to generate long-distance dependency features. These long-distance dependency features possess global correlation and are used for subsequent cross-layer information interaction and multimodal fusion.

[0077] In the implementation, multi-scale aligned feature sequences can be divided according to fixed window lengths (e.g., 16 frames, 32 frames) to ensure adaptability to sequences of different lengths. The generation of local query vectors, key vectors, and value vectors is achieved using single-layer linear projection matrix parameterization, with the matrix dimension dynamically adjusted according to the feature sequence dimension. The scaled dot product result is calculated via matrix multiplication, efficiently completing large-scale parallel attention matrix operations. The normalization exponential function uses the softmax function, and the normalization dimension is chosen to match the window length. Local weighted aggregation is quickly performed using matrix multiplication to obtain a local window-level contextual representation.

[0078] After concatenating the local attention output sequence, it is input into the global attention module, and the same operation process is repeated, but the query, key, and value vector range covers the entire concatenated result. Global attention processing can support multi-head parallel computation, allowing different heads to focus on different long-range dependency patterns. The computational complexity can be further optimized, for example, by using sparse attention or approximate attention algorithms to reduce the amount of computation, adapting to scenarios with ultra-long sequence inputs.

[0079] The implementation can also incorporate masking mechanisms to ensure consistency of causal order or context, such as preventing future information leakage in language sequences. In dynamic sequences, the local window length and global attention receptive field can be adaptively adjusted to improve real-time performance and generalization ability.

[0080] Example: In the healthcare field, surgical robots can accurately understand minute tissue structures and semantic details of neighboring frames in a continuous sequence of endoscopic images through local attention processing. Global attention processing ensures that the entire surgical procedure is context-aware, guaranteeing accurate control at critical nodes based on historical global information. In the language modality, the robot can achieve contextual consistency in understanding long paragraphs of surgical instructions, avoiding misjudgments caused by relying solely on local instructions.

[0081] In the fintech business, intelligent customer service systems perform local and global attention processing on video frame sequences of customer facial expressions, text sequences of consultation language, and sequences of interactive actions in video conferencing scenarios. This enables the system to simultaneously capture short-term fluctuations in customer attitudes and long-term conversational emotional trends, thereby improving the overall service experience and risk perception capabilities.

[0082] In robot operation scenarios, assembly robots use local attention processing to focus on the microscopic details of parts and adjacent operation instructions, and global attention processing to understand the entire assembly task process and historical context, avoiding errors in the sequence of actions and ensuring the smooth completion of complex assembly tasks.

[0083] In human-computer interaction scenarios, interactive robots combine local attention and global attention, enabling them to respond quickly to short user commands while maintaining consistency in long conversation histories, thus providing a more natural interactive experience that aligns with human habits.

[0084] This embodiment, through the collaborative work of local and global attention processing, can fully model short-range correlations and long-range dependencies within a sequence, addressing the problem of insufficient ability of traditional models to capture long sequence information. It overcomes the technical bottlenecks of gradient vanishing and contextual information fragmentation, ensuring that features at different time points and scales in visual, linguistic, and action multimodal sequences can be effectively correlated. Long-range dependency features, as inputs for subsequent cross-layer information interaction and multimodal fusion, greatly improve the consistency and expressiveness of global information, especially effectively alleviating the semantic fragmentation and contextual loss problems of traditional concatenation-based fusion in multimodal and multi-scale data processing.

[0085] S50, perform cross-layer information interaction on the long-distance dependent features to generate comprehensive multi-scale features;

[0086] In this embodiment, cross-layer information interaction for long-distance dependent features refers to distinguishing the feature according to different scales or levels, separating the feature sequence corresponding to each scale level, and each feature sequence represents descriptive information of different granularities in a multi-scale context. Long-distance dependent features usually originate from the collaborative results of local and global attention mechanisms in the preceding processing steps, and already possess global context awareness capabilities, but the collaborative relationships between multi-scale levels have not yet been explicitly modeled.

[0087] For the feature sequences at each feature level, a multi-head attention mechanism is used to fuse the correlations between different modalities within that level, capturing complementary information from visual, linguistic, and action modalities within a unified level. The multi-head attention mechanism can distribute attention to the correlations at different positions or dimensions within the same feature sequence, generating locally fused features. These locally fused features preserve fine-grained connections within cross-modal contexts.

[0088] To enhance cross-scale information exchange between adjacent layers, upsampling operations are performed on the local fusion features at smaller scales to increase their resolution to match that of the larger scale layers. Upsampling can be achieved through linear interpolation, transposed convolution, or learnable interpolation modules. Downsampling operations are performed on the local fusion features at larger scales to reduce their resolution to match that of the smaller scale layers. Downsampling can be achieved through average pooling, max pooling, or strided convolution.

[0089] After adjustment, the upsampled features and adjacent larger-scale local fusion features are input into the first gated fusion unit. In this unit, a first weighting coefficient is determined. This weighting coefficient can be automatically learned using a sigmoid activation function or a normalization layer, and is used to quantify the relative importance between the two scale inputs. The features are then weighted and fused according to the weighting coefficients to obtain the first cross-layer fusion feature. Similarly, the downsampled features and adjacent smaller-scale local fusion features are input into the second gated fusion unit, a second weighting coefficient is determined, and the features are fused using the same method to generate the second cross-layer fusion feature.

[0090] Gated fusion units ensure that information flows proportionally during cross-scale interactions, preventing low-resolution or high-resolution information from biasing the overall feature representation. Finally, all cross-layer fusion results are integrated to generate comprehensive multi-scale features with multi-scale global consistency and cross-modal comprehensive representation capabilities.

[0091] In practical implementation, a fixed scale can be selected for partitioning. For example, the visual modality can be divided into layers using convolutions with different receptive fields, the language modality into layers based on phrases, sentences, and paragraphs, and the action modality into layers based on short and long time windows. The multi-head attention module generates query, key, and value vectors for different heads through independent linear mapping matrices, allowing each head to learn relevance in different dimensions.

[0092] Upsampling can use bilinear interpolation to increase the resolution to the target layer. Downsampling can use 2x2 average pooling to gradually reduce the feature resolution. The first gated fusion unit concatenates two feature streams and inputs them into a small neural network to calculate the fusion weights. The weights are constrained to the 0-1 range using sigmoid activation to ensure physical interpretability. The second gated fusion unit employs the exact same implementation strategy to ensure interactive symmetry in both the up and down directions. All cross-layer interaction module parameters can be jointly trained end-to-end with the task objective to adapt to different data modalities and task requirements.

[0093] In dynamic applications, the window length and resolution ratio of each scale can be adjusted to support a trade-off between real-time performance and performance. For example, the number of scales can be dynamically reduced on low-computing-power edge devices, while the scale granularity can be increased in high-precision applications.

[0094] Example: In the field of healthcare, surgical robots can use cross-layer information interaction to integrate small blood vessels, tissue texture features and overall anatomical structures at multiple visual scales, combine phrase-level surgical instructions and paragraph-level operation goals in the language modality, and combine minor adjustments and overall path planning in the action modality to ensure that the robot can complete complex operations accurately and stably.

[0095] In the fintech business, intelligent risk control systems can process the subtle changes in facial expressions and the wide range of behavioral patterns of video customers through cross-layer information interaction. By combining short text transaction instructions with long text contract constraints, and by combining instantaneous interactive actions with historical operation sequences, a highly credible user behavior assessment can be formed for precise fraud prevention.

[0096] In robot operation scenarios, assembly robots can integrate screw detail textures with the global assembly layout in real time through cross-layer information interaction, ensuring both operational accuracy and global consistency, and avoiding assembly errors caused by a lack of contextual support for local details.

[0097] In human-computer interaction scenarios, social robots, through cross-layer information interaction, associate the details and tone of short sentences in the dialogue with the context of long historical dialogues. This allows the robot to adjust its tone and response strategy based on the user's current statement and past context, providing a smoother and more natural interactive experience.

[0098] This embodiment significantly improves the integrity of multi-scale contextual information through cross-layer information interaction, solving the problem that traditional feature processing models can only be computed independently within the same scale. It breaks down the disconnect between feature representations at different scales, achieving collaborative expression and dynamic complementarity of information at different granularities. This results in more comprehensive, refined, and consistent contextual interpretability of the generated integrated multi-scale features, providing high-quality input for subsequent multimodal dynamic fusion and task decision output. This processing mechanism demonstrates good adaptability and scalability when modeling long-distance associations across spatial, temporal, and semantic dimensions of multimodal sequence data.

[0099] S60, Based on the comprehensive multi-scale features, dynamically fuse multi-modal information to generate the final fused features;

[0100] In this embodiment, dynamic fusion of multimodal information based on comprehensive multi-scale features refers to extracting the visual, linguistic, and action multimodal features contained within the comprehensive multi-scale features as independent components. The comprehensive multi-scale features originate from multi-level context and cross-modal interaction processing, inherently containing global consistency in multimodal expression and complementarity of information at different scales. Dynamic fusion refers to dynamically adjusting the weights of different modal information according to the context, avoiding information bias caused by fixed weights, and ensuring that visual, linguistic, and action information play their due role under different task conditions.

[0101] From the comprehensive multi-scale features, the visual feature component is a high-dimensional tensor formed after prior visual channel and multi-scale processing, which usually corresponds to the spatial structure description of the image region; the linguistic feature component is the output after prior semantic encoding and multi-level syntactic semantic aggregation, representing the contextual logical relationship in the text data; the action feature component is the expression after sensor-collected motion data is processed by multi-scale temporal convolution, reflecting the physical motion state. The three types of components are arranged in the same dimension and then input into a specially designed gating unit.

[0102] The gating unit contains multiple parallel weight paths to independently learn the importance of visual, linguistic, and action components within the current task context. Dynamic weight coefficients are automatically learned through a small fully connected network and activation functions, reflecting the contribution of each modality. Weight learning relies on the global context of the integrated multi-scale features; therefore, dynamic weight coefficients are not statically defined but rather context-sensitive parameters driven by task and input data conditions.

[0103] Based on the learned dynamic weight coefficients, the visual, linguistic, and action feature components are weighted separately to generate weighted visual, linguistic, and action features, respectively. The intramodal integrity of each feature is preserved, and the intensity ratio is adjusted. Finally, the three weighted features are weighted and summed or concatenated and mapped to a unified output space, fusing them into a single high-dimensional vector representation, forming the final fused feature. The final fused feature maintains visual spatial distribution information, linguistic semantic integrity, and action temporal continuity, with different modal information embedded in a dynamically balanced manner, improving the comprehensiveness and contextual consistency of the expression.

[0104] In practical implementation, multi-scale features can be mapped into visual, linguistic, and action feature components of the same dimension through a multi-channel linear projection layer. The gating unit adopts a multi-head gating design, with each path corresponding to the visual, linguistic, and action channels. Internally, it is processed by a single-layer perceptron, and the output is a corresponding dynamic weight coefficient. The weight coefficient is normalized between 0 and 1 through softmax or sigmoid activation to ensure numerical stability and interpretability.

[0105] The training of dynamic weights can be jointly optimized with downstream task objectives. For example, in robotics tasks, the task success rate can be used as the optimization objective, and the weight parameters and gating network weights can be adjusted together via backpropagation. In different tasks, task-conditional encoding can be added to the gating unit to make the dynamic weights sensitive to task changes. For example, visual weights can be increased in purely vision-driven tasks, and language weights can be increased in voice command tasks.

[0106] For devices with limited computing resources, the number of weight parameters in the gating unit can be reduced, for example, by using shared parameters or weight sparsity constraints. For real-time systems, dynamic weights from recent historical inputs can be cached to form a smooth output within a time window, reducing instability caused by frequent changes.

[0107] Example: In the field of healthcare, surgical robots can dynamically fuse multimodal information, combining local details from camera visual data, semantic targets of intraoperative surgeon voice commands, and motion states of robotic arm action sequences to achieve precise response to surgeon intentions and scene changes.

[0108] In the fintech business, intelligent customer service systems can dynamically integrate multimodal information, combining facial expressions in videos, emotional changes implied in voice commands, and historical click behavior sequences to generate globally consistent customer profiles for risk assessment and service recommendations.

[0109] In robot operation scenarios, assembly robots can dynamically adjust the fusion ratio of visual, linguistic, and motion information. For example, when an assembly task includes both voice guidance and motion adjustment, the weight of linguistic and motion features is automatically increased, enabling the robot to better respond to dynamic commands and real-time changes in assembly details.

[0110] In human-computer interaction scenarios, family companion robots can dynamically integrate the user's current facial expression visual information, voice content, and action behavior prediction to generate response strategies that are highly consistent with the current interaction intent, providing a warmer, more delicate, and natural interactive experience.

[0111] This embodiment effectively solves the information redundancy and bias problems caused by traditional fixed-weight fusion by dynamically fusing multimodal information. It enables visual, linguistic and action information to participate in the expression in the optimal proportion under different tasks and contexts, ensuring that the final fused features have higher expressive efficiency and context relevance. It is particularly suitable for handling complex dependencies across space, time and semantics in multimodal sequence data, and improves the accuracy and adaptability of subsequent task decisions and outputs.

[0112] S70, the final fused features are input into the task decision network to obtain the target task result.

[0113] In this embodiment, the final fused feature is input into the task decision network to obtain the target task result. This refers to the high-dimensional representation of the fused visual, linguistic, and action information generated after previous processing, which is directly fed into the task decision module as input to achieve the target output of the downstream task. The final fused feature is a multi-dimensional tensor with a unified semantic embedding space and multimodal consistency, ensuring that it can be fully utilized by subsequent models.

[0114] A task decision network is a neural network architecture that supports multiple tasks, capable of processing multimodal, high-dimensional input data and mapping it to the corresponding task output space. This network architecture typically includes multiple parts such as an input encoding layer, a feature transformation layer, and a decision output layer. The input encoding layer takes the final fused features as input and performs necessary dimensionality adjustments and normalization to ensure numerical stability. The feature transformation layer extracts the contextual relationships and patterns from the final fused features through multiple layers of nonlinear mapping, such as multilayer perceptrons, residual blocks, or Transformer encoding modules. The decision output layer designs different output structures depending on the task objective.

[0115] The task decision network supports multi-task modes. For example, when the task type is robot operation control, the network outputs a sequence of fine-grained motion parameters, such as position offsets, angle adjustment values, and speed control commands. When the task type is semantic understanding, the network outputs natural language descriptive text, such as a summary of the current scene or object recognition results. Internally, the task decision network achieves multi-task adaptation through a conditional encoding mechanism. Conditional encoding can be based on task identifiers, environmental context scalars, or historical behavioral features. Through conditional encoding, the task decision network adaptively adjusts the decision paths for different tasks within the same structure.

[0116] To ensure efficiency and real-time performance, task decision networks can be designed as lightweight deep neural network structures, such as using depthwise separable convolutions, attention pruning mechanisms, or weight quantization methods to reduce memory and computational resource consumption. Furthermore, task decision networks can be deployed in a distributed manner across different devices; for example, the input encoding layer can run on edge computing units, while subsequent feature transformation layers and decision output layers can run on high-performance servers to achieve real-time multimodal task decision-making.

[0117] In practical implementation, a multi-head Transformer network can be used as the task decision network. The input encoding layer adjusts the final fused feature dimension to meet the input dimension requirements of the Transformer through linear transformation and adds task-related positional encoding. The feature transformation layer uses stacked multi-head self-attention modules and feedforward network modules to extract deep cross-modal relationships in the final fused features. The decision output layer dynamically selects the output path according to the task type. For example, for robot control tasks, it outputs a set of real-valued vectors representing control command parameters; for semantic understanding tasks, it outputs a text sequence processed by the decoder.

[0118] The task decision network can be optimized through end-to-end training, with training data including multi-task labels and supervision targets. During training, input data can include robot operation datasets, scene description datasets, motion trajectory datasets, etc., to ensure that the task decision network can generalize across different tasks. For scene-adaptive tasks, environmental context embeddings can be introduced into the input encoding layer to improve the robustness of the task decision network in different environments.

[0119] In resource-constrained scenarios, techniques such as network pruning or knowledge distillation can be used to compress the trained task decision network into a version suitable for edge device deployment. In multi-task parallel processing scenarios, the task decision network can adopt a multi-output structure with a shared backbone and task-specific branches to improve resource utilization efficiency and task collaboration.

[0120] Example: In the healthcare business, surgical robots can input the final fused features of camera images, doctor's voice commands, and the motion status of surgical instruments into the task decision network, and output precise surgical auxiliary motion control parameters to automatically adjust the operating trajectory of the robotic arm and the posture of the stabilizer, reducing human intervention and errors.

[0121] In the fintech business, intelligent customer service robots can input the final fusion features of video footage, customer verbal communication, and mouse / gesture actions into a task decision network, and output intelligent response content or next step recommendation services to achieve efficient understanding of customer intentions and personalized services.

[0122] In robot operation scenarios, industrial robots can input the final fused features of multi-angle visual images, operator voice commands, and end effector motion states into the task decision network, and output grasping path, posture adjustment, and dynamic obstacle avoidance action parameters, enabling the robot to flexibly handle complex industrial assembly tasks.

[0123] In human-computer interaction scenarios, companion robots can input the final fusion features of user facial expressions, natural language communication, and user action patterns into the task decision network, and output emotionally adaptive interactive behaviors and voice responses to improve the friendliness and personalization of the interaction.

[0124] This embodiment, by inputting the final fused features into the task decision network, enables adaptive fusion and utilization of multimodal information at the decision level, improves the understanding of complex multimodal inputs, and enhances the accuracy and real-time performance of multi-task outputs. In particular, when facing cross-modal, cross-scale, and long-term tasks, it can effectively alleviate the shortcomings of existing methods in multi-task adaptation and long-distance dependency representation, and support unified and efficient decision-making under multi-task and multi-environment conditions.

[0125] This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as intelligent agent autonomous decision-making, fintech, and healthcare. It discloses a method, apparatus, device, and medium for processing multimodal sequence data, comprising: acquiring raw visual sequence data, raw language text sequence data, and raw motion sensor sequence data and generating an initial feature sequence; extracting multi-scale feature levels from the initial feature sequence and combining them into a multi-scale feature pyramid set; performing cross-modal feature alignment on each feature level of the multi-scale feature pyramid set to generate a multi-scale aligned feature sequence; performing local attention processing and global attention processing on the multi-scale aligned feature sequence to generate long-distance dependent features; performing cross-layer information interaction on the long-distance dependent features to generate comprehensive multi-scale features; dynamically fusing multimodal information based on the comprehensive multi-scale features to generate a final fused feature; and finally inputting the fused feature into a task decision network to obtain the target task result. This invention effectively addresses the problem of insufficient information correlation between different modalities at different scales in multimodal long sequence data by using multi-scale feature pyramid sets, cross-modal feature alignment, combination of local and global attention, and cross-layer information interaction. It achieves refined modeling and dynamic fusion of multi-scale features of different modalities, thereby improving the processing effect and accuracy of multimodal tasks in complex scenarios.

[0126] In one embodiment, step S10 includes:

[0127] S101, capture a continuous image stream of the target scene through an image acquisition device, generate raw visual sequence data, perform frame rate normalization processing on the raw visual sequence data to generate a visual sequence, and perform size normalization processing on the visual sequence to generate a standardized visual sequence.

[0128] S102, the speech signal is collected by the speech recognition device and converted into text, or the text input stream is received to generate the original data of the language text sequence, the encoding format of the original data of the language text sequence is uniformly processed to generate the language text sequence, and the language text sequence is segmented and part-of-speech tagging is performed to generate the segmented tagging sequence.

[0129] S103: Physical motion signals are acquired through an inertial measurement unit, force sensor, or joint encoder to generate raw data of motion sensor sequence. The raw data of motion sensor sequence is processed by sampling rate synchronization to generate motion sensor sequence. The motion sensor sequence is then processed by noise filtering and numerical standardization to generate standardized motion sequence.

[0130] S104, The standardized visual sequence is processed by a convolutional neural network to generate a visual feature vector sequence;

[0131] S105, The word segmentation and annotation sequence is processed by the word embedding model to generate a word embedding vector sequence, and position encoding information is added to the word embedding vector sequence to generate a language feature vector sequence;

[0132] S106, The standardized action sequence is processed through a fully connected network to generate an action feature vector sequence;

[0133] S107, the visual feature vector sequence, language feature vector sequence and action feature vector sequence are concatenated in time stamp order to generate an initial feature sequence.

[0134] In this embodiment, the process of acquiring raw visual sequence data, raw language text sequence data, and raw motion sensor sequence data, and generating an initial feature sequence, involves a comprehensive preprocessing and feature extraction process for multimodal data. This ensures that subsequent fusion and analysis can be performed within a unified semantic space and temporal coordinate system. A continuous image stream of the target scene is captured using an image acquisition device, resulting in high-dimensional, unstructured pixel matrix data that directly reflects objects, backgrounds, environmental textures, and other content in the visual space. After acquisition, the raw visual sequence data undergoes frame rate normalization. Specifically, this is achieved by first statistically analyzing the distribution of inter-frame time intervals in the input sequence, and then using a temporal resampling algorithm to make the frame intervals more consistent. For example, linear interpolation can be used to fill missing frames, or redundant frames can be extracted, thus ensuring that the visual sequence is aligned with other modal data in the temporal dimension. Size normalization is indispensable in standardizing the visual sequence. By fixing the input size (e.g., 224×224 pixels), scale deviations caused by different acquisition resolutions, shooting distances, or angles are eliminated. This can be achieved using bilinear interpolation or region average pooling to uniformly map the raw frames to a predefined input tensor size.

[0135] The acquisition of language text sequences involves two types of input sources. The first type is speech signals collected by speech recognition devices, whose processing includes acoustic feature extraction (such as Mel-frequency cepstral coefficients, MFCC), decoding and matching (based on joint probability optimization of acoustic and language models), and conversion into text strings. The second type is directly received text input streams. Unified encoding format processing ensures consistency in character encoding across text sequences from different sources. For example, all input text is uniformly converted to UTF-8 encoding format to avoid inconsistencies caused by multilingual and multi-terminal input. Word segmentation and part-of-speech tagging are performed using natural language processing toolchains (such as sequence labelers based on Conditional Random Fields, CRF) to parse the text and generate sequences containing word units and their corresponding part-of-speech tags, providing refined syntactic support for subsequent semantic modeling.

[0136] Acquiring raw motion sensor sequence data involves raw physical signals output from inertial measurement units, force sensors, or joint encoders. These signals may have inconsistent sampling rates and varying noise levels along the time axis. Sampling rate synchronization processing uses timestamp alignment and interpolation resampling algorithms to ensure that the outputs of different sensors are aligned on a unified time baseline, avoiding temporal misalignment issues in multi-source motion data. Noise filtering uses bandpass filters (such as Kalman filters or low-pass filters) to remove high-frequency noise and equipment jitter interference. Numerical standardization processing uses linear transformations (such as Z-score standardization) to ensure that the outputs of each sensor are distributed within a standard normal distribution range with zero mean and unit variance, improving the model's robustness to differences in cross-sensor data distribution.

[0137] The processing of standardized visual sequences employs convolutional neural networks to map spatial pixel matrices into sequences of visual feature vectors. This is implemented using stacked layers of deep convolutional networks, such as ResNet or a lightweight MobileNet, to extract multi-level visual features including local texture, shape contours, and object semantics. Word segmentation and annotation sequences are processed through word embedding models, mapping them into high-dimensional, dense semantic vector sequences. This can be achieved using pre-trained word vectors (such as Word2Vec or GloVe) or context-dynamic word vectors (such as BERT embeddings), with positional encoding added to these embedding vectors. This allows the model to perceive the order and contextual relationships of words within the text sequence during the input phase. Standardized action sequences are processed through fully connected networks, mapping temporal action sensor signals into dense sequences of action feature vectors. This process can be implemented using stacked multilayer perceptrons, with ReLU activation units in the hidden layers and the output layer's dimensions adjustable according to task requirements.

[0138] The process of concatenating visual feature vector sequences, language feature vector sequences, and action feature vector sequences in timestamp order represents a strict alignment of multimodal features along the temporal dimension. Specifically, this involves aligning the three types of feature sequences using a timestamp-based sliding window, interpolating to fill in missing modalities, and forming a complete tensor sequence with aligned time steps, which serves as the initial feature sequence. This initial feature sequence possesses a unified temporal reference and semantic embedding space, laying a structured and highly consistent data foundation for subsequent cross-modal and cross-scale fusion.

[0139] This embodiment, through the above-described process, standardizes, extracts, and aligns the raw data of visual, linguistic, and action modalities through preprocessing, feature extraction, and temporal alignment, ultimately generating an initial feature sequence that can be directly used for subsequent multimodal modeling. This eliminates problems such as differences in multi-source data acquisition conditions, inconsistent encoding formats, mismatched sampling rates, and timeline misalignment, providing a unified and high-quality input foundation for subsequent feature fusion, interaction, and decision-making, and improving the robustness, accuracy, and applicability of the multimodal sequence data processing system.

[0140] In one embodiment, step S20 above includes:

[0141] S201, Separate the visual feature component sequence, the language feature component sequence, and the action feature component sequence from the initial feature sequence;

[0142] S202, extract the edge texture feature level of the visual feature component sequence through the shallow convolutional layer of the feature pyramid network;

[0143] S203, extract the object semantic feature level of the visual feature component sequence through the deep convolutional layer of the feature pyramid network;

[0144] S204, the edge texture feature level and the object semantic feature level are fused to generate a visual multi-scale feature level;

[0145] S205, The language feature component sequence is processed by the aggregation module to generate a phrase-level language feature hierarchy;

[0146] S206, The aggregation module aggregates the phrase-level language feature layers to generate a sentence-level language feature layer;

[0147] S207, The aggregation module aggregates the sentence-level language feature layers to generate a paragraph-level language feature layer;

[0148] S208, combine the phrase-level language feature layer, the sentence-level language feature layer, and the paragraph-level language feature layer to generate a multi-scale language feature layer;

[0149] S209, extract the action detail feature level of the action feature component sequence through small time window convolution of the hierarchical temporal convolutional network;

[0150] S210, extract the action trend feature level of the action feature component sequence through large temporal window convolution of the hierarchical temporal convolutional network;

[0151] S211, combine the action detail feature level and the action trend feature level to generate an action multi-scale feature level;

[0152] S212, combine the visual multi-scale feature layers, language multi-scale feature layers, and action multi-scale feature layers to generate a multi-scale feature pyramid set.

[0153] In this embodiment, the operation of extracting multi-scale feature levels from the initial feature sequence and combining them into a multi-scale feature pyramid set aims to construct a hierarchical and scale-based representation structure to meet the needs of subsequent cross-modal and cross-scale collaborative processing, addressing the diversity of data features across different modalities and scales. First, visual feature component sequences, language feature component sequences, and action feature component sequences are separated from the initial feature sequence. This separation process is based on modality identifiers, using indexing or channel segmentation to directly decompose the multimodal concatenation tensor into individual modality component sequences. This allows each modality feature to independently enter its corresponding multi-scale feature processing module, ensuring that information between modalities is not confused.

[0154] For the visual feature component sequence, shallow convolutional layers of a feature pyramid network are used to extract edge texture features. These shallow convolutional layers capture fine-grained relationships between local pixels through a smaller receptive field, such as low-level features like boundaries, corners, and textures. In practice, multiple progressively layered two-dimensional convolutional operations combined with batch normalization and non-linear activation units can be used to ensure stable output features and good gradient fluidity. The deep convolutional layer extraction process for the visual feature component sequence focuses on global contextual relationships within a larger receptive field, used to encode higher-order semantic features such as object category, positional relationships, and scene layout. This process is achieved through multiple downsampling, channel expansion, and non-linear mapping. The fusion of the edge texture feature layer and the object semantic feature layer can be achieved by lateral connection of the feature pyramid. High-resolution, low-semantic-intensity shallow features and low-resolution, high-semantic-intensity deep features are integrated through upsampling and element-wise weighted summation to form a multi-scale visual feature layer. This structure has the ability to simultaneously express local and global visual information.

[0155] The processing of the language feature component sequences is completed by the aggregation module. First, a phrase-level language feature layer is generated. Phrase-level language features are aggregated into word embedding vectors of consecutive words using a sliding window approach with a fixed number of words, and weighted averaging or convolutional encoding is used to express local semantic combinations. The sentence-level language feature layer continues to aggregate longer semantic units based on the phrase-level language features through windowing or pooling, strengthening cross-phrase grammatical dependencies. The paragraph-level language feature layer further aggregates multiple sentence-level language features to form a semantic abstraction of logical and thematic consistency in long texts. Finally, by combining the phrase-level, sentence-level, and paragraph-level language feature layers, a multi-granularity, multi-contextuality multi-scale language feature layer is formed.

[0156] The motion feature component sequence is processed through small-window convolutions in a hierarchical temporal convolutional network to extract motion detail features. Small-window convolutions use short kernels that slide along the time dimension, making them sensitive to local motion changes, such as small, rapid changes in a single joint. Conversely, large-window convolutions are used to extract motion trend features, with kernels covering a longer time period, expressing the dynamic trends and rhythmic patterns of slow, continuous motion. The combination of motion detail features and motion trend features can be achieved through channel concatenation or element-wise weighted summation, enabling the multi-scale motion feature layers to simultaneously express both detailed and global motion.

[0157] Finally, the visual multi-scale feature layers, language multi-scale feature layers, and action multi-scale feature layers are combined into a multi-scale feature pyramid set. The combination process uses multimodal tensor concatenation or alignment mapping, so that the multi-scale layers of the three modalities are represented in a unified data structure while retaining their independent hierarchical information, laying a structured representation foundation for cross-modal fusion and multi-scale interaction.

[0158] This embodiment effectively addresses the problem of insufficient information representation capabilities of different modalities at different scales in existing models by performing multimodal separation, multi-scale feature decomposition, and multi-level aggregation on the initial feature sequence, and unifying them into a multi-scale feature pyramid set. Layered extraction and fusion of edge texture and object semantic features enable the visual modality to consider both local details and global semantics; phrase, sentence, and paragraph-level language feature extraction allows the language modality to encompass semantic details and overall text logic; and layered extraction and combination of action details and action trends enable the action modality to simultaneously express short-term rapid changes and long-term dynamic trends. The multi-scale feature pyramid set, as a structured input, provides a unified foundation for the alignment, interaction, and fusion of different modalities and scales in subsequent steps, improving the ability to represent and process complex multimodal and multi-scale information, and enhancing the system's robustness and generalization ability to cross-modal sequence data in dynamic environments.

[0159] In one embodiment, step S30 above includes:

[0160] S301, construct cross-modal mapping functions for the visual multi-scale feature hierarchy, language multi-scale feature hierarchy and action multi-scale feature hierarchy of the multi-scale feature pyramid set;

[0161] S302, the visual multi-scale feature layer, the language multi-scale feature layer, and the action multi-scale feature layer are projected into a unified semantic space through the cross-modal mapping function;

[0162] S303, within the unified semantic space, determine the first cosine similarity between the visual multi-scale feature level and the language multi-scale feature level, the second cosine similarity between the visual multi-scale feature level and the action multi-scale feature level, and the third cosine similarity between the language multi-scale feature level and the action multi-scale feature level.

[0163] S304, Adjust the parameters of the cross-modal mapping function based on the first cosine similarity, the second cosine similarity and the third cosine similarity;

[0164] S305, the visual multi-scale feature layers, language multi-scale feature layers, and action multi-scale feature layers are projected through the adjusted cross-modal mapping function to generate a cross-modal aligned multi-scale feature sequence.

[0165] In this embodiment, cross-modal feature alignment is performed on each feature level in the multi-scale feature pyramid set. The goal is to establish the correspondence between visual, linguistic, and action modalities at the same semantic scale, and to solve the problem of information inconsistency and difficulty in fusion caused by the heterogeneity of different modalities in the semantic expression space. First, cross-modal mapping functions are constructed for the visual multi-scale feature level, the linguistic multi-scale feature level, and the action multi-scale feature level, respectively. The cross-modal mapping function can be implemented through a parameterized neural network structure, such as a single-layer or multi-layer feedforward fully connected network. The feature distribution characteristics of each modality are specifically modeled, with the aim of mapping the multi-scale features of each modality to a latent expression space with consistent dimensions and distribution. This latent space is defined as a unified semantic space. The visual multi-scale feature level includes edge texture and semantic object features, the linguistic multi-scale feature level includes phrase, sentence, and paragraph-level semantic expressions, and the action multi-scale feature level includes detailed dynamics and global trend expressions. The design of the mapping function needs to ensure that it can adapt to the heterogeneity of various features in terms of statistical distribution and scale differences.

[0166] By processing each modality through a cross-modal mapping function, visual multi-scale feature layers, linguistic multi-scale feature layers, and action multi-scale feature layers are all projected into a unified semantic space, enabling the three modalities to have an aligned representational basis within this space. In the unified semantic space, cosine similarity is used as a method to measure the semantic relevance between features of different modalities. Cosine similarity can be obtained through standard inner product and normalization calculations, quantifying the first cosine similarity between visual and linguistic multi-scale feature layers, the second cosine similarity between visual and action multi-scale feature layers, and the third cosine similarity between linguistic and action multi-scale feature layers. These similarity values ​​reflect the degree of semantic consistency between different modalities at the current scale.

[0167] Based on three sets of cosine similarity values, the parameters of the cross-modal mapping function are dynamically adjusted. Parameter adjustment can be achieved through gradient descent optimization with the goal of minimizing the cosine distance between modalities, ensuring that the projected features of different modalities are positioned as close as possible in a unified semantic space, thereby strengthening semantic consistency between different modalities. The adjusted cross-modal mapping function is then applied to all feature levels of the multi-scale feature pyramid set to complete the final cross-modal alignment operation.

[0168] Finally, the cross-modal mapping function, after parameter adaptive optimization, reprojects the visual multi-scale feature layers, linguistic multi-scale feature layers, and action multi-scale feature layers, generating a cross-modal aligned multi-scale feature sequence. This aligned feature sequence exhibits cross-modal and cross-scale consistency in a unified semantic space, serving as a unified foundation for subsequent cross-modal fusion and global modeling.

[0169] This embodiment addresses the alignment difficulties caused by inconsistent semantic space distributions of visual, linguistic, and action features across different modal levels of the multi-scale feature pyramid set by constructing cross-modal mapping functions and adaptively adjusting parameters. After projection into a unified semantic space and similarity-driven optimization, the visual, linguistic, and action modalities exhibit higher semantic consistency at a unified scale. This ensures that subsequent models can fully utilize complementary information between different modalities during multi-modal fusion and cross-scale global inference, reducing information bias and redundancy, and improving the overall system's modeling accuracy and adaptability to complex multi-modal inputs.

[0170] In one embodiment, step S40 above includes:

[0171] S401, the multi-scale aligned feature sequence is divided into multiple local windows according to a preset window length;

[0172] S402, generate a local query vector, a local key vector, and a local value vector for the multi-scale aligned feature sequence within each local window;

[0173] S403, determine the local scaling dot product of the local query vector and the local key vector within each local window;

[0174] S404, Apply a normalized exponential function to the local scaling dot product result within each local window to generate local attention weights;

[0175] S405, the local value vectors within the local window are weighted and aggregated according to the local attention weights to generate a local attention output sequence;

[0176] S406, The local attention output sequence is spliced ​​together in the time dimension to form a local window output sequence;

[0177] S407, Generate a global query vector, a global key vector, and a global value vector for the local window output sequence;

[0178] S408, determine the global scaled dot product result of the global query vector and the global key vector;

[0179] S409, Apply a normalized exponential function to the global scaling dot product result to generate global attention weights;

[0180] S410, the global value vector is weighted according to the global attention weight to generate long-distance dependency features.

[0181] In this embodiment, local and global attention processing is performed on the multi-scale aligned feature sequence to generate long-distance dependency features, aiming to simultaneously capture local fine-grained dependencies and long-distance interaction patterns across the global scope. First, the multi-scale aligned feature sequence is divided into multiple local windows according to a preset window length. This division can be achieved using time step indexing or frame numbering, and the window length can be set to a fixed length or dynamically adjusted according to task requirements. The division into local windows confines subsequent computation to a smaller range, reducing computational complexity while enhancing local context resolution capabilities.

[0182] Within each local window, local query vectors, local key vectors, and local value vectors are generated from the multi-scale aligned feature sequence. This process can be implemented using a linear projection layer or a learnable weight matrix. The mapping dimension can be preset or adaptively adjusted based on the input data to ensure comparability of different feature dimensions in the same representation space. The locally scaled dot product between the local query vector and the local key vector is divided by the square root of the vector dimension after element-wise inner product calculation to mitigate numerical instability caused by the increase in vector dimension.

[0183] The local scaling dot product result is normalized by applying a normalized exponential function (typically using the Softmax function) to each local window, mapping the scaling dot product result to a probability distribution between 0 and 1, thus obtaining the local attention weights. These local attention weights can be viewed as the importance distribution of features at each time step within the window. The local value vectors are then weighted and aggregated according to these weights to obtain the local attention output sequence. This weighted aggregation is performed using matrix multiplication; the product of the weight distribution and the value vector matrix is ​​the output sequence.

[0184] The output sequences of each local window are concatenated along the time dimension to form a local window output sequence. This operation can be regarded as the temporal restoration of local context encoding, ensuring that the input in the subsequent global processing stage maintains a complete temporal structure.

[0185] For the local window output sequence, a global query vector, a global key vector, and a global value vector are further generated. The generation of the global representation is similar to that of the local representation, relying on linear mapping operations within the global scope. The global scaled dot product of the global query vector and the global key vector is calculated, and then scaled by dividing by the square root of the global feature dimension after the inner product calculation to alleviate numerical issues. A normalized exponential function is applied to the global scaled dot product result to generate global attention weights, which are used to measure the relative importance of the global context at different time steps.

[0186] The global value vector is weighted based on the global attention weights, and a global weighted aggregation operation is performed to obtain long-range dependency features. The final output long-range dependency features take into account both local contextual information and global contextual relationships, providing a unified basis for global contextual representation for subsequent cross-layer information interaction and multimodal fusion steps.

[0187] This embodiment divides the multi-scale aligned feature sequence into local windows and performs local attention computation independently within each window, which can fully capture fine-grained dependencies in the local context and enhance the model's ability to resolve local spatiotemporal features. Based on local encoding, a global context representation is formed by concatenating the representations along the time dimension, and further, long-distance dependencies across time and modality are captured through global attention computation. This effectively overcomes the gradient vanishing problem caused by the increasing sequence length in traditional recurrent neural networks and the excessive computational complexity of Transformer-type models in long sequence tasks. The overall operation enables simultaneous modeling of interactions between different times and modalities at both the local detail and global relationship scales, improving the model's ability to express and resolve complex multimodal sequence inputs, thereby enhancing the accuracy and robustness of subsequent cross-layer interactions and task decision networks.

[0188] In one embodiment, step S50 above includes:

[0189] S501, the long-distance dependent features are separated into feature sequences of multiple feature levels according to the feature level;

[0190] S502, For the feature sequence of each feature level, the multi-head attention mechanism is used to fuse the multimodal features within the feature level to generate the local fused features of the feature level;

[0191] S503: For each pair of local fusion features at adjacent feature levels, the resolution of the smaller scale feature is adjusted to the resolution of the adjacent larger scale feature through an upsampling operation to obtain the upsampled feature, and the resolution of the larger scale feature is adjusted to the resolution of the adjacent smaller scale feature through a downsampling operation to obtain the downsampled feature.

[0192] S504, the upsampled features and adjacent larger-scale local fusion features are input into the first gated fusion unit, and a first weight coefficient is determined in the first gated fusion unit;

[0193] S505, Based on the first weighting coefficient, the upsampled feature is fused with the adjacent larger-scale local fusion feature to generate the first cross-layer fusion feature;

[0194] S506, the downsampled features and adjacent smaller-scale local fusion features are input into the second gated fusion unit, and a second weighting coefficient is determined in the second gated fusion unit;

[0195] S507, Based on the second weighting coefficient, the downsampling feature is fused with the adjacent smaller-scale local fusion feature to generate a second cross-layer fusion feature;

[0196] S508, merge the first cross-layer fusion feature and the second cross-layer fusion feature to generate a comprehensive multi-scale feature.

[0197] In this embodiment, cross-layer information interaction is performed on long-range dependency features to generate comprehensive multi-scale features. First, the long-range dependency features need to be separated according to feature hierarchy. This separation operation is achieved through scale-based labeling or grouping mechanisms. Different levels can be extracted based on predefined hierarchical identifiers in the original multi-scale feature pyramid set, ensuring that each feature sequence contains only long-range dependency expressions from the same level. This step is the foundation for subsequent independent fusion of each level, guaranteeing that the multi-scale representation has a clear structure at different resolutions and semantic depths.

[0198] For each feature sequence at each feature level, a multi-head attention mechanism is used for fusion. This mechanism uses parallel attention heads to calculate similarity weights for different subspaces within the feature sequence, enabling features from different modalities to interact based on different attention perspectives at the same feature level, capturing fine-grained cross-modal relationships. The outputs of each attention head are concatenated along the feature dimension and then linearly transformed back to the output dimension, forming locally fused features at the feature level.

[0199] After independent fusion of each feature level, cross-layer information interaction is performed on the local fused features of each pair of adjacent feature levels. For smaller-scale features, their resolution is adjusted to match that of adjacent larger-scale features through upsampling operations. Upsampling operations can be implemented using bilinear interpolation or learnable deconvolution to ensure that smaller-scale features are spatially aligned with larger-scale features. For larger-scale features, their resolution is adjusted to match that of adjacent smaller-scale features through downsampling operations. Downsampling operations can be implemented using average pooling, max pooling, or strided convolution. Upsampling and downsampling operations ensure that features at different levels are comparable in both spatial and temporal resolution, facilitating subsequent fusion.

[0200] The upsampled features and adjacent larger-scale local fusion features are input into the first gated fusion unit. The first gated fusion unit uses learnable parameters to determine a first weight coefficient, which is used to adjust the fusion ratio of the upsampled features and the larger-scale features. The weight coefficient is typically output through a sigmoid activation function, ensuring the weights are between 0 and 1. Based on the first weight coefficient, the upsampled features and the larger-scale features are weighted and summed to generate the first cross-layer fusion feature.

[0201] Similarly, the downsampled features and adjacent smaller-scale local fusion features are input into the second gated fusion unit. A second weighting coefficient is determined in the second gated fusion unit to adjust the fusion ratio between the downsampled features and the smaller-scale features. The second cross-layer fusion feature is generated through weighted fusion using the second weighting coefficient.

[0202] Finally, the first and second cross-layer fusion features are fused using a weighted or concatenated method to form a comprehensive multi-scale feature. This comprehensive multi-scale feature, while retaining the independent expression of each scale, further introduces cross-layer information interaction to achieve the sharing and enhancement of multi-scale context, providing a comprehensive representation with multi-resolution and multi-semantic levels, and laying a unified feature foundation for subsequent multi-modal fusion and task decision-making.

[0203] This embodiment enhances the internal consistency of features at different resolutions by independently separating and fusing long-distance dependent features at different scale feature levels. Based on this, upsampling and downsampling operations are used to align features at adjacent scales in terms of resolution, ensuring lossless information transmission. A gating unit adaptively adjusts the feature fusion weights between different scales, improving the flexibility and adaptability of cross-layer interactions. The final fused comprehensive multi-scale features balance detail and global representation, exhibiting stronger robustness and generalization ability, enabling subsequent modules to achieve higher accuracy in multimodal information parsing and task decision-making when facing complex multi-scale data.

[0204] In one embodiment, step S60 above includes:

[0205] S601, visual feature components, language feature components and action feature components are separated from the integrated multi-scale features;

[0206] S602, input the visual feature components, language feature components and action feature components into the gating unit, and determine the dynamic weight coefficients of the visual feature components, language feature components and action feature components in the gating unit.

[0207] S603, The visual feature components are weighted according to the dynamic weight coefficients of the visual feature components to generate weighted visual features;

[0208] S604, The language feature components are weighted according to the dynamic weight coefficients of the language feature components to generate weighted language features;

[0209] S605, The motion feature components are weighted according to the dynamic weight coefficients of the motion feature components to generate weighted motion features;

[0210] S606, the weighted visual features, weighted language features, and weighted action features are fused to generate the final fused features.

[0211] In this embodiment, the process of dynamically fusing multimodal information based on comprehensive multi-scale features to generate the final fused features first involves accurately separating visual feature components, linguistic feature components, and action feature components from the comprehensive multi-scale features. This separation operation relies on the structured feature encoding maintained by the multi-scale features in the preceding processing stage, where each modality's features have an independent dimension or position index in the tensor representation. By clearly defining the index range, the corresponding visual tensor, linguistic tensor, and action tensor can be efficiently extracted directly from memory, ensuring modal consistency and data integrity in subsequent weighted processing. Visual feature components typically contain the local and global structure of the image spatial context, linguistic feature components express the semantic relationships of words, phrases, and sentences in natural language, and action feature components reflect dynamic changes such as motion trajectory, acceleration, and angular velocity in a time series. This separation stage is not only a logical division but also provides the foundation for multimodal weight calculation.

[0212] After separation, the three modal components are input in parallel to the gating unit. The gating unit employs a learnable multilayer perceptron, extracting the global representation of each modal component through global mean pooling or statistical summarization of the input features. These representations are then mapped to a scalar space via fully connected layers, and finally normalized to continuous dynamic weight coefficients in the interval [0,1] using an activation function such as the Sigmoid function. The dynamic weight coefficients of the visual feature components characterize the importance of the visual modality under the current input conditions, while the dynamic weight coefficients of the language and action feature components characterize the relative importance of the language and action modalities, respectively. This dynamic weight determination process, combined with the current task context, allows the model to adaptively adjust the proportion of each modality in the overall representation based on the complexity and responsiveness of the input content, avoiding the risk of rigidity or overfitting caused by fixed weights for different modalities.

[0213] Subsequently, for each modal component, a weighting operation is performed using its corresponding dynamic weight coefficients. This weighting operation is expanded element-wise along the tensor dimension, generating weighted visual features, weighted linguistic features, and weighted action features through multiplication of scalar weight coefficients with the entire modal tensor. This weighting not only adjusts the modal contribution globally but also finely adjusts the intensity of modal information at the element level, improving the accuracy of the weighted features.

[0214] After obtaining the weighted visual, linguistic, and action features, they need to be fused to generate a unified final fused feature. Fusion can be achieved in various ways, with the optimal approach being concatenation along the channel dimension followed by adjustment of dimensional consistency using a unified linear mapping. This method stacks weighted visual, linguistic, and action feature tensors along the channel dimension to form a high-dimensional joint feature, which is then compressed or expanded into a standardized, unified output dimension using a fully connected linear layer. Compared to simple summation, this fusion method maximizes the preservation of the unique information expression and complementarity of each modality while avoiding alignment problems caused by dimensional inconsistencies between different modalities.

[0215] Example: In the field of healthcare, a multimodal intelligent sensing system for surgical assistance is constructed. This system is used to assist surgeons in performing complex surgical procedures in real time and provide decision support.

[0216] First, the system acquires raw data from visual sequences, speech / text sequences, and motion sensor sequences. Visual data is captured by multi-angle high-definition cameras installed in the operating room, capturing a continuous stream of images of the patient's surgical area in real time. Frame rate normalization eliminates differences in frame rates between different cameras, ensuring consistent visual sequences. Subsequently, size normalization converts images of different resolutions into standardized visual sequences, guaranteeing uniformity and comparability in subsequent visual feature extraction. Speech data is acquired through a speech recognition device integrated into the operating table, including real-time voice communication between surgical team members or surgical plan text entered by the surgeon. The speech / text undergoes standardized encoding to generate a speech / text sequence, which is then segmented and tagged with parts of speech to form a semantically structured sequence. Motion data originates from the inertial measurement unit (IMU) worn by the surgeon and force sensors on surgical instruments, collecting physical motion signals from the surgeon's hand movements. Motion sensor data undergoes sampling rate synchronization to ensure consistency in the timeline of different types of motion data. Noise filtering and numerical normalization are then applied to obtain standardized motion sequences.

[0217] A standardized visual sequence is input into a convolutional neural network to extract a sequence of visual feature vectors. These visual features cover the texture, instrument shape, and tissue boundary information of the surgical area. A word segmentation and annotation sequence is input into a word embedding model to obtain word embedding vectors and add positional encoding, forming a sequence of linguistic feature vectors to ensure the semantic accuracy of linguistic information in the temporal dimension. A standardized action sequence is processed by a fully connected network to generate a sequence of action feature vectors reflecting the surgeon's force, speed, and trajectory characteristics. These three types of features are concatenated in timestamp order to form an initial feature sequence, which serves as the basis for subsequent unified input.

[0218] On the initial feature sequence, visual, linguistic, and action feature components are first separated. Visual feature components are processed through shallow convolutional layers of a feature pyramid network to extract edge texture features of the surgical region, and deep convolutional layers to extract semantic features of the organism (e.g., semantic labels for scalpels, blood vessels, and organs). These two types of visual feature layers are fused to generate a multi-scale visual feature layer. Linguistic feature components are aggregated from bottom to top through an aggregation module, progressively generating phrase-level, sentence-level, and paragraph-level linguistic feature layers, reflecting the contextual logic of medical terminology and surgical plans. Action feature components are processed through small-time-window convolutions in a hierarchical temporal convolutional network to extract detailed surgical action features, and through large-time-window convolutions to extract surgical operation trend features; these two are combined to form a multi-scale action feature layer. The combination of these three types of multi-scale feature layers constitutes a multi-scale feature pyramid set, fully expressing the multimodal and multi-granular information of the surgical scene.

[0219] To achieve cross-modal collaborative perception, the system constructs cross-modal mapping functions for the visual, linguistic, and action multi-scale feature levels of the multi-scale feature pyramid set, projecting them onto a unified semantic space. Within this space, the cosine similarity between visual and linguistic, visual and action, and linguistic and action multi-scale feature levels is calculated sequentially to quantitatively assess the semantic associations between each modality. Based on the similarity, the parameters of the cross-modal mapping functions are dynamically adjusted to ensure consistent alignment of features from different modalities in the semantic space, ultimately outputting a cross-modal aligned multi-scale feature sequence.

[0220] On the multi-scale aligned feature sequence, local windows are divided according to a preset window length. Local query, key, and value vectors are generated for the feature sequence within each local window. Local attention weights are calculated using scaling dot products and a normalized exponential function to guide the weighted aggregation of local value vectors, resulting in a local attention output sequence. The local window outputs are then concatenated along the time dimension to form the overall local window output sequence. Global query, key, and value vectors are further generated for this sequence. Global scaling dot products are calculated, and global attention weights are generated using a normalized exponential function. After weighted aggregation of the global value vectors, long-distance dependency features are output, capturing key information associations across time periods throughout the entire surgical process.

[0221] Long-distance dependent features are separated according to feature hierarchy. For each level, a multi-head attention mechanism is used to fuse visual, linguistic, and action multimodal features within that level to generate local fused features. For adjacent levels, smaller-scale local fused features are upsampled to a larger scale, and larger-scale local fused features are downsampled to a smaller scale, resulting in upsampled and downsampled features. The upsampled and larger-scale local fused features are input into a first gating fusion unit to determine the first weight coefficients and are fused to generate the first cross-layer fused feature. The downsampled and smaller-scale local fused features are input into a second gating fusion unit to determine the second weight coefficients and are fused to generate the second cross-layer fused feature. The two are finally fused to generate a comprehensive multi-scale feature, forming a rich interactive expression between multimodal and multi-scale modes.

[0222] Based on the comprehensive multi-scale features, visual, linguistic, and action feature components are separated and input into the gating unit. Dynamic weight coefficients for each of the three components are determined, and each component is weighted according to the weights, outputting weighted visual features, weighted linguistic features, and weighted action features respectively. The three weighted features are concatenated along the channel dimension and compressed and integrated using a unified linear mapping to finally generate a unified final fused feature, which serves as the input to the task decision network.

[0223] The task decision network utilizes final fusion features to output various types of results for surgical scenarios: In real-time surgical assistance, it outputs automated instrument recommendations and surgical path prompts based on a comprehensive understanding of vision, language, and movement. In surgical record archiving, it outputs surgical semantic summaries and standardized surgical operation logs. In intraoperative human-computer interaction, it identifies the surgeon's intentions and provides operational suggestions and reminders through voice feedback.

[0224] In the fintech field, we are building a multimodal intelligent customer service and risk assessment system, which is applied to efficient interaction between financial institutions and users and personalized risk control decisions.

[0225] First, the system acquires raw data from visual sequences, language text sequences, and motion sensor sequences. Visual data originates from continuous video frames captured in real-time by camera equipment in scenarios such as financial service counters, video customer service terminals, or remote account opening. Frame rate standardization eliminates differences in acquisition frequencies across different hardware devices, unifies the temporal dimension of the visual sequences, and then size normalization adjusts the video frames to a uniform resolution for subsequent feature processing. Language data is collected from conversations between users and customer service representatives or smart terminals via speech recognition devices, or from natural language text input via mobile devices and web pages recording user-submitted requests. Language text undergoes unified encoding format processing to ensure compatibility with character set standards from different input sources, and further performs word segmentation and part-of-speech tagging to extract structured semantic units of the language content. Motion data is collected by user behavior sensors on smart teller machines or video terminals, covering behavioral information such as clicks, gesture trajectories, and body movements. Motion sensor data undergoes sampling rate synchronization, noise filtering, and numerical standardization to ensure consistent data timing and quality. The standardized visual, linguistic, and action data are respectively input into a convolutional neural network, a word embedding model, and a fully connected network to generate a visual feature vector sequence, a linguistic feature vector sequence with added positional encoding, and an action feature vector sequence. These are then concatenated in timestamp order to form an initial feature sequence, ensuring consistent temporal correlation among the three modalities.

[0226] In the initial feature sequence, visual, linguistic, and action feature components are separated and processed separately. Visual features are extracted through shallow convolutional layers of the feature pyramid network to obtain detailed features of the user image (such as facial details and document clarity), and through deep convolutional layers to extract semantic features (such as document type and user facial expression). The two are then fused to obtain a multi-scale visual feature hierarchy. Linguistic features are progressively formed into phrase-level, sentence-level, and paragraph-level multi-scale linguistic feature hierarchies through an aggregation module, reflecting the multi-layered semantic logic of user questions or requests. Action features are extracted through small time window convolutions to obtain detailed features (such as the fineness of mouse trajectory and operation response speed), and through large time window convolutions to extract trend features (such as the smoothness of interaction behavior and abnormal pause patterns). The three multi-scale feature hierarchies are finally combined to form a multi-scale feature pyramid set.

[0227] The visual, linguistic, and action multi-scale feature levels in the multi-scale feature pyramid are projected onto a unified semantic space through a constructed cross-modal mapping function. In this space, cosine similarity is calculated between visual and linguistic features, visual and action features, and linguistic and action features to quantify the degree of association between user information across different modalities, such as the consistency between facial expressions and vocal emotions, and semantic content and operational actions. The system dynamically adjusts the parameters of the cross-modal mapping function to ensure the alignment consistency of multi-modal features, outputting a cross-modal aligned multi-scale feature sequence.

[0228] Multi-scale aligned feature sequences are divided into local windows, and local query, key, and value vectors are calculated window by window. Local attention weights are calculated using scaling dot products and normalized exponential functions, guiding the weighted aggregation of local value vectors to form a local attention output sequence. After the sequence is concatenated by time, global query, key, and value vectors are generated. Global scaling dot products are calculated to generate global attention weights, which guide the weighted aggregation of global value vectors to form long-distance dependency features. This captures potential contextual dependencies with large time spans in multiple rounds of user interaction, such as cross-scenario information associations in account opening identity verification and risk control investigations.

[0229] Long-range dependent features are separated hierarchically. For each level, a multi-head attention mechanism is used to fuse multimodal features, outputting local fused features. For local fused features of adjacent levels, smaller-scale features are upsampled to adjust resolution to align with larger-scale features, and larger-scale features are downsampled to adjust resolution to align with smaller-scale features, resulting in upsampled and downsampled features respectively. The upsampled features and larger-scale local fused features are input into the first gating fusion unit to determine the first weight coefficients and fuse them to generate the first cross-layer fused feature. The downsampled features and smaller-scale local fused features are input into the second gating fusion unit to determine the second weight coefficients and fuse them to generate the second cross-layer fused feature. The two are further fused to obtain comprehensive multi-scale features containing rich modal information and scale context interaction features.

[0230] The system separates visual, linguistic, and action feature components from comprehensive multi-scale features, inputs them into a gating unit, calculates dynamic weight coefficients for the three types of features, and dynamically adjusts the weights based on the current interaction context and user state to generate weighted visual, linguistic, and action features. These three features are then fused to generate the final fused feature, which serves as the input to the financial customer service task decision network.

[0231] The task decision-making network assesses risk levels, determines customer identity verification results, and generates service recommendations based on the final fusion characteristics. In remote account opening, the system outputs an automated decision result indicating whether customer identity verification is approved or rejected. In customer risk assessment, the system outputs the current interaction risk level of the customer in real time and provides visual prompts based on historical multimodal data (visual, verbal, and behavioral). In intelligent financial service consultation, the system utilizes long-distance dependency and dynamically weighted multimodal fusion results to generate natural language response suggestions and drive real-time feedback from the virtual customer service interface.

[0232] In the field of robot control, a multimodal robot autonomous perception and decision-making control system is constructed to enable robots to perceive, understand, and make dynamic task decisions based on multimodal inputs such as vision, language, and motion in complex task environments.

[0233] First, the robot captures a continuous stream of images from its workspace using image acquisition devices mounted on its body, acquiring raw visual sequence data. The visual sequence undergoes frame rate normalization to ensure a uniform acquisition frame rate and eliminate sampling differences across different acquisition scenarios. Subsequently, size normalization is performed to adjust the images to a fixed resolution to suit subsequent processing modules. Raw language / text sequence data is acquired by the robot's onboard voice acquisition device, either from live voice commands or text commands transmitted via a network interface. This data undergoes encoding format standardization to ensure input consistency, followed by word segmentation and part-of-speech tagging to transform the language input into structured semantic units. Raw motion sensor sequence data is acquired through the robot's built-in inertial measurement unit, joint encoders, or force sensors, including information on the robotic arm's motion state, gripping pressure changes, and posture adjustments. This data undergoes sampling rate synchronization, noise filtering, and numerical normalization to improve data accuracy and temporal consistency. The standardized visual, language / text, and motion sensor sequences are then converted into visual, language, and motion feature vector sequences, respectively, using convolutional neural networks, word embedding models, and fully connected networks. These sequences are then concatenated in timestamp order to form an initial feature sequence, ensuring data alignment in the temporal dimension.

[0234] The robot control system separates visual feature components, linguistic feature components, and motion feature components from the initial feature sequence. Visual features are extracted at the edge texture level through shallow convolutional layers of a feature pyramid network to identify the detailed contours and boundaries of objects in the environment; semantic features are extracted at the object level through deep convolutional layers to identify object categories, attributes, and scene layout. These two are then fused to form a multi-scale visual feature level. Linguistic features are extracted layer by layer at the phrase, sentence, and paragraph levels through an aggregation module to construct a multi-scale linguistic feature level. Motion features are extracted through small time-window convolutions in a hierarchical temporal convolutional network to extract detailed motion features, such as minor joint adjustments and end-effector corrections; and through large time-window convolutions to extract motion trend features, such as path stability and overall movement trends, forming a multi-scale motion feature level. These three multi-scale feature levels are combined to form a multi-scale feature pyramid set.

[0235] The robotic system constructs a cross-modal mapping function for each feature level of the multi-scale feature pyramid set, projecting visual, linguistic, and action-based multi-scale features onto a shared semantic space. Within this space, the system calculates the cosine similarity between visual and linguistic, visual and action, and linguistic and action information to measure the consistency of multimodal information associations. Based on the similarity scores, the system adjusts the parameters of the cross-modal mapping function to improve the alignment of multimodal representations in the space. The updated mapping function outputs a multi-scale aligned feature sequence after cross-modal alignment.

[0236] The robotic system performs local and global attention processing on the aligned multi-scale aligned feature sequence. The sequence is divided into local windows of a preset window length. For each local window, a local query vector, key vector, and value vector are calculated. The local scaling dot product is obtained, and local attention weights are generated using a normalized exponential function. These weighted local value vectors are then aggregated to generate the local attention output sequence. This sequence is concatenated along the temporal dimension to generate global query vectors, key vectors, and value vectors. Global scaling dot products and global attention weights are calculated to weight the global value vectors, ultimately outputting long-distance dependency features. This captures the robot's long-range contextual dependencies across time during the task flow, such as the association of cross-step voice commands and object tracking under continuous image scene changes.

[0237] The robot system further separates long-distance dependent features hierarchically and fuses multimodal features within each level through a multi-head attention mechanism to form local fusion features. For each pair of adjacent level local fusion features, the smaller-scale features are upsampled to adjust their resolution and align with the larger-scale features, while the larger-scale features are downsampled to adjust their resolution and align with the smaller-scale features. These are then input into the first and second gated fusion units, respectively, to dynamically determine the first and second weighting coefficients. Based on these weighting coefficients, the upsampled and downsampled features are weighted and fused with the adjacent scale local fusion features to generate the first and second cross-level fusion features. The two are finally fused to obtain a comprehensive multi-scale feature, ensuring the robot's information interaction and consistent expression across different perception scales.

[0238] The robot control system separates visual, linguistic, and action feature components from comprehensive multi-scale features and inputs them into a gating unit, where the weight coefficients of each modality are dynamically calculated. These weight coefficients reflect the importance of each modality in the current robot task; for example, visual features have a higher weight during precise grasping, while linguistic features have a higher weight during task confirmation. The adjusted weights guide the generation of weighted visual, linguistic, and action features. The three types of weighted features are then fused to form the final fused feature, which serves as the input to the task decision network.

[0239] The task decision network utilizes the final fusion features to make control decisions and outputs robot control commands in real time, such as adjusting the robotic arm path planning, adjusting the gripping force of the end effector, correcting the direction of visual navigation, and providing conversational task confirmation feedback. The robot can accurately execute multi-round human-robot interaction commands in dynamic and complex environments, autonomously perceiving and adapting to environmental changes. For example, service robots in financial service halls can simultaneously perceive customer body movements, voice inquiries, and environmental visual changes; or robots in manufacturing workshops can simultaneously identify workpiece positions, worker voice commands, and their own movement status.

[0240] This embodiment dynamically fuses visual, linguistic, and action modal information from multi-scale features, enabling the processing results to automatically adjust modal importance and fully utilize modal complementarity for different task contexts. This avoids information redundancy and modal conflicts, enhancing the expressive power of the fused features. The introduction of weighted operations and dynamic gating mechanisms allows the system to robustly generate high-quality, multi-dimensional final fused features even in scenarios with varying visual saliency, fluctuating linguistic context complexity, and unstable dynamic action data. The final fused features provide a consistent, rich, and dynamically optimized input foundation for subsequent task decision-making and reasoning, significantly improving the performance of multimodal processing tasks in complex environments. This includes enhanced adaptability and robustness in applications such as robot fine manipulation, human-computer interaction, semantic understanding, and behavior prediction.

[0241] In one embodiment, a multimodal sequence data processing apparatus is provided, which corresponds one-to-one with the multimodal sequence data processing methods described in the above embodiments. (Refer to...) Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the multimodal sequence data processing device of the present invention. The modules include a multimodal input encoding module 10, a multi-scale feature extraction module 20, a cross-modal alignment module 30, a local and global attention module 40, a cross-layer interactive fusion module 50, a dynamic multimodal fusion module 60, and a task decision module 70. Detailed descriptions of each functional module are as follows:

[0242] The multimodal input encoding module 10 is used to acquire original visual sequence data, original language text sequence data, and original motion sensor sequence data, and generate an initial feature sequence based on the original visual sequence data, original language text sequence data, and original motion sensor sequence data.

[0243] The multi-scale feature extraction module 20 is used to extract multi-scale feature levels from the initial feature sequence and combine the multi-scale feature levels into a multi-scale feature pyramid set.

[0244] The cross-modal alignment module 30 is used to perform cross-modal feature alignment for each feature level of the multi-scale feature pyramid set, generating a multi-scale aligned feature sequence;

[0245] Local and global attention module 40 is used to perform local attention processing and global attention processing on the multi-scale aligned feature sequence to generate long-distance dependent features;

[0246] The cross-layer interaction fusion module 50 is used to perform cross-layer information interaction on the long-distance dependent features to generate comprehensive multi-scale features.

[0247] The dynamic multimodal fusion module 60 is used to dynamically fuse multimodal information based on the comprehensive multi-scale features to generate the final fused features;

[0248] The task decision module 70 is used to input the final fused features into the task decision network to obtain the target task result.

[0249] In one embodiment, the multimodal input encoding module 10 is specifically used for:

[0250] A continuous image stream of the target scene is captured by an image acquisition device to generate raw visual sequence data. The raw visual sequence data is then subjected to frame rate normalization to generate a visual sequence. Finally, the visual sequence is subjected to size normalization to generate a standardized visual sequence.

[0251] The system collects speech signals through a speech recognition device and converts them into text, or receives a text input stream to generate raw data of a language text sequence. The raw data of the language text sequence is then processed to unify the encoding format to generate a language text sequence. Finally, the language text sequence is processed to perform word segmentation and part-of-speech tagging to generate a word segmentation and tagging sequence.

[0252] Physical motion signals are acquired by an inertial measurement unit, force sensor, or joint encoder to generate raw data of motion sensor sequence. The raw data of motion sensor sequence is then processed by sampling rate synchronization to generate motion sensor sequence. The motion sensor sequence is then processed by noise filtering and numerical standardization to generate standardized motion sequence.

[0253] The standardized visual sequence is processed by a convolutional neural network to generate a sequence of visual feature vectors.

[0254] The word segmentation and annotation sequence is processed by a word embedding model to generate a word embedding vector sequence. Position encoding information is added to the word embedding vector sequence to generate a language feature vector sequence.

[0255] The standardized action sequence is processed by a fully connected network to generate an action feature vector sequence;

[0256] The visual feature vector sequence, language feature vector sequence, and action feature vector sequence are concatenated in timestamp order to generate an initial feature sequence.

[0257] In one embodiment, the multi-scale feature extraction module 20 is specifically used for:

[0258] Separate the visual feature component sequence, the language feature component sequence, and the action feature component sequence from the initial feature sequence;

[0259] Edge texture feature levels of the visual feature component sequence are extracted through shallow convolutional layers of the feature pyramid network;

[0260] The object semantic feature level of the visual feature component sequence is extracted through the deep convolutional layers of the feature pyramid network;

[0261] By fusing the edge texture feature level with the object semantic feature level, a visual multi-scale feature level is generated;

[0262] The language feature component sequence is processed by the aggregation module to generate a phrase-level language feature hierarchy.

[0263] The aggregation module aggregates the phrase-level language feature layers to generate a sentence-level language feature layer.

[0264] The aggregation module aggregates the sentence-level language feature levels to generate paragraph-level language feature levels.

[0265] By combining the phrase-level language feature layer, the sentence-level language feature layer, and the paragraph-level language feature layer, a multi-scale language feature layer is generated;

[0266] The action detail feature level of the action feature component sequence is extracted by small time window convolution of a hierarchical temporal convolutional network;

[0267] The action trend feature hierarchy of the action feature component sequence is extracted by large temporal window convolution of a hierarchical temporal convolutional network.

[0268] The action detail feature level and the action trend feature level are combined to generate a multi-scale action feature level;

[0269] The visual multi-scale feature hierarchy, the language multi-scale feature hierarchy, and the action multi-scale feature hierarchy are combined to generate a multi-scale feature pyramid set.

[0270] In one embodiment, the cross-modal alignment module 30 is specifically used for:

[0271] Construct cross-modal mapping functions for the visual multi-scale feature hierarchy, language multi-scale feature hierarchy, and action multi-scale feature hierarchy of the multi-scale feature pyramid set;

[0272] The cross-modal mapping function projects visual multi-scale feature layers, linguistic multi-scale feature layers, and action multi-scale feature layers into a unified semantic space.

[0273] Within the unified semantic space, determine the first cosine similarity between the visual multi-scale feature level and the linguistic multi-scale feature level, the second cosine similarity between the visual multi-scale feature level and the action multi-scale feature level, and the third cosine similarity between the linguistic multi-scale feature level and the action multi-scale feature level.

[0274] The parameters of the cross-modal mapping function are adjusted based on the first cosine similarity, the second cosine similarity, and the third cosine similarity.

[0275] The visual multi-scale feature layers, language multi-scale feature layers, and action multi-scale feature layers are projected by the adjusted cross-modal mapping function to generate a cross-modal aligned multi-scale feature sequence.

[0276] In one embodiment, the local and global attention module 40 is specifically used for:

[0277] The multi-scale aligned feature sequence is divided into multiple local windows according to a preset window length;

[0278] Within each local window, a local query vector, a local key vector, and a local value vector are generated for the multi-scale aligned feature sequence.

[0279] Determine the local scaled dot product of the local query vector and the local key vector within each local window;

[0280] The local scaling dot product result is applied to a normalized exponential function within each local window to generate local attention weights;

[0281] The local value vectors within the local window are weighted and aggregated according to the local attention weights to generate a local attention output sequence.

[0282] The local attention output sequences are concatenated in the time dimension to form a local window output sequence;

[0283] Generate a global query vector, a global key vector, and a global value vector for the output sequence of the local window;

[0284] Determine the result of the global scaled dot product of the global query vector and the global key vector;

[0285] Apply a normalized exponential function to the global scaling dot product result to generate global attention weights;

[0286] The global value vector is weighted according to the global attention weights to generate long-distance dependency features.

[0287] In one embodiment, the cross-layer interaction fusion module 50 is specifically used for:

[0288] The long-distance dependency features are separated into feature sequences of multiple feature levels according to feature levels;

[0289] For the feature sequence of each feature level, a multi-head attention mechanism is used to fuse the multimodal features within the feature level to generate the local fused features of the feature level.

[0290] For each pair of local fusion features at adjacent feature levels, the resolution of the smaller-scale feature is adjusted to the resolution of the adjacent larger-scale feature through an upsampling operation to obtain the upsampled feature, and the resolution of the larger-scale feature is adjusted to the resolution of the adjacent smaller-scale feature through a downsampling operation to obtain the downsampled feature.

[0291] The upsampled features and adjacent larger-scale local fusion features are input into the first gated fusion unit, and a first weight coefficient is determined in the first gated fusion unit;

[0292] Based on the first weighting coefficient, the upsampled feature is fused with the adjacent larger-scale local fusion feature to generate the first cross-layer fusion feature;

[0293] The downsampled features and adjacent smaller-scale local fusion features are input into the second gated fusion unit, and a second weighting coefficient is determined in the second gated fusion unit;

[0294] The second cross-layer fusion feature is generated by fusing the downsampled feature with the adjacent smaller-scale local fusion feature according to the second weighting coefficient.

[0295] The first cross-layer fusion feature and the second cross-layer fusion feature are fused to generate a comprehensive multi-scale feature.

[0296] In one embodiment, the dynamic multimodal fusion module 60 is specifically used for:

[0297] Visual feature components, language feature components, and action feature components are separated from the integrated multi-scale features;

[0298] The visual feature components, language feature components, and action feature components are input into a gating unit, and the dynamic weighting coefficients of the visual feature components, language feature components, and action feature components are determined in the gating unit.

[0299] The visual feature components are weighted according to the dynamic weight coefficients of the visual feature components to generate weighted visual features;

[0300] The language feature components are weighted according to the dynamic weight coefficients of the language feature components to generate weighted language features;

[0301] The motion feature components are weighted according to the dynamic weight coefficients of the motion feature components to generate weighted motion features;

[0302] The weighted visual features, weighted language features, and weighted action features are fused to generate the final fused features.

[0303] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used for communication with external user terminals via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a multimodal sequence data processing method on the server side.

[0304] In one embodiment, a computer device is provided, which may be a user terminal, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements the functions or steps of a multimodal sequence data processing method on the user side.

[0305] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:

[0306] Acquire raw visual sequence data, raw language text sequence data, and raw motion sensor sequence data, and generate an initial feature sequence based on the raw visual sequence data, raw language text sequence data, and raw motion sensor sequence data;

[0307] Multi-scale feature levels are extracted from the initial feature sequence, and the multi-scale feature levels are combined into a multi-scale feature pyramid set.

[0308] Cross-modal feature alignment is performed for each feature level of the multi-scale feature pyramid set to generate a multi-scale aligned feature sequence;

[0309] Local attention and global attention processing are performed on the multi-scale aligned feature sequence to generate long-distance dependent features;

[0310] Cross-layer information interaction is performed on the long-distance dependent features to generate comprehensive multi-scale features;

[0311] Based on the comprehensive multi-scale features, multi-modal information is dynamically fused to generate the final fused features;

[0312] The final fused features are input into the task decision network to obtain the target task result.

[0313] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0314] Acquire raw visual sequence data, raw language text sequence data, and raw motion sensor sequence data, and generate an initial feature sequence based on the raw visual sequence data, raw language text sequence data, and raw motion sensor sequence data;

[0315] Multi-scale feature levels are extracted from the initial feature sequence, and the multi-scale feature levels are combined into a multi-scale feature pyramid set.

[0316] Cross-modal feature alignment is performed for each feature level of the multi-scale feature pyramid set to generate a multi-scale aligned feature sequence;

[0317] Local attention and global attention processing are performed on the multi-scale aligned feature sequence to generate long-distance dependent features;

[0318] Cross-layer information interaction is performed on the long-distance dependent features to generate comprehensive multi-scale features;

[0319] Based on the comprehensive multi-scale features, multi-modal information is dynamically fused to generate the final fused features;

[0320] The final fused features are input into the task decision network to obtain the target task result.

[0321] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0322] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0323] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0324] It should be noted that if any software tools or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The embodiments described above are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for processing multimodal sequence data, characterized in that, Includes the following steps: Acquire raw visual sequence data, raw language text sequence data, and raw motion sensor sequence data, and generate an initial feature sequence based on the raw visual sequence data, raw language text sequence data, and raw motion sensor sequence data; Multi-scale feature levels are extracted from the initial feature sequence, and the multi-scale feature levels are combined into a multi-scale feature pyramid set. Cross-modal feature alignment is performed for each feature level of the multi-scale feature pyramid set to generate a multi-scale aligned feature sequence; Local attention and global attention processing are performed on the multi-scale aligned feature sequence to generate long-distance dependent features; Cross-layer information interaction is performed on the long-distance dependent features to generate comprehensive multi-scale features; Based on the comprehensive multi-scale features, multi-modal information is dynamically fused to generate the final fused features; The final fused features are input into the task decision network to obtain the target task result.

2. The multimodal sequence data processing method as described in claim 1, characterized in that, Acquire raw visual sequence data, raw language text sequence data, and raw motion sensor sequence data, and generate an initial feature sequence based on the raw visual sequence data, raw language text sequence data, and raw motion sensor sequence data, including: A continuous image stream of the target scene is captured by an image acquisition device to generate raw visual sequence data. The raw visual sequence data is then subjected to frame rate normalization to generate a visual sequence. Finally, the visual sequence is subjected to size normalization to generate a standardized visual sequence. The system collects speech signals through a speech recognition device and converts them into text, or receives a text input stream to generate raw data of a language text sequence. The raw data of the language text sequence is then processed to unify the encoding format to generate a language text sequence. Finally, the language text sequence is processed to perform word segmentation and part-of-speech tagging to generate a word segmentation and tagging sequence. Physical motion signals are acquired by an inertial measurement unit, force sensor, or joint encoder to generate raw data of motion sensor sequence. The raw data of motion sensor sequence is then processed by sampling rate synchronization to generate motion sensor sequence. The motion sensor sequence is then processed by noise filtering and numerical standardization to generate standardized motion sequence. The standardized visual sequence is processed by a convolutional neural network to generate a sequence of visual feature vectors. The word segmentation and annotation sequence is processed by a word embedding model to generate a word embedding vector sequence. Position encoding information is added to the word embedding vector sequence to generate a language feature vector sequence. The standardized action sequence is processed by a fully connected network to generate an action feature vector sequence; The visual feature vector sequence, language feature vector sequence, and action feature vector sequence are concatenated in timestamp order to generate an initial feature sequence.

3. The multimodal sequence data processing method as described in claim 1, characterized in that, Extracting multi-scale feature levels from the initial feature sequence and combining the multi-scale feature levels into a multi-scale feature pyramid set, including: Separate the visual feature component sequence, the language feature component sequence, and the action feature component sequence from the initial feature sequence; Edge texture feature levels of the visual feature component sequence are extracted through shallow convolutional layers of the feature pyramid network; The object semantic feature level of the visual feature component sequence is extracted through the deep convolutional layers of the feature pyramid network; By fusing the edge texture feature level with the object semantic feature level, a visual multi-scale feature level is generated; The language feature component sequence is processed by the aggregation module to generate a phrase-level language feature hierarchy. The aggregation module aggregates the phrase-level language feature layers to generate a sentence-level language feature layer. The aggregation module aggregates the sentence-level language feature levels to generate paragraph-level language feature levels. By combining the phrase-level language feature layer, the sentence-level language feature layer, and the paragraph-level language feature layer, a multi-scale language feature layer is generated; The action detail feature level of the action feature component sequence is extracted by small time window convolution of a hierarchical temporal convolutional network; The action trend feature hierarchy of the action feature component sequence is extracted by large temporal window convolution of a hierarchical temporal convolutional network. The action detail feature level and the action trend feature level are combined to generate a multi-scale action feature level; The visual multi-scale feature hierarchy, the language multi-scale feature hierarchy, and the action multi-scale feature hierarchy are combined to generate a multi-scale feature pyramid set.

4. The multimodal sequence data processing method as described in claim 1, characterized in that, Perform cross-modal feature alignment for each feature level of the multi-scale feature pyramid set to generate a multi-scale aligned feature sequence, including: Construct cross-modal mapping functions for the visual multi-scale feature hierarchy, language multi-scale feature hierarchy, and action multi-scale feature hierarchy of the multi-scale feature pyramid set; The cross-modal mapping function projects visual multi-scale feature layers, linguistic multi-scale feature layers, and action multi-scale feature layers into a unified semantic space. Within the unified semantic space, determine the first cosine similarity between the visual multi-scale feature level and the linguistic multi-scale feature level, the second cosine similarity between the visual multi-scale feature level and the action multi-scale feature level, and the third cosine similarity between the linguistic multi-scale feature level and the action multi-scale feature level. The parameters of the cross-modal mapping function are adjusted based on the first cosine similarity, the second cosine similarity, and the third cosine similarity. The visual multi-scale feature layers, language multi-scale feature layers, and action multi-scale feature layers are projected by the adjusted cross-modal mapping function to generate a cross-modal aligned multi-scale feature sequence.

5. The multimodal sequence data processing method as described in claim 1, characterized in that, Local attention processing and global attention processing are performed on the multi-scale aligned feature sequence to generate long-distance dependent features, including: The multi-scale aligned feature sequence is divided into multiple local windows according to a preset window length; Within each local window, a local query vector, a local key vector, and a local value vector are generated for the multi-scale aligned feature sequence. Determine the local scaled dot product of the local query vector and the local key vector within each local window; The local scaling dot product result is applied to a normalized exponential function within each local window to generate local attention weights; The local value vectors within the local window are weighted and aggregated according to the local attention weights to generate a local attention output sequence. The local attention output sequences are concatenated in the time dimension to form a local window output sequence; Generate a global query vector, a global key vector, and a global value vector for the output sequence of the local window; Determine the result of the global scaled dot product of the global query vector and the global key vector; Apply a normalized exponential function to the global scaling dot product result to generate global attention weights; The global value vector is weighted according to the global attention weights to generate long-distance dependency features.

6. The multimodal sequence data processing method as described in claim 1, characterized in that, Cross-layer information interaction is performed on the long-distance dependent features to generate comprehensive multi-scale features, including: The long-distance dependency features are separated into feature sequences of multiple feature levels according to feature levels; For the feature sequence of each feature level, a multi-head attention mechanism is used to fuse the multimodal features within the feature level to generate the local fused features of the feature level. For each pair of local fusion features at adjacent feature levels, the resolution of the smaller-scale feature is adjusted to the resolution of the adjacent larger-scale feature through an upsampling operation to obtain the upsampled feature, and the resolution of the larger-scale feature is adjusted to the resolution of the adjacent smaller-scale feature through a downsampling operation to obtain the downsampled feature. The upsampled features and adjacent larger-scale local fusion features are input into the first gated fusion unit, and a first weight coefficient is determined in the first gated fusion unit; Based on the first weighting coefficient, the upsampled feature is fused with the adjacent larger-scale local fusion feature to generate the first cross-layer fusion feature; The downsampled features and adjacent smaller-scale local fusion features are input into the second gated fusion unit, and a second weighting coefficient is determined in the second gated fusion unit; The second cross-layer fusion feature is generated by fusing the downsampled feature with the adjacent smaller-scale local fusion feature according to the second weighting coefficient. The first cross-layer fusion feature and the second cross-layer fusion feature are fused to generate a comprehensive multi-scale feature.

7. The multimodal sequence data processing method as described in claim 1, characterized in that, Based on the aforementioned comprehensive multi-scale features, multi-modal information is dynamically fused to generate the final fused features, including: Visual feature components, language feature components, and action feature components are separated from the integrated multi-scale features; The visual feature components, language feature components, and action feature components are input into a gating unit, and the dynamic weighting coefficients of the visual feature components, language feature components, and action feature components are determined in the gating unit. The visual feature components are weighted according to the dynamic weight coefficients of the visual feature components to generate weighted visual features; The language feature components are weighted according to the dynamic weight coefficients of the language feature components to generate weighted language features; The motion feature components are weighted according to the dynamic weight coefficients of the motion feature components to generate weighted motion features; The weighted visual features, weighted language features, and weighted action features are fused to generate the final fused features.

8. A multimodal sequence data processing device, characterized in that, The multimodal sequence data processing device includes: A multimodal input encoding module is used to acquire raw visual sequence data, raw language text sequence data, and raw motion sensor sequence data, and generate an initial feature sequence based on the raw visual sequence data, raw language text sequence data, and raw motion sensor sequence data. A multi-scale feature extraction module is used to extract multi-scale feature levels from the initial feature sequence and combine the multi-scale feature levels into a multi-scale feature pyramid set. The cross-modal alignment module is used to perform cross-modal feature alignment for each feature level of the multi-scale feature pyramid set, generating a multi-scale aligned feature sequence; The local and global attention module is used to perform local and global attention processing on the multi-scale aligned feature sequence to generate long-distance dependent features. The cross-layer interaction fusion module is used to perform cross-layer information interaction on the long-distance dependent features to generate comprehensive multi-scale features; The dynamic multimodal fusion module is used to dynamically fuse multimodal information based on the comprehensive multi-scale features to generate the final fused features; The task decision module is used to input the final fused features into the task decision network to obtain the target task result.

9. A computer device, characterized in that, The computer device includes a memory, a processor, and a multimodal sequence data processing program stored in the memory and executable on the processor, wherein the multimodal sequence data processing program, when executed by the processor, implements the steps of the multimodal sequence data processing method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The storage medium stores a multimodal sequence data processing program, which, when executed by a processor, implements the steps of the multimodal sequence data processing method as described in any one of claims 1-7.

Citation Information

Cited By

  • Interaction method and system based on multi-modal data

    CN121255027A

  • Network request risk detection method, system and server

    CN121309233A

  • Network request risk detection method, system and server

    CN121309233B

  • Pathogenesis record generation method and system based on mobile terminal and multi-modal data

    CN121483473A

  • Method and system for generating a medical history based on a mobile terminal and multi-modal data

    CN121483473B