Task instruction generation method and device based on cross-modal fusion, equipment and medium

By decoding and denoising the input video, identifying keyframes and extracting spatial and temporal features, and combining text and action features for cross-modal fusion, the problem of insufficient accuracy in generating task instructions in dynamic scenes by VLM is solved, achieving higher accuracy and adaptability.

CN120932052AActive Publication Date: 2025-11-11PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202511185261.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-22
Publication Date
2025-11-11
Estimated Expiration
2045-08-22

AI Technical Summary

Technical Problem

Existing visual language models (VLMs) rely on static image input, making it difficult to effectively extract temporal features and multimodal information from videos, resulting in insufficient accuracy and responsiveness in generating task instructions in dynamic scenes.

Method used

The input video is decoded and denoised to generate a frame sequence. Multiple key frames are identified, and the spatial and temporal features of the key frames are obtained. Cross-modal fusion is performed by combining text semantic features and action features to generate fused features and a perception vector, and finally, task instructions are generated.

Benefits of technology

By effectively utilizing the dynamic temporal information and multi-source perception information of video, the expressive power of perception vectors is improved, and the accuracy and context adaptability of generated task instructions are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932052A_ABST
    Figure CN120932052A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, can be applied to business scenes such as financial science and technology and medical health, and discloses a task instruction generation method, device and equipment based on cross-modal fusion, and a medium, and the method comprises the steps: carrying out the decoding and noise reduction of an input video, generating a frame sequence, and recognizing a plurality of key frames based on the inter-frame similarity; extracting spatial features of the key frames to form a sequence, and generating video spatial-temporal features in combination with time features; performing semantic preprocessing on the input text to obtain text semantic features, and acquiring motion sensor signals to obtain motion features; fusing the video spatio-temporal features, the text semantic features and the action features to generate fused features; and generating a perception vector based on the fusion feature and outputting a task instruction. According to the method, multi-modal fusion is realized through key frame extraction and space-time fusion mechanisms in combination with text semantic features and action features, and the perception expression ability and the task instruction generation accuracy are improved by using time sequence information and multi-source perception input of the video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device and storage medium for generating task instructions based on cross-modal fusion. Background Technology

[0002] In traditional Vision-Language Models (VLMs), visual information mostly comes from static images, and the model achieves cross-modal understanding by aligning images with text. However, static images can only reflect scene information at a specific point in time and cannot present the development process of continuous events, lacking the ability to model the temporal dimension. This limitation causes the model to perform poorly when dealing with tasks involving dynamic changes and temporal relationships, especially in applications that require understanding semantics such as behavioral trajectories, action causality, and event evolution.

[0003] In the fintech sector, scenarios such as intelligent customer service, risk control analysis, and financial regulatory assistance are increasingly incorporating video-based human-computer interaction methods. For example, remote identity verification, video interviews, and financial service operation guidance not only involve voice and text but also crucial dynamic information such as user actions, facial expressions, and operational sequences. Visual-language models relying solely on image and text processing struggle to capture abnormal patterns or fraudulent tendencies in user behavior, and cannot accurately model user intent and behavioral logic, easily leading to misjudgments or recognition delays.

[0004] In the healthcare field, typical applications such as remote assisted diagnosis and treatment, surgical guidance, and rehabilitation training assessment exhibit clear temporal evolution characteristics in patients' movements, gestures, and body postures. Existing VLM models cannot effectively model these dynamic processes closely related to medical intentions, making it difficult to support accurate assessment and feedback of patient status, thus limiting the diagnostic accuracy and interaction efficiency of intelligent assistance systems.

[0005] Furthermore, in general visual understanding scenarios such as intelligent security, autonomous driving, and robotic collaboration, the model's insufficient ability to perceive dynamic environments has become a key factor restricting the level of system intelligence. Due to the lack of effective identification of key frames and modeling of temporal relationships in video sequences, traditional VLM models face problems such as incomplete perceptual information and inconsistent semantic understanding during multimodal fusion, making them unable to accurately respond to complex scenes. Summary of the Invention

[0006] The main objective of this invention is to provide a method, apparatus, device, and storage medium for generating task instructions based on cross-modal fusion. This invention aims to solve the technical problem that existing visual language models generally rely on static image input, making it difficult to effectively extract temporal features from videos and deeply fuse them with multimodal information, resulting in insufficient accuracy and responsiveness of task instructions generated in dynamic scenes.

[0007] To achieve the above objectives, the present invention provides a task instruction generation method based on cross-modal fusion, comprising:

[0008] Decoding and noise reduction are performed on the input video to generate a frame sequence, and multiple key frames are identified based on the similarity relationship between adjacent frames in the frame sequence.

[0009] The spatial features of the multiple key frames are obtained to form a spatial feature sequence. Temporal features are extracted based on the spatial feature sequence, and the spatial feature sequence and the temporal features are fused to form video spatiotemporal features.

[0010] Semantic preprocessing is performed on the input text to generate text semantic features, and motion sensor signals are collected to obtain motion features;

[0011] Cross-modal fusion is performed on the video spatiotemporal features, the text semantic features, and the action features to generate fused features;

[0012] A perception vector is generated based on the fused features, and task instructions are generated based on the perception vector.

[0013] Furthermore, to achieve the above objectives, the present invention provides a task instruction generation apparatus based on cross-modal fusion, comprising:

[0014] The video preprocessing module is used to perform decoding and noise reduction on the input video to generate a frame sequence, and to identify multiple key frames based on the similarity relationship between adjacent frames in the frame sequence.

[0015] The spatiotemporal feature extraction module is used to obtain the spatial features of the multiple key frames to form a spatial feature sequence, extract temporal features based on the spatial feature sequence, and fuse the spatial feature sequence with the temporal features to form video spatiotemporal features;

[0016] The multimodal feature acquisition module is used to perform semantic preprocessing on the input text to generate text semantic features and to collect motion sensor signals to obtain motion features;

[0017] The cross-modal fusion module is used to perform cross-modal fusion on the video spatiotemporal features, the text semantic features, and the action features to generate fused features;

[0018] The perception decision module is used to generate a perception vector based on the fused features and to generate task instructions based on the perception vector.

[0019] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and a task instruction generation program based on cross-modal fusion stored in the memory and executable on the processor, wherein when the task instruction generation program based on cross-modal fusion is executed by the processor, it implements the steps of the task instruction generation method based on cross-modal fusion as described above.

[0020] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing a task instruction generation program based on cross-modal fusion, wherein when the task instruction generation program based on cross-modal fusion is executed by a processor, it implements the steps of the task instruction generation method based on cross-modal fusion as described above.

[0021] Beneficial Effects: This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as fintech and healthcare. It discloses a method, apparatus, device, and medium for generating task instructions based on cross-modal fusion, comprising: decoding and denoising an input video to generate a frame sequence, and identifying multiple keyframes based on the similarity relationship between adjacent frames in the frame sequence; acquiring spatial features of the multiple keyframes to form a spatial feature sequence, extracting temporal features based on the spatial feature sequence, and fusing the spatial feature sequence and temporal features to generate video spatiotemporal features; performing semantic preprocessing on the input text to generate text semantic features, and simultaneously acquiring motion sensor signals to obtain motion features; performing cross-modal fusion of video spatiotemporal features, text semantic features, and motion features to generate fused features; generating a perception vector based on the fused features, and generating task instructions based on the perception vector. This invention achieves multimodal fusion by extracting video keyframes and fusing spatiotemporal features, combined with text semantic features and motion features, thereby effectively utilizing the dynamic temporal information and multi-source perception information of the video, improving the expressive power of the perception vector, and ultimately improving the accuracy and context adaptability of the generated task instructions. Attached Figure Description

[0022] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings:

[0023] Figure 1 This is a schematic diagram of an application environment for a task instruction generation method based on cross-modal fusion according to an embodiment of the present invention;

[0024] Figure 2 This is a flowchart illustrating an embodiment of the task instruction generation method based on cross-modal fusion of the present invention;

[0025] Figure 3This is a schematic diagram of the functional modules of a preferred embodiment of the task instruction generation device based on cross-modal fusion of the present invention;

[0026] Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention;

[0027] Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0028] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0029] The task instruction generation method based on cross-modal fusion provided in this invention can be applied to, for example... Figure 1 In this application environment, the user terminal communicates with the server via a network. The server can decode and denoise the input video through the user terminal to generate a frame sequence, and identify multiple keyframes based on the similarity relationship between adjacent frames in the frame sequence; acquire the spatial features of multiple keyframes to form a spatial feature sequence, extract temporal features based on the spatial feature sequence, and fuse the spatial feature sequence and temporal features to generate video spatiotemporal features; perform semantic preprocessing on the input text to generate text semantic features, and simultaneously collect motion sensor signals to obtain motion features; perform cross-modal fusion of video spatiotemporal features, text semantic features, and motion features to generate fused features; generate a perception vector based on the fused features, and generate task instructions based on the perception vector. This invention achieves multimodal fusion by combining video keyframe extraction and spatiotemporal feature fusion mechanism with text semantic features and motion features, thereby effectively utilizing the dynamic temporal information and multi-source perception information of the video, improving the expressive power of the perception vector, and ultimately improving the accuracy and context adaptability of the generated task instructions. The user terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster composed of multiple servers. The present invention will now be described in detail through specific embodiments.

[0030] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the task instruction generation method based on cross-modal fusion provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0031] like Figure 2 As shown, the task instruction generation method based on cross-modal fusion proposed in this invention includes the following steps:

[0032] S10, Decoding and noise reduction are performed on the input video to generate a frame sequence, and multiple key frames are identified based on the similarity relationship between adjacent frames in the frame sequence.

[0033] In this embodiment, the input video refers to a data sequence containing consecutive image frames, which can be encoded in formats such as MP4, AVI, and H.264. After the system receives the input video, it first needs to undergo decoding, that is, restoring the compressed video stream to the original image frames. The decoding operation is usually based on multimedia frameworks such as FFmpeg, obtaining the frame sequence by restoring the pixel matrix frame by frame. To reduce noise interference in the image, further noise reduction processing is performed. Noise reduction can use spatial domain filtering (such as median filtering) or temporal domain filtering methods (such as multi-frame fusion). The filtering parameters can be dynamically set according to the image sharpness and the adaptive noise level estimation results. The frame sequence generated by the above processing can provide high-fidelity time-series image data.

[0034] The similarity relationship between adjacent frames is used to assess the degree of change in an image over time. Similarity calculation can be performed based on structural similarity index (SSIM), cosine similarity, or deep feature matching. The system traverses each pair of adjacent frames in the frame sequence, calculates their similarity value, and compares this similarity value with a set threshold to determine whether the current frame has significant changes. To enhance the accuracy of recognition, similarity judgment can be combined with image gradient change or optical flow analysis to further improve temporal perception capabilities.

[0035] After performing similarity analysis on adjacent frames, the system iterates through all frames, marking those with similarity below a threshold or exhibiting significant structural changes as candidate keyframes. Keyframe identification relies not only on the magnitude of change in a single pair of frames but also on integrating local fluctuation trends across multiple frames within a sliding window to ensure a balanced distribution of keyframes over time. The data structures used in this step typically include a frame index, a similarity matrix, and a tag list, forming the base frame set for downstream feature extraction and fusion.

[0036] During decoding, GPU-accelerated video codecs can be used to improve processing throughput, especially suitable for high-definition video input scenarios. In the noise reduction stage, learning-based noise reduction using convolutional neural networks (such as DnCNN) can be chosen, suitable for image optimization in complex environments. For similarity calculation, ResNet or Vision Transformer can be used to extract inter-frame semantic features before similarity measurement, to better adapt to heterogeneous video sources. In the keyframe recognition stage, inter-frame motion estimation results or intra-frame object detection scores can be combined as supplementary features to improve the robustness of the judgment.

[0037] Example Description: In the healthcare field, this can be used for motion sequence analysis of intraoperative endoscopic images. After decoding and denoising the intraoperative video, the system extracts keyframes of the surgical tool movement trajectory, which helps in subsequent motion recognition and risk prediction.

[0038] In the fintech business, it can be applied to the identification and analysis of customer interaction actions in bank hall videos. By filtering key frames of sudden changes in interaction behavior, it can capture abnormal service signals and assist in behavior compliance judgment and risk interception.

[0039] This embodiment generates a frame sequence by performing decoding and noise reduction on the input video, and identifies multiple keyframes based on the similarity relationship between adjacent frames in the frame sequence. This effectively improves the quality of video information expression and the efficiency of feature extraction. Decoding and noise reduction ensure the clarity and stability of the frame data, while the keyframe identification strategy compresses the amount of data processed subsequently and enhances the system's ability to model temporal changes, thus laying a high-quality temporal foundation for subsequent spatial feature extraction and multimodal fusion.

[0040] S20, obtain the spatial features of the multi-frame keyframes to form a spatial feature sequence, extract temporal features based on the spatial feature sequence, and fuse the spatial feature sequence with the temporal features to form video spatiotemporal features;

[0041] In this embodiment, keyframe images typically contain representative scene information, and the frame sequence contains rich structural, textural, and semantic features. Obtaining the spatial features of keyframes refers to feature encoding of the content of each frame image in a two-dimensional plane, often accomplished using a convolutional neural network (CNN) structure. Spatial features express static attributes such as objects, background, boundaries, and textures in an image in the form of high-dimensional vectors or feature maps, exhibiting spatial locality and receptive field structure.

[0042] Each keyframe undergoes encoding processing via a uniformly structured feature extraction network. The extracted spatial features of each frame are then arranged according to their original temporal order, forming a spatial feature sequence. This sequence maintains the consistency of the image's temporal structure and provides a foundation for modeling temporal dependencies. To ensure the integrity of temporal information, the sequence length is not compressed; instead, the order of each frame is preserved in its total order.

[0043] Extracting temporal features from a constructed spatial feature sequence typically employs a network architecture capable of temporal modeling. Long Short-Term Memory (LSTM) networks or Gated Recurrent Units (GRUs) are suitable for modeling long dependencies in temporal data. In higher-order semantic modeling, the Transformer architecture can also be used to process spatial feature sequences and generate sequence-level temporal feature representations. Temporal features are used to reflect dynamic patterns such as inter-frame change patterns, action phase transitions, and state evolution trajectories, and possess cross-frame global modeling capabilities.

[0044] The fusion of spatial and temporal feature sequences is not a simple concatenation; it requires aligning their dimensions, semantic levels, and information distributions. Feature alignment networks can be used to align, resample, or normalize spatial and temporal vectors, ensuring consistency across channels. Fusion methods include channel-level concatenation, weighted summation, and attention-weighted combination. The fused spatiotemporal features serve as a unified encoded representation, preserving intra-frame image structure while reflecting inter-frame evolution, providing integrated semantic support for downstream tasks.

[0045] When extracting spatial features, structures such as ResNet, MobileNet, or Swin Transformer can be used, with the model selection based on the computing power of the application environment. If there are high requirements for deployment on edge devices, a lightweight feature extraction module can be selected. After the spatial feature sequence is constructed, the features of each frame can be uniformly encoded into a fixed-length vector while maintaining the temporal order. In actual systems, the sequence construction is completed through tensor stacking.

[0046] When extracting temporal features, a multi-layer LSTM structure or a bidirectional GRU structure can be used, and the processing order can be forward, backward, or a combination of both. For processing long video sequences, a segmented modeling strategy can be adopted, dividing the spatial feature sequence into time windows, and merging the temporal features of each window after independent modeling. If the system has strong computing resources, a Transformer-based temporal modeling approach can also be used, utilizing a self-attention mechanism to extract global temporal dependencies.

[0047] During the fusion phase, an adaptive attention mechanism can be used to assign dynamic weights to features at different temporal or spatial locations, enabling the fused vector to possess both region selectivity and temporal awareness in its representation. After fusion, a multilayer perceptron can be used to compress or enhance the features to adapt to the input requirements of subsequent task modules.

[0048] Example description: In healthcare business scenarios, it can be applied to auxiliary diagnosis and treatment video analysis to extract spatial lesion features from the imaging equipment output images in key frames, and at the same time model the temporal evolution of the diagnosis and treatment process to identify whether there are abnormal steps in the surgical procedure.

[0049] In fintech businesses, it can be used to understand interactive actions in customer identity verification video streams. It can identify customer holding posture and background consistency through spatial features, and determine whether the sequence of actions conforms to the preset compliance process through temporal features, thereby assisting in customer intent identification and risk control.

[0050] This embodiment acquires spatial features from keyframes and constructs a spatial feature sequence. It then combines this with sequence-level temporal modeling to generate temporal features, fusing the two to form a unified spatiotemporal feature set. This enables a comprehensive depiction of both static information and dynamic changes in video input. Spatial features handle structural recognition and semantic segmentation tasks, while temporal features reflect the continuity of actions and stage changes. The fusion of these two features significantly improves the accuracy of multimodal fusion and action decision-making in downstream tasks, avoiding the scene fragmentation and action ambiguity problems caused by static representations.

[0051] S30: Perform semantic preprocessing on the input text to generate text semantic features, and collect motion sensor signals to obtain motion features;

[0052] In this embodiment, semantic preprocessing refers to the structuring, standardization, and vectorization of text data so that the subsequent perception module can understand the text semantics and fuse it with other modal information. Input text typically originates from user language commands, dialogue content, prompts, etc., and its structure may contain various linguistic phenomena, such as polysemous words, ambiguous sentences, domain-specific terms, and context-dependent expressions. To ensure the accurate transmission of text semantics during the encoding process, semantic preprocessing operations typically include basic language processing steps such as word segmentation, part-of-speech tagging, named entity recognition, and dependency parsing.

[0053] The core of semantic preprocessing is to transform discrete text representations into continuous semantic vector representations with context-aware capabilities. In practice, the preprocessing results are fed into a context modeling network, such as a pre-trained language model based on the Transformer architecture, to capture dependencies between words and semantic representations at the sentence level. The generated text semantic features are fixed-length or variable-length vector structures that can express abstract semantic information at the word, phrase, or sentence level. These features will participate in subsequent multimodal fusion processing, therefore requiring good discriminativeness and composability.

[0054] The acquisition of motion features relies on real-time acquisition of motion sensor signals, which can include time-series data such as acceleration, angular velocity, pressure, displacement, and resistance changes. These sensors are typically deployed at key points of human movement, wearable devices, input terminals, or interactive devices. The data is in a continuous stream structure, exhibiting time correlation and noise interference. After signal acquisition, preprocessing operations in the time or frequency domain are required, including filtering, normalization, resampling, and segmentation, to ensure signal stability and inputability.

[0055] Processing motion sensor signals to obtain motion features typically employs a two-stage approach. The first stage uses a temporal convolutional structure to extract motion patterns within a local time window, including short-term amplitude changes, periodic fluctuations, and adjacent rate of change. The second stage introduces a temporal memory network, such as a gated recurrent unit or a variant of LSTM, to capture behavioral trends, motion patterns, and behavioral intentions across time periods. This two-stage structure can jointly model local and global motion patterns while maintaining the integrity of the temporal structure.

[0056] The acquisition of textual semantic features and action features provides two non-visual modal inputs—language and action—for subsequent multimodal alignment and decision-making. These two modalities differ in their expression, temporal structure, and noise properties; therefore, the design must ensure that they possess an alignment foundation at the abstract level to support the robustness and expressiveness of cross-modal information fusion operations.

[0057] Semantic preprocessing can be accomplished using a language understanding model based on the BERT architecture. The text input is first segmented into sub-word units by a tokenizer, then embedded into an encoder to generate multi-layered semantic feature vectors. For longer instructions or dialogues with context, dialogue optimization models such as DialogBERT can be used to capture the referential and logical relationships between sentences. In practical deployments, simplified pre-trained models such as DistilBERT or ALBERT can be used to reduce resource consumption and improve response speed.

[0058] The acquisition frequency, sampling dimension, and preprocessing strategy of motion sensor signals need to be adapted to the specific hardware device. For a triaxial accelerometer, the data format is a three-channel floating-point sequence, and the sampling frequency can be set to 50Hz or 100Hz. Filtering strategies can include a moving average filter or wavelet denoising algorithm to improve signal stability. Motion data segments can be divided using fixed or sliding window truncation methods, with the window length set according to task latency requirements. In the feature extraction stage, a one-dimensional convolutional network can be used to obtain the motion change trend within a time segment, and then a two-layer GRU structure can be used to model long-term time dependencies.

[0059] In a multi-tasking environment, to address the issue of significant differences in action dimensions across different scenarios, an adaptive dimension mapping module can be introduced to uniformly map action data collected by different devices to the same feature dimension, thereby enhancing the system's adaptability in heterogeneous sensor environments.

[0060] Example description: In healthcare scenarios, this can be applied to the interactive process of intelligent rehabilitation training. Patients describe their feelings through voice and perform limb movements at the same time. The system determines user needs based on semantic analysis and judges whether the rehabilitation movements are standardized and whether the range of motion is up to standard through action recognition, thereby helping to generate personalized training suggestions.

[0061] In fintech businesses, it can be deployed in remote identity authentication systems. Users interact with text and actions according to prompts. The text part is used for semantic understanding and identity matching, while the action part is used for instruction execution compliance verification, such as serialized behaviors like nodding, holding a certificate, and turning around, which improves authentication security and the naturalness of interaction.

[0062] This embodiment generates high-quality semantic vectors by semantic preprocessing the input text and simultaneously collects and models the temporal signals of the motion sensor. It can establish a structured input channel between the two non-visual modalities of language and action, providing clear and structurally complete semantic and action representations for multimodal fusion. This significantly enhances the system's ability to accurately perceive user intent and behavioral state, avoiding problems such as one-sided scene understanding and incomplete behavior recognition caused by relying solely on visual information.

[0063] S40, perform cross-modal fusion on the video spatiotemporal features, the text semantic features, and the action features to generate fused features;

[0064] In this embodiment, cross-modal fusion is a process that connects, compares, and collaboratively models data from different modalities at the feature level. The goal is to construct a unified fusion representation that is highly consistent in expression, semantically complete, and usable for subsequent inference tasks. In this step, the input includes three different modalities: video spatiotemporal features, text semantic features, and action features. These three features have been standardized into structured vector representations during generation, but they still exhibit significant differences in source data attributes, temporal organization, expression granularity, and noise distribution. Therefore, to achieve effective fusion, vector alignment operations need to be performed on the three types of features first to eliminate dimensional and scale differences between modalities.

[0065] Vector alignment typically involves two key actions: first, mapping three types of features to the same vector space through linear transformation or adaptive projection mechanisms; and second, constructing the intermodal relationship structure using cross-attention mechanisms or shared embedding representations. After alignment, the fusion process inputs these aligned features into a fusion network for semantic interaction modeling. The fusion network can employ a multi-head self-attention mechanism to achieve cross-extraction of information, where each attention head is responsible for capturing the coupling relationships and complementary information between different modalities, and finally, the outputs of all attention heads are integrated to form a fused feature vector.

[0066] The generated fusion features not only include spatiotemporal dynamic information from videos, linguistic intent information from texts, and behavioral execution information from actions, but also possess a certain degree of robustness and generalization ability, adapting to diverse task objectives and scene inputs. This fusion processing method requires that each modality feature possess composability, semantic consistency, and contextual relevance to ensure that the final fused representation does not fail or deviate due to heterogeneous inputs.

[0067] In practical implementation, video spatiotemporal features, text semantic features, and action features are each normalized and projected onto the same embedding space through three independent linear transformation modules to ensure consistent vector dimensions. Subsequently, the three types of embedding features are combined in a concatenated or weighted manner to form a fusion input matrix. This matrix is ​​then fed into a cross-modal fusion module containing a multi-head attention layer and a feedforward neural network layer for modeling.

[0068] Each attention head uses independent queries, key-value transformations, and modeling of dependencies between different modalities, such as language-visual alignment between text and video, instruction-execution mapping between text and action, and image-behavior collaboration between video and action. The outputs of multiple attention heads are concatenated and input into a residual connection structure and a layer normalization structure to improve the stability and expressiveness of the features.

[0069] In the process of generating fused features, a modality weight adjustment mechanism can be introduced to dynamically adjust the importance of different modalities according to the current task scenario. For example, in a task that emphasizes behavior recognition, the weight of action features is increased, while in a task that aims to understand scene intent, the proportion of textual semantic features is increased, thereby improving the task adaptability and decision effectiveness of the fused representation.

[0070] Example description: In the field of healthcare, this system can be used in patient rehabilitation training monitoring scenarios. It acquires training videos to generate video spatiotemporal features, identifies patients' verbal descriptions to generate text semantic features, and records motion sensor data to generate motion features. By fusing the generated features, it determines whether the current training state meets the predetermined goals and assists physicians in dynamically adjusting training strategies.

[0071] In fintech business scenarios, this technology can be deployed during remote auditing or behavioral authentication processes to acquire customer operation video and voice interaction commands, while simultaneously recording the trajectory of operational actions. The fused features generated after fusing these three modalities are used to identify the compliance and intent of user behavior. Figure 1 This consistency improves the accuracy of system risk identification and user profile construction.

[0072] This embodiment effectively mitigates issues such as semantic offset, dimensional inconsistency, and structural mismatch among multimodal inputs by unifying and merging video spatiotemporal features, text semantic features, and action features. This enhances the model's comprehensive understanding of multimodal cues in complex task environments. The generation method of fused features strengthens the complementarity and structural correlation of semantic-level information, which helps improve the accuracy and stability of subsequent task vector and instruction generation modules, avoiding the risk of information fragmentation or misleading information.

[0073] S50, generate a perception vector based on the fused features, and generate task instructions based on the perception vector.

[0074] In this embodiment, the fused features include cross-modal integrated information extracted and uniformly encoded from video, text, and motion signals. Its structural representation has been organized into a high-dimensional vector with semantic cooperative relationships. To convert this vector into a structured expression usable for instruction generation, it needs to be further compressed, reconstructed, and semantically integrated using a neural network structure with context modeling capabilities to form a task-aware vector. This awareness vector is an intermediate semantic representation whose function is to accurately express the intent, environmental state, and behavioral context embodied in the current input based on multimodal fusion.

[0075] The process of constructing perceptual vectors can be accomplished using a deep model based on an attention mechanism. The fused features are fed into a perceptual network with global modeling capabilities. This network possesses a multi-head attention mechanism, which can simultaneously model information interactions between different dimensions and feature enhancement at key semantic locations. After attention-weighted processing, the fused features generate a more structurally focused intermediate vector representation, exhibiting stronger decision-orientedness and task adaptability.

[0076] Based on the feature patterns of the perceptual vectors, the system selects an appropriate decision module to generate task instructions. The structure of the task instructions matches the current system task library and typically consists of a set of parameterized instruction units, such as action sequence labels, target recognition instructions, and semantic response templates. The decision module can employ a classifier, sequence generator, or a hybrid control policy network, with different strategies automatically adapting to task requirements. The process of generating task instructions is not merely a simple mapping; rather, it involves selecting the task with the optimal execution probability from multiple task candidates based on the contextual expression and semantic concentration in the perceptual vectors, thus achieving linkage between task reasoning and expression mapping.

[0077] The fused features are first fed into a perceptual network containing a multi-head attention layer and a feedforward network. This network architecture is adapted from the standard Transformer model, introducing an adjustable attention scale coefficient to accommodate the semantic density of different modalities. During the model's forward pass, the perceptual network uses attention heads to capture residual structural information across modalities, outputting a set of attention-weighted features.

[0078] These features, after being processed through residual connections and normalization, are compressed to a fixed dimension via a fully connected network, forming a unified perceptual vector. This vector structure contains multiple sub-dimensions, each representing a certain type of task attribute indicator or intent orientation, used to drive the subsequent task inference model to perform structural mapping.

[0079] To generate specific task instructions, the system predefines multiple task classifier or generator modules, each with an independent input interface and parameter structure. By analyzing the similarity distribution between the feature structure of the perceptual vector and the task label, the optimal instruction generation module is dynamically selected, and the perceptual vector is input into that module to construct a complete task instruction. In some task environments, positional encoding or historical interaction states can also be appended to the perceptual vector to enhance contextual continuity.

[0080] Example Explanation: In healthcare scenarios, when patients describe their symptoms via voice input and record related behaviors using gestures or body-sensing operations, the system generates fusion features based on video and motion data, then converts them into perceptual vectors. These perceptual vectors drive the instruction module to generate instructions such as "It is recommended to perform a basic gait test" or "It is recommended to conduct an initial screening for skeletal abnormalities," which are used to assist doctors in making judgments or to automatically queue for initial examinations.

[0081] In fintech business scenarios, after the video behavior, language interaction and action trajectory of users when performing remote operations are uniformly modeled, task instructions such as "enter risk control verification process" or "execute fast settlement operation" are generated through perception vectors, thereby realizing system guidance for compliance, anti-fraud or transaction operations and having high-precision decision support capabilities.

[0082] This embodiment utilizes a processing mechanism that fuses features to generate perceptual vectors and then generates task instructions based on these vectors. This mechanism enables efficient conversion from multimodal semantic representations to structured behavioral targets, enhancing the system's responsiveness to complex semantic inputs. The perceptual vectors possess semantic abstraction capabilities, maintaining information consistency and accuracy across different task contexts. The generated task instructions offer advantages such as context awareness, adaptive reasoning, and consistent instruction structure, significantly improving the system's task understanding and execution capabilities in multi-source perceptual environments.

[0083] This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as fintech and healthcare. It discloses a method, apparatus, device, and medium for generating task instructions based on cross-modal fusion, comprising: decoding and denoising an input video to generate a frame sequence, and identifying multiple keyframes based on the similarity relationship between adjacent frames in the frame sequence; acquiring spatial features of the multiple keyframes to form a spatial feature sequence, extracting temporal features based on the spatial feature sequence, and fusing the spatial feature sequence and temporal features to generate video spatiotemporal features; performing semantic preprocessing on the input text to generate text semantic features, and simultaneously acquiring motion sensor signals to obtain motion features; performing cross-modal fusion of video spatiotemporal features, text semantic features, and motion features to generate fused features; generating a perception vector based on the fused features, and generating task instructions based on the perception vector. This invention achieves multimodal fusion by extracting video keyframes and fusing spatiotemporal features, combined with text semantic features and motion features, thereby effectively utilizing the dynamic temporal information and multi-source perception information of the video, improving the expressive power of the perception vector, and ultimately improving the accuracy and context adaptability of the generated task instructions.

[0084] In one embodiment, step S10 above includes:

[0085] S101 uses a high-efficiency video coding standard to perform decoding operations on the input video and generate the original image frame sequence;

[0086] S102, the original image frame sequence is denoised frame by frame by a convolutional autoencoder to generate a denoised frame sequence;

[0087] S103, Based on the dynamic time warping module, determine the similarity value between each frame in the noise reduction frame sequence and the next frame;

[0088] S104, when the similarity value between a certain frame and its next frame is lower than a preset threshold, the certain frame is marked as a candidate keyframe;

[0089] S105, Identify consecutively adjacent candidate keyframe groups in a chronologically ordered sequence of candidate keyframes;

[0090] S106, For each consecutive adjacent candidate keyframe group, only the earliest candidate keyframe in the group is retained;

[0091] S107, all retained candidate keyframes are used as the final multi-frame keyframes.

[0092] In this embodiment, the input video consists of a time series of consecutive frames. Its original encoding format may employ various compression standards. To accurately extract valid keyframe information, a complete image decoding operation must first be performed. The decoding process uses efficient video coding standards, including but not limited to H.265 / HEVC, AV1, or future deep learning-optimized adaptive frame prediction structures. This operation aims to restore the compressed data blocks in the encoded video to a frame-by-frame original image sequence in the image domain, providing a pixel-accurate input foundation for subsequent image analysis and processing.

[0093] The decoded raw image frame sequence contains noise and blur introduced during compression encoding, which may lead to the loss of structural details and texture information, especially during inter-frame prediction and transform quantization. Therefore, image denoising is necessary to enhance spatial structural representation. The image denoising module employs a convolutional autoencoder architecture. Each frame of the image is input into the encoder for feature compression, and then restored pixel by pixel by the decoder. Skip connections and residual enhancement mechanisms are used to maintain the integrity of edge and texture information restoration. This module maintains the temporal order and spatial consistency of image frames during processing, ensuring that the denoised image can be used for temporal comparative analysis.

[0094] Based on the denoised frame sequence, a dynamic time warping module is introduced to calculate inter-frame similarity in order to identify the temporal dynamic features and content differences between each frame. This module extracts multi-scale feature maps of adjacent frames through a bidirectional sliding window structure and uses methods such as cosine similarity, structural similarity (SSIM), or contrastive loss function to measure the degree of consistency of information between frames. Specifically, the vector distance between the current frame and its successor frame in the feature space is standardized to form a similarity matrix for batch judgment and threshold determination.

[0095] A frame is labeled as a candidate keyframe when its similarity to its successor is below a set threshold. This threshold can be pre-trained or dynamically adjusted based on the actual task scenario to ensure that keyframes are labeled only when there are significant scene changes or abrupt changes in motion trends, reducing redundant computation. After all candidate keyframes are arranged in their original chronological order, adjacent dense regions need to be further filtered to remove redundant frames. The system identifies consecutively arranged groups of candidate keyframes, i.e., several frames are consecutively labeled as candidate keyframes, and considers them to collectively represent a short-term abrupt change or the start of an action segment.

[0096] In each group of consecutive candidate keyframes, only the frame with the earliest timestamp is retained as the representative frame of the change event. This strategy ensures that the keyframe distribution has temporal sparsity and content representativeness, thereby avoiding the burden of high-frequency redundant frames on subsequent processing modules. The final set of all retained candidate keyframes constitutes a multi-frame keyframe sequence. This sequence has strong representational power and the ability to sparsely express dynamic information, providing a compact and complete input foundation for subsequent spatial feature extraction and temporal modeling.

[0097] This workflow emphasizes a layered selection mechanism from pixel-level decoding to semantic mutation. Each stage is based on explicit data structure processing modules and adjustable parameter mechanisms to ensure the repeatability and cross-task adaptability of keyframe extraction. Noise reduction and inter-frame difference calculation establish a robust connection between spatial feature preservation and temporal feature modeling. The keyframe selection strategy is based on an event detection strategy rather than static sampling, thereby enhancing the responsiveness to temporal dynamics in complex scenes.

[0098] This embodiment generates a frame sequence by decoding and denoising the input video, and identifies multiple keyframes based on the similarity relationship between adjacent frames. This effectively removes redundant information from the video while maintaining image quality, retaining only representative frames with significant semantic changes. This process restores image clarity while capturing inter-frame structural changes and extracting the frame sequence that contributes most to scene understanding. By constructing a dynamic similarity judgment and candidate frame selection mechanism, adaptive compression representation of the frame sequence is achieved, reducing redundant calculations in downstream feature extraction, improving spatial modeling efficiency, and enhancing the semantic carrying capacity of keyframes in tasks such as action recognition and event inference. The final generated multiple keyframes possess temporal sparsity, expressive completeness, and task orientation, providing a structurally clear visual input foundation for cross-modal fusion.

[0099] In one embodiment, step S20 above includes:

[0100] S201 employs an improved sliding window attention mechanism architecture to process each key frame and generate spatial features for each key frame.

[0101] S202, arrange all spatial features in chronological order to form a spatial feature sequence;

[0102] S203, input the spatial feature sequence into a long short-term memory network, and extract the temporal feature vector through the long short-term memory network;

[0103] S204, Connect all spatial features in the spatial feature sequence to form a long vector of spatial features;

[0104] S205, concatenate the spatial feature long vector and the temporal feature vector to generate a fused input vector;

[0105] S206, the fused input vector is processed through a fully connected layer, and nonlinear activation processing is performed on the processing result of the fully connected layer to generate video spatiotemporal features.

[0106] In this embodiment, to fully extract the structural information of multiple keyframes in both spatial and temporal dimensions, spatial structure modeling needs to be performed separately for each keyframe. When processing each keyframe image, to simultaneously consider both the image's local detail perception capability and global context modeling capability, an improved architecture with structural enhancements is adopted based on the standard sliding window attention mechanism. Traditional sliding window attention mechanisms, due to their completely independent computation between local windows, often fail to effectively model cross-regional semantic relationships in images, especially when action chains span multiple spatial regions. Therefore, this mechanism introduces two improvements: firstly, it introduces cross-window communication paths between local windows, specifically by introducing shared query vectors in the edge regions of adjacent windows, allowing edge pixels to establish global connections across windows; secondly, it designs an inter-layer weight sharing mechanism, that is, during the stacking of multi-layer window attention, it retains the attention weights of some key channels for cross-layer initialization, thereby maintaining the sensitivity of lower layers to edges and textures while enhancing the robust aggregation capability of higher layers for semantic structures. This improvement not only significantly enhances the representation quality of spatial feature extraction but also maintains the locality and linear complexity of the processing computation, providing more structured input features for subsequent temporal modeling.

[0107] This architecture achieves accurate perception of texture, edges, and object boundaries in local image regions by setting multiple local windows on the image and performing self-attention computation within each window. Building upon this, to enhance information interaction between different regions, a cross-window connection mechanism is introduced, sharing attention weights between sliding windows, thereby achieving global semantic linkage while maintaining local sensitivity. The output of each keyframe after processing is a set of high-dimensional spatial feature vectors, fully expressing the structural features and semantic relationships of salient regions in the image.

[0108] After acquiring the spatial features of all keyframes, these features need to be arranged in chronological order according to the original frame sequence to form a spatial feature sequence. The temporal consistency and spatial continuity of this sequence lay the foundation for subsequent temporal feature extraction. Based on this sequence, a Long Short-Term Memory (LSTM) network is introduced to learn the temporal dependencies between frames. The LSTM structure captures the dynamic change patterns between keyframes through a gating mechanism, effectively modeling both short-term abrupt changes and long-term trends, thereby outputting a set of temporal feature vectors. This process not only reflects the evolutionary trajectory of keyframes on the timeline but also supplements the continuity of action and rhythmic changes that spatial features cannot cover.

[0109] Simultaneously, to enhance the direct contribution of spatial information to the fusion result, the system connects all single-frame spatial features in the spatial feature sequence to form a unified long spatial feature vector. This long vector preserves inter-frame positional information and enhances structural integrity in the spatial dimension. Subsequently, the system concatenates this long spatial feature vector with the aforementioned temporal feature vector to generate a fusion input vector, thereby constructing an information structure that is mutually referential and complementary in the spatiotemporal dimensions.

[0110] The fused input vector serves as the input to a fully connected neural network, which performs weighted, combined, and nested transformations on different feature dimensions of the input vector to output a high-order semantic representation. To avoid the insufficient expressive power of linear models, the system introduces non-linear activation functions, such as ReLU or Swish functions, into the fully connected output to enhance the model's ability to fit complex feature distributions. The final result is a video spatiotemporal feature. This feature representation simultaneously includes local details of keyframe images, dynamic relationships between frames, and semantic aggregation trends reflected in multi-frame combinations, providing a visual representation that combines continuity and separability for downstream multimodal fusion.

[0111] This embodiment constructs a temporally aligned spatial feature sequence and introduces a Long Short-Term Memory (LSTM) network to extract temporal features. Then, it performs vector-level fusion and nonlinear mapping of the two types of features, achieving unified encoding of spatiotemporal structural information in videos. This process significantly improves the system's ability to perceive dynamic relationships between multiple frames, giving the output video spatiotemporal features stronger expressive and discriminative power in terms of action changes, object trajectories, and scene transitions. Simultaneously, the sliding window attention mechanism enhances the local modeling capability of single-frame images, and combined with the nonlinear mapping capability of fully connected layers, it constructs a representation structure with greater hierarchical depth and consistency across time dimensions, effectively improving the accuracy and robustness of subsequent task perception.

[0112] In one embodiment, step S30 above includes:

[0113] S301, Perform word segmentation and part-of-speech tagging on the input text to generate preprocessed text;

[0114] S302, Input the preprocessed text into the pre-trained language model to generate text semantic features;

[0115] S303, acquires time-series signals from the motion sensor;

[0116] S304, Use a temporal convolutional network to process the time series signal and extract local temporal features;

[0117] S305, the local timing features are input into the gated loop unit to generate action features.

[0118] In this embodiment, to obtain linguistic information representation and action execution features that can be used for cross-modal fusion, semantic preprocessing is first performed on the input text. The first step of semantic preprocessing is word segmentation and part-of-speech tagging. Word segmentation divides continuous text into basic linguistic units, often implemented using methods such as maximum matching algorithms based on dictionary matching, Hidden Markov Models (HMMs) based on statistical learning, or BiLSTM-CRF models based on deep neural networks. Part-of-speech tagging determines the language category of each segmented unit based on its lexical and syntactic functions in the context, such as noun, verb, adjective, etc., facilitating the extraction of semantic role relationships by the subsequent language model. The preprocessed text output at this stage not only retains the basic semantic structure of the language but also possesses lexical and syntactic encoded information, providing an accurate contextual basis for semantic modeling.

[0119] The processed text is input into a pre-trained language model to generate semantic features. The pre-trained language model typically employs a Transformer architecture, using a self-attention mechanism to model dependencies between words globally, thereby generating context-aware language representations. Specific models can utilize structures such as BERT, RoBERTa, or DeBERTa. Depending on the training task, masked language modeling or next-sentence prediction tasks can be used to fine-tune the model. In this process, word vectors are encoded through multiple Transformer layers to obtain a high-dimensional feature representation containing contextual semantic dependencies, syntactic structure information, and abstract conceptual relationships—that is, text semantic features. This representation possesses good cross-task transfer capabilities, supporting subsequent co-modeling with video and action modalities.

[0120] In addition to text, motion sensor data is also needed to obtain the physical behavioral characteristics of users or devices. Motion sensors may include modules such as accelerometers, gyroscopes, and inertial measurement units (IMUs), which acquire time-series signals through continuous sampling. These signals typically record the motion state along each axis in three-dimensional vector form, exhibiting strong time dependence and periodic fluctuation characteristics. To extract local behavioral patterns, a temporal convolutional network is first used to process the time-series signal. The temporal convolutional network extracts change patterns within a continuous time window through one-dimensional convolution operations and progressively expands the receptive field through multiple convolutional layers to capture dynamic changes at different time scales. The local temporal features output by the convolutional layers preserve the rhythmicity and short-term trends of the actions, providing a foundation for higher-level sequence modeling.

[0121] Subsequently, to model long-term dependencies in the action sequence, local temporal features are input into a gated recurrent unit (GRU). GRU is a lightweight recurrent neural network architecture with update and reset gate mechanisms, effectively capturing temporal dynamics in long-term sequences while reducing the gradient vanishing problem. Through the state recursion structure of the GRU, the model can learn behavioral evolution paths across multiple time steps, and the output action feature vector contains continuous changes, behavioral patterns, and periodic trends throughout the entire action process. This feature possesses strong temporal and physical semantic expressive power, providing accurate support for the action dimension in multimodal fusion.

[0122] Textual semantic features and action features constitute two independent but complementary modal dimensions: one reflects intent and language structure, while the other characterizes physical execution behavior. After completing this stage of processing, the system possesses two feature foundations that express language content and action performance, providing stable input for subsequent multimodal joint modeling.

[0123] This embodiment introduces a text processing path based on semantic modeling and a motion sensor data processing path based on temporal modeling. This allows for the acquisition of textual semantic features expressing intent and motion features reflecting the execution process, while maintaining the integrity of information within each modality. Word segmentation and part-of-speech tagging in the text processing flow ensure a clear basic structure for language input, while the pre-trained language model provides high-order semantic representations shared across tasks. The motion signal processing flow combines temporal convolutional networks with gated recurrent units to achieve synergy between local pattern recognition and long-term sequence modeling, significantly improving the ability to understand complex behavioral patterns. This structure allows the features of both textual intent and motion performance to complement each other in subsequent fusion, laying a solid foundation for the system's response accuracy and adaptability in dynamic environments.

[0124] In one embodiment, step S40 above includes:

[0125] S401, Perform vector alignment processing on the video spatiotemporal features, the text semantic features and the action features to obtain an aligned feature vector;

[0126] S402, input the aligned feature vector into the multi-head self-attention mechanism;

[0127] S403, The aligned feature vector is processed by the multi-head self-attention mechanism to generate a fused feature vector;

[0128] S404, the fused feature vector is used as the fused feature.

[0129] In this embodiment, the video spatiotemporal features, text semantic features, and action features originate from the visual processing module, language processing module, and sensor perception module, respectively. They possess different dimensions, scales, representational spaces, and statistical distributions. Therefore, directly concatenating or fusing these features can lead to an imbalance in information interaction between modalities and may even cause fusion degradation. Consequently, before performing multimodal fusion, vector alignment processing must first be performed on the features of the three modalities to construct a unified representational space.

[0130] Vector alignment can be achieved through a shared embedding space construction method. Specifically, this involves projecting video spatiotemporal features, text semantic features, and action features using three different linear transformation layers to unify their dimensions and adjust their distribution trends to converge. The weight matrix of the linear transformation can be automatically learned during end-to-end training, aiming to eliminate inconsistencies in the spatial structure of different modal features. Layer normalization and nonlinear activation functions, such as ReLU or GELU, can also be used in this process to enhance the ability to model nonlinear relationships. The processed aligned feature vector set not only has dimensional consistency but also possesses statistically co-expressive power.

[0131] The aligned feature vector set is input into a multi-head self-attention mechanism for deep interaction modeling. This mechanism, built on the Transformer framework, works by using multiple attention heads in parallel to compute attention weight matrices in different subspaces and capture the relationships between modalities within each subspace. With h attention heads, each head maps the aligned feature vector set to a query vector Q, a key vector K, and a value vector V, respectively, and calculates an attention score matrix. This matrix is ​​then normalized using a softmax function and multiplied by the value vector to generate a weighted representation. The outputs from multiple attention heads are concatenated into a fused feature representation, which is further integrated into a unified expression vector through linear mapping. This structure can capture heterogeneous interaction patterns between multiple modalities, such as the potential correlation between language descriptions and action changes, or the indicative connection between object motion trends and language targets in video frames.

[0132] The result of multi-head attention processing is a fused feature vector, which expresses the complex relationship between video, text, and action in the semantic space. This vector is not merely a simple concatenation of the three modalities at the semantic level, but rather the result of repeated interactions across multiple attention channels, possessing the expressive capabilities of cross-modal alignment, semantic co-construction, and structural collaboration. It can be directly used to generate the final task-related perceptual decision representation.

[0133] This fusion process not only overcomes the information independence problem in traditional modal fusion, but also achieves a unified representation of multi-dimensional semantic information by explicitly modeling the semantic alignment path between different modalities. It is especially suitable for handling complex task scenarios with temporality, abstract language descriptions and dynamic action signals.

[0134] This embodiment effectively addresses the differences in dimensionality, distribution, and semantic representation of video, language, and action features by introducing a collaborative fusion structure of vector alignment and multi-head self-attention. The vector alignment stage unifies the representation spaces of each modality, providing a structural foundation for fusion. The multi-head attention mechanism learns the potential dependencies between different modalities in multiple semantic subspaces, achieving cross-modal semantic alignment and dynamic semantic reconstruction. This processing path not only significantly improves the consistency of fused feature representation and contextual relevance but also enhances the model's ability to perceive collaborative changes between modalities in complex scenes, making the subsequently generated perceptual vectors more robust, generalizable, and adaptable to decision-making.

[0135] In one embodiment, step S50 above includes:

[0136] S501, The fused features are input into a perceptual network based on the Transformer architecture;

[0137] S502, The fused features are processed through the multi-head self-attention mechanism of the perception network to generate attention-weighted features;

[0138] S503, perform feature integration on the attention-weighted features to generate a perceptual vector;

[0139] S504, Select the decision module based on the task type;

[0140] S505, The perception vector is input into the selected decision module, and the decision module generates task instructions.

[0141] In this embodiment, the fused features, as a unified representation integrating video spatiotemporal information, linguistic semantic representation, and action signals, need to be further converted into structured instruction information that can be used for downstream task execution. This process first introduces a perceptual network based on a Transformer structure, aiming to fully capture the potential high-order semantic dependencies and task-related structures within the fused features. The perceptual network consists of multiple Transformer encoder layers, each including a multi-head self-attention mechanism and a feedforward network module, possessing the ability to perform nonlinear mapping, feature reweighting, and semantic modeling on the input representation.

[0142] After the fused features are input into the perceptual network, information attention weights are first redistributed through a multi-head self-attention mechanism. Under this mechanism, the perceptual network constructs multiple combinations of query vectors, key vectors, and value vectors, and calculates an attention matrix within each combination, generating dynamic attention weights through softmax normalization. These attention weights are used to perform a weighted summation of the value vectors, outputting a reconstructed representation of information across multiple attention channels. Finally, a linear transformation integrates these representations into a unified attention-weighted feature representation. With parallel modeling across multiple attention heads, this representation captures the multi-level semantic dependencies and structural information within the fused features.

[0143] After obtaining attention-weighted features, the perceptual network ensures feature stability through residual connections and layer normalization, and combines a feedforward neural network to further process and project the higher-order representations. The representation vector output by this process is the perceptual vector, which has a highly compressed information expression capability and can reflect the correspondence between the input fused features and the task intent. The perceptual vector is the internal decision-making basis of the task perceptual model, carrying not only a compressed expression of the original multimodal information, but also containing the key semantic paths extracted through the Transformer network.

[0144] After the perception vector is generated, different decision modules need to be selected based on the specific task type to adapt to diverse task response requirements. Task types can be set by an external system or automatically identified through a task classifier in the fused features, including semantic question-answering instructions, action planning instructions, and visual tracking instructions. The distinction between task types determines the structural selection and mapping path of the decision modules. For example, the language generation module uses a sequence generation model to construct instruction text, the action control module uses an instruction decoder to predict joint angle sequences or path trajectories, and the perception guidance module outputs attention areas or warning markers through attention distribution.

[0145] The perception vector serves as input, and after processing by the selected decision module, it outputs a structured task instruction. This task instruction has a clear execution objective and operable parameters, supporting further parsing by the execution system or control of device responses. The entire generation process employs an end-to-end mapping mechanism to ensure continuity, semantic consistency, and decision accuracy from multimodal input to task instruction, making it particularly suitable for complex scenarios where tasks highly rely on multimodal cross-information.

[0146] This embodiment introduces a perceptual network structure based on the Transformer architecture, combined with a multi-head self-attention mechanism and a task-type adaptive decision module, to achieve efficient mapping of fused features to perceptual vectors and structured output of task instructions. This mechanism effectively improves the model's ability to model complex multimodal relationships, enhances the completeness of semantic expression, and improves the relevance of instruction generation. The perceptual network deeply mines high-order semantic dependencies in fused features, and the multi-head mechanism models multi-source information paths in parallel across multiple subspaces, significantly improving the accuracy of task recognition and the operability of generated task instructions. By adaptively switching the decision module according to the task type, the model achieves generalization and flexible response to multi-task scenarios such as language generation, action planning, and visual interaction.

[0147] In one embodiment, after step S50 above, the method further includes:

[0148] S601, Collect action execution deviation data of the task instruction during the execution process;

[0149] S602, Analyze the semantic accuracy data of the language response to the task instruction;

[0150] S603 records data on the accuracy of video scene understanding;

[0151] S604 collects performance index data of the video preprocessing module, feature extraction module, and fusion module;

[0152] S605, integrate the action execution deviation data, language response semantic accuracy data, video scene understanding correctness index data and performance index data into multi-dimensional feedback information;

[0153] S606, The multi-dimensional feedback information is combined with the input video, input text and motion sensor signals to form training samples;

[0154] S607, the network parameters of the video preprocessing module, feature extraction module and fusion module are updated through the backpropagation module;

[0155] S608, dynamically adjust the keyframe extraction threshold, cross-modal fusion attention parameters, and perception network weight parameters based on the multi-dimensional feedback information.

[0156] In this embodiment, the execution results of task instructions in a real environment are typically affected by various factors, including scene changes, missing modal information, and model errors. Therefore, constructing a multi-dimensional evaluation and update mechanism based on execution feedback is crucial for achieving closed-loop perception decision-making and dynamic adaptation. This step, after generating task instructions, further collects and analyzes execution feedback data, and uses the feedback information for training and parameter optimization of multiple modules, enabling continuous evolution of system performance.

[0157] After the task instructions are generated, the first step is to collect motion execution deviation data. This means that during the process of the task instructions driving the execution entity (such as a robot or embedded control unit) to complete the physical action, the offset between the target motion trajectory and the actual executed action is recorded. This deviation can be obtained through pose tracking, inertial sensor feedback, or visual positioning error assessment, and is usually quantified as indicators such as Euclidean distance, angular deviation, or time delay, used to measure the accuracy of the instructions and the consistency of execution.

[0158] Simultaneously, the analysis examines the semantic matching degree between the language responses guided by task instructions and the target intent, generating semantic accuracy data for the language responses. This analysis, based on semantic similarity models or natural language understanding evaluation systems, assesses whether the generated language correctly expresses the perceived target or instruction requirements, and can employ multi-dimensional indicators such as BLEU, ROUGE, and BERTScore. This type of data reflects potential semantic offset issues in the mapping path from perceptual vectors to language, serving as a crucial feedback source for semantic control accuracy.

[0159] Further data on the correctness of video scene understanding will be recorded to evaluate the system's ability to understand scene structure, object relationships, and dynamic changes when processing input video information. This can be obtained through methods such as object detection accuracy, behavior recognition consistency, or event tracking completeness. Reference benchmarks will be generated by combining manual annotation or multi-model comparisons to serve as the evaluation criteria for visual understanding capabilities.

[0160] During system operation, performance metrics data for the video preprocessing module, feature extraction module, and fusion module also need to be collected. This performance data includes computation time, memory usage, inference throughput, and model convergence curves, aiming to reflect the system's stability and responsiveness under different hardware resources and input loads. Combining performance metrics with structural optimization or resource scheduling algorithms can assist the system in making strategic decisions regarding parameter fine-tuning or model replacement.

[0161] The four types of feedback data mentioned above are integrated into a multi-dimensional feedback information with consistent structure and unified dimensions, uniformly representing the system's performance characteristics in each stage of perception, understanding, expression, and execution. To achieve the system's self-correction capability, this multi-dimensional feedback information is linked and matched with the initial input video frame sequence, text data, and motion sensor signals to form training samples. This process not only preserves the task context information but also ensures that the feedback results accurately correspond to the original input path, enhancing the targeted nature of model updates.

[0162] The training samples are then input into the backpropagation module to perform parameter gradient calculation and network weight update operations, updating the network parameters of the video preprocessing module (e.g., keyframe extraction structure), feature extraction module (e.g., sliding attention and temporal modeling structure), and fusion module (e.g., multi-head attention layer and nonlinear activation path), respectively. The parameter update process constructs a loss function based on the feedback error signal and uses gradient descent or variant optimization algorithms to complete iterative correction, achieving a sensitive response of the model to the feedback signal.

[0163] In addition to training the network structure layers, the keyframe extraction threshold, cross-modal fusion attention parameters, and perceptual network weight parameters are dynamically adjusted based on feedback information. The keyframe extraction threshold affects the accuracy of temporal modeling and the redundancy compression effect, and its adjustment is based on the fluctuation range of inter-frame similarity and the distribution of scene complexity. The fusion attention parameters affect the participation strength and coupling mode of each modality's information, and the attention mapping weight matrix can be adjusted according to the modal deviation characteristics in the feedback. The perceptual network weight parameters control the decision path for task instruction generation, and its fine-tuning can alleviate the systematic error in low-precision instruction generation. The overall adjustment is driven by multi-dimensional feedback data, realizing a dynamic parameter tuning closed loop from input understanding to output control.

[0164] Example Description: In the healthcare field, to assist surgical robots in performing complex collaborative tasks, the system first acquires input video for preoperative simulation. This video is simultaneously captured from multiple angles, recording the surgeon's hand movements, tool swings, and real-time changes in the surgical site. The system uses an efficient video coding standard to decode the input video, generating a sequence of original image frames. Subsequently, a trained convolutional autoencoder is used to perform frame-by-frame noise reduction on the frame sequence to eliminate low-quality frames caused by lighting interference, instrument reflections, and camera shake, generating a clean, denoised frame sequence.

[0165] To extract representative key moments from the video, the system uses a dynamic time warping module to calculate the similarity between each frame in the denoised frame sequence and its subsequent frames. When the similarity between a frame and its successor is below a set threshold, it indicates that the frame may be a point of change in visual state, i.e., a candidate keyframe. The candidate keyframe sequence, sorted by time, is further grouped for identification. If consecutive points of change exist, only the earliest frame in each group is retained as the final keyframe to avoid redundancy.

[0166] The system applies an attention architecture based on an improved sliding window mechanism to extract spatial structural features frame-by-frame from the aforementioned multi-frame keyframes. For example, in laparoscopic images, the window mechanism can focus on key areas such as blood vessel edges, tissue surfaces, and tool contact points, and enhance the understanding of the global surgical scene through cross-window connections. The extracted spatial features are then used to construct a spatial feature sequence in chronological order and input into a long short-term memory network to capture the instrument movement trajectory and tissue deformation trends between frames, forming a temporal feature vector. To enhance the overall spatiotemporal coupling capability, the system concatenates all spatial features to form a long spatial feature vector, which is then spliced ​​with the temporal features to generate a fused input vector. This fused input vector is mapped through a fully connected layer and processed by nonlinear activation to ultimately form complete video spatiotemporal features, capable of characterizing intraoperative action patterns, operational rhythm, and environmental conditions.

[0167] In parallel, the system receives medical text descriptions synchronized with the video, such as doctor's instructions transcribed from speech or preoperative preparation materials. The system performs word segmentation and part-of-speech tagging to construct a structured text representation and generates semantic features using a pre-trained language model in the medical field, such as BioBERT or ClinicalBERT. Furthermore, the system collects data from inertial motion sensors worn by the surgeon to obtain time-series signals of their upper limb movements. It then uses a temporal convolutional network to identify micro-motion features, such as wrist rotation angular velocity and changes in grip strength, before inputting these features into a gated recurrent unit to extract continuous motion features.

[0168] Subsequently, the system aligns the video spatiotemporal features, text semantic features, and action feature execution vectors to ensure consistency across time, semantic space, and data scale. These are then fed into a multi-head self-attention mechanism for cross-modal fusion. During fusion, the self-attention mechanism dynamically allocates attention weights based on the importance of each modal feature to the overall task, achieving a mapping between video keyframes and terminology commands and action rhythms. The resulting fused features comprehensively characterize the visual presentation, language objectives, and execution strategies of the current task.

[0169] The system integrates feature inputs into a Transformer-based perceptual network, extracts key task-related features through a multi-head self-attention mechanism, generates attention-weighted features, and integrates them to obtain the final perceptual vector. Based on the task type (e.g., hemostasis, suturing, tissue separation), the system selects the corresponding decision module, inputs the perceptual vector into that module for inference, outputs operation execution instructions, and drives the surgical robot to complete the next action.

[0170] During execution, the system continuously collects task feedback information, including spatial deviations in action execution (e.g., the robotic arm failing to accurately reach the target position); semantic accuracy of verbal responses (e.g., the system's response deviation to the doctor's voice prompts); and the correctness indicators of video scene understanding (e.g., misidentification of surgical tools or target areas). In addition, the system also collects performance indicators such as processing latency, accuracy, and module output stability of the video preprocessing module, feature extraction module, and fusion module in real time. All feedback data is uniformly integrated into multi-dimensional feedback information and used with the original input video, text, and motion sensor signals to form training samples, updating the network parameters of the aforementioned modules through a backpropagation mechanism. The system also dynamically adjusts the keyframe extraction threshold, cross-modal fusion attention parameters, and the weight structure of the perception network to better adapt to individual doctor operating styles and task scene characteristics.

[0171] In intraoperative trials, the system demonstrated good adaptability to different surgeons' operating habits, quickly focusing on key operational steps, extracting information appropriately, and dynamically providing feedback and adjustments, effectively improving surgical precision and robot response efficiency.

[0172] In the smart counter service scenario of fintech business, the system first receives live video for customer interaction. This video records changes in the customer's behavioral state in the counter area, such as standing, sitting, signing, and submitting materials. The system decodes the input video using an efficient video coding standard, extracts the image frame sequence, and performs frame-by-frame noise reduction using a convolutional autoencoder to filter out interference factors such as glass reflections, personnel obstruction, and low-light blur, generating a clear, denoised frame sequence. Based on the similarity between temporally adjacent frames, the system identifies frames where the customer's behavior changes state and selects representative image frames as candidate keyframes. If multiple frames are concentrated in a continuous period of change, only the earliest one is retained to ensure that the keyframes are representative of the behavioral nodes.

[0173] The system employs a sliding window attention network to process each keyframe, perceiving spatial structural information such as document edges, customer body orientation, and equipment placement changes within the counter. Introducing a cross-window attention mechanism enhances responsiveness to interactive information between different areas of the counter layout. After completing spatial structural modeling of all keyframes, the system arranges these spatial features chronologically to form a sequence and inputs it into a long short-term memory network to identify temporal behavioral patterns such as "customers submitting materials and waiting" and "employees leaving the window and returning." It extracts the rhythm and contextual relationships between frame operations and concatenates long vectors of spatial features with temporal features to construct a spatiotemporal fusion vector. Through fully connected layers, structural reconstruction and nonlinear activation mapping are performed to obtain the spatiotemporal feature representation of the video.

[0174] The system synchronously collects customer language input, such as when a user says "I'm here to change my account opening mobile number" or "My transfer hasn't arrived," and the text is transcribed into speech in real time. The system performs word segmentation and part-of-speech tagging on the text and inputs it into a domain-specific language model to generate semantic vectors, recognizing semantic intent, core keywords, and sentence structure. Simultaneously, motion sensors embedded in the counter operation equipment collect user hand movements, such as clicking the touchscreen or a pen touching the surface. Local micro-motion features are extracted using a temporal convolutional network, and a gated recurrent structure is used to identify the complete hand movement sequence as motion features.

[0175] The system performs vector alignment on the three types of information mentioned above to eliminate the effects of dimensionality differences and temporal offsets. This data is then input into a multi-head self-attention structure, combining the internal weights of each modality with cross-modal alignment weights to calculate a set of fused feature vectors. These vectors encompass behavioral change trends in the video, task instruction intent in the text, and evidence of user physical interaction in the action sequence, providing a three-dimensional information foundation for subsequent execution responses.

[0176] The system integrates feature inputs to build a perception network on a Transformer architecture, identifying the most representative task perception vector in the current counter context, such as "need to verify customer identity and retrieve the mobile number change interface." Based on the perceived task type in the scenario (e.g., identity verification, information modification, anomaly query), the system selects the corresponding decision module, inputs the perception vector for policy reasoning, outputs operation instructions, and drives the system to jump to the relevant business interface or link with external terminals for operation.

[0177] The system continuously collects feedback information during task execution. For example, user requests for error correction from the voice response system indicate inaccurate language comprehension; camera detection of customers repeatedly performing the same action suggests that the instruction was not executed correctly; system module response latency under high load reflects performance bottlenecks. The feedback data, including action execution deviations, semantic accuracy, visual recognition correctness, and system processing performance, are uniformly integrated into multi-dimensional feedback information and used with the original input data to form training samples. The parameters of the video processing structure, feature extraction model, and fusion mechanism are updated through backpropagation. Furthermore, the similarity threshold during keyframe extraction, the attention strategy for cross-modal fusion, and the weight parameters in the perception model are dynamically adjusted based on the current task feedback results to optimize the stability and accuracy of subsequent task responses.

[0178] In practical deployments, this system can be widely applied to scenarios such as smart branches, remote video tellers, and mobile terminal interaction. It assists the system in understanding customer task intentions in real time, identifying actual execution status, and making multimodal fusion decisions, providing support for highly automated and differentiated responses in intelligent customer service and counter services. Compared to traditional task decision-making systems based on static images or single voice recognition, this multimodal fusion process demonstrates stronger contextual understanding capabilities and task reasoning flexibility.

[0179] This embodiment constructs a multi-source feedback mechanism based on dimensions such as action execution deviation, linguistic semantic accuracy, scene understanding correctness, and module performance. It then jointly constructs training samples with the original input and optimizes the perception and decision-making processes through backpropagation and parameter adjustment. This achieves a closed-loop self-optimization path oriented towards task execution results. This mechanism not only improves the adaptability of video preprocessing, feature extraction, and multimodal fusion but also enhances the system's self-correction capability and robustness in complex tasks. Especially under dynamic conditions of multiple tasks and scenarios, this feedback adjustment path supports dynamic reconstruction of model parameters and adaptation of fusion strategies, significantly improving the consistency of task execution, the accuracy of instruction generation, and the model's sustainable evolution capability.

[0180] In one embodiment, a task instruction generation apparatus based on cross-modal fusion is provided, which corresponds one-to-one with the task instruction generation method based on cross-modal fusion described in the above embodiments. (Refer to...) Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the task instruction generation device based on cross-modal fusion of the present invention. The module includes a video preprocessing module 10, a spatiotemporal feature extraction module 20, a multimodal feature acquisition module 30, a cross-modal fusion module 40, and a perception decision module 50. Detailed descriptions of each functional module are as follows:

[0181] The video preprocessing module 10 is used to perform decoding and noise reduction processing on the input video to generate a frame sequence, and to identify multiple key frames based on the similarity relationship between adjacent frames in the frame sequence.

[0182] The spatiotemporal feature extraction module 20 is used to obtain the spatial features of the multi-frame key frames to form a spatial feature sequence, extract temporal features based on the spatial feature sequence, and fuse the spatial feature sequence with the temporal features to form video spatiotemporal features;

[0183] The multimodal feature acquisition module 30 is used to perform semantic preprocessing on the input text to generate text semantic features and to acquire motion sensor signals to obtain motion features;

[0184] The cross-modal fusion module 40 is used to perform cross-modal fusion on the video spatiotemporal features, the text semantic features, and the action features to generate fused features;

[0185] The perception decision module 50 is used to generate a perception vector based on the fused features and to generate task instructions based on the perception vector.

[0186] In one embodiment, the video preprocessing module 10 is specifically used for:

[0187] The input video is decoded using an efficient video coding standard to generate a sequence of original image frames.

[0188] The original image frame sequence is denoised frame by frame by a convolutional autoencoder to generate a denoised frame sequence.

[0189] The similarity value between each frame and the next frame in the denoised frame sequence is determined based on the dynamic time warping module.

[0190] When the similarity value between a frame and its next frame is lower than a preset threshold, the frame is marked as a candidate keyframe.

[0191] In a chronologically ordered sequence of candidate keyframes, identify groups of consecutively adjacent candidate keyframes.

[0192] For each consecutive adjacent candidate keyframe group, only the earliest candidate keyframe within the group is retained;

[0193] All retained candidate keyframes are used as the final multi-frame keyframes.

[0194] In one embodiment, the spatiotemporal feature extraction module 20 is specifically used for:

[0195] An improved sliding window attention mechanism is used to process each keyframe and generate the spatial features of each keyframe.

[0196] Arrange all spatial features in chronological order to form a spatial feature sequence;

[0197] The spatial feature sequence is input into a long short-term memory network, and the time feature vector is extracted through the long short-term memory network.

[0198] Connect all spatial features in the spatial feature sequence to form a long vector of spatial features;

[0199] The spatial feature long vector and the temporal feature vector are concatenated to generate a fused input vector;

[0200] The fused input vector is processed by a fully connected layer, and nonlinear activation processing is performed on the processing result of the fully connected layer to generate video spatiotemporal features.

[0201] In one embodiment, the multimodal feature acquisition module 30 is specifically used for:

[0202] Perform word segmentation and part-of-speech tagging on the input text to generate preprocessed text;

[0203] The preprocessed text is input into a pre-trained language model to generate text semantic features;

[0204] Acquire time-series signals from motion sensors;

[0205] The time-series signal is processed using a temporal convolutional network to extract local temporal features;

[0206] The local temporal features are input into the gated loop unit to generate action features.

[0207] In one embodiment, the cross-modal fusion module 40 is specifically used for:

[0208] The video spatiotemporal features, the text semantic features, and the action features are vector aligned to obtain an aligned feature vector.

[0209] The aligned feature vector is input into a multi-head self-attention mechanism;

[0210] The aligned feature vector is processed by the multi-head self-attention mechanism to generate a fused feature vector;

[0211] The fused feature vector is used as the fused feature.

[0212] In one embodiment, the perception decision module 50 is specifically used for:

[0213] The fused features are input into a perceptual network based on the Transformer architecture;

[0214] The fused features are processed through the multi-head self-attention mechanism of the perceptual network to generate attention-weighted features;

[0215] The attention-weighted features are integrated to generate a perceptual vector;

[0216] Select the decision module based on the task type;

[0217] The perception vector is input into the selected decision module, which then generates task instructions.

[0218] In one embodiment, the perception decision module 50 is specifically used for:

[0219] Collect action execution deviation data during the execution of the task instructions;

[0220] Analyze the semantic accuracy data of the language responses to the task instructions;

[0221] Record data on the accuracy of video scene comprehension.

[0222] Collect performance metrics data for the video preprocessing module, feature extraction module, and fusion module;

[0223] The action execution deviation data, language response semantic accuracy data, video scene understanding correctness index data, and performance index data are integrated into multi-dimensional feedback information;

[0224] The multi-dimensional feedback information is combined with the input video, input text, and motion sensor signals to form training samples;

[0225] The network parameters of the video preprocessing module, feature extraction module, and fusion module are updated through the backpropagation module;

[0226] The keyframe extraction threshold, cross-modal fusion attention parameters, and perceptual network weight parameters are dynamically adjusted based on the multi-dimensional feedback information.

[0227] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used for communication with external user terminals via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a task instruction generation method based on cross-modal fusion on the server side.

[0228] In one embodiment, a computer device is provided, which may be a user terminal, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a task instruction generation method based on cross-modal fusion on the user side.

[0229] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:

[0230] Decoding and noise reduction are performed on the input video to generate a frame sequence, and multiple key frames are identified based on the similarity relationship between adjacent frames in the frame sequence.

[0231] The spatial features of the multiple key frames are obtained to form a spatial feature sequence. Temporal features are extracted based on the spatial feature sequence, and the spatial feature sequence and the temporal features are fused to form video spatiotemporal features.

[0232] Semantic preprocessing is performed on the input text to generate text semantic features, and motion sensor signals are collected to obtain motion features;

[0233] Cross-modal fusion is performed on the video spatiotemporal features, the text semantic features, and the action features to generate fused features;

[0234] A perception vector is generated based on the fused features, and task instructions are generated based on the perception vector.

[0235] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0236] Decoding and noise reduction are performed on the input video to generate a frame sequence, and multiple key frames are identified based on the similarity relationship between adjacent frames in the frame sequence.

[0237] The spatial features of the multiple key frames are obtained to form a spatial feature sequence. Temporal features are extracted based on the spatial feature sequence, and the spatial feature sequence and the temporal features are fused to form video spatiotemporal features.

[0238] Semantic preprocessing is performed on the input text to generate text semantic features, and motion sensor signals are collected to obtain motion features;

[0239] Cross-modal fusion is performed on the video spatiotemporal features, the text semantic features, and the action features to generate fused features;

[0240] A perception vector is generated based on the fused features, and task instructions are generated based on the perception vector.

[0241] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0242] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0243] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0244] It should be noted that if any software tools or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The embodiments described above are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A task instruction generation method based on cross-modal fusion, characterized in that, Includes the following steps: Decoding and noise reduction are performed on the input video to generate a frame sequence, and multiple key frames are identified based on the similarity relationship between adjacent frames in the frame sequence. The spatial features of the multiple key frames are obtained to form a spatial feature sequence. Temporal features are extracted based on the spatial feature sequence, and the spatial feature sequence is fused with the temporal features to form video spatiotemporal features. Semantic preprocessing is performed on the input text to generate text semantic features, and motion sensor signals are collected to obtain motion features; Cross-modal fusion is performed on the video spatiotemporal features, the text semantic features, and the action features to generate fused features; A perception vector is generated based on the fused features, and a task instruction is generated based on the perception vector.

2. The task instruction generation method based on cross-modal fusion as described in claim 1, characterized in that, The input video is decoded and denoised to generate a frame sequence, and multiple keyframes are identified based on the similarity relationship between adjacent frames in the frame sequence, including: The input video is decoded using an efficient video coding standard to generate a sequence of original image frames. The original image frame sequence is denoised frame by frame by a convolutional autoencoder to generate a denoised frame sequence. The similarity value between each frame and the next frame in the denoised frame sequence is determined based on the dynamic time warping module. When the similarity value between a frame and its next frame is lower than a preset threshold, the frame is marked as a candidate keyframe. In a chronologically ordered sequence of candidate keyframes, identify groups of consecutively adjacent candidate keyframes. For each consecutive adjacent candidate keyframe group, only the earliest candidate keyframe within the group is retained; All retained candidate keyframes are used as the final multi-frame keyframes.

3. The task instruction generation method based on cross-modal fusion as described in claim 1, characterized in that, The process involves acquiring spatial features from the multiple keyframes to form a spatial feature sequence, extracting temporal features based on the spatial feature sequence, and fusing the spatial feature sequence with the temporal features to form video spatiotemporal features, including: An improved sliding window attention mechanism is used to process each keyframe and generate the spatial features of each keyframe. Arrange all spatial features in chronological order to form a spatial feature sequence; The spatial feature sequence is input into a long short-term memory network, and the time feature vector is extracted through the long short-term memory network. Connect all spatial features in the spatial feature sequence to form a long vector of spatial features; The spatial feature long vector and the temporal feature vector are concatenated to generate a fused input vector; The fused input vector is processed by a fully connected layer, and nonlinear activation processing is performed on the processing result of the fully connected layer to generate video spatiotemporal features.

4. The task instruction generation method based on cross-modal fusion as described in claim 1, characterized in that, Semantic preprocessing is performed on the input text to generate text semantic features, and motion sensor signals are collected to obtain motion features, including: Perform word segmentation and part-of-speech tagging on the input text to generate preprocessed text; The preprocessed text is input into a pre-trained language model to generate text semantic features; Acquire time-series signals from motion sensors; The time-series signal is processed using a temporal convolutional network to extract local temporal features; The local temporal features are input into the gated loop unit to generate action features.

5. The task instruction generation method based on cross-modal fusion as described in claim 1, characterized in that, Cross-modal fusion is performed on the video spatiotemporal features, the text semantic features, and the action features to generate fused features, including: The video spatiotemporal features, the text semantic features, and the action features are vector aligned to obtain an aligned feature vector. The aligned feature vector is input into a multi-head self-attention mechanism; The aligned feature vector is processed by the multi-head self-attention mechanism to generate a fused feature vector; The fused feature vector is used as the fused feature.

6. The task instruction generation method based on cross-modal fusion as described in claim 1, characterized in that, Based on the fused features, a perception vector is generated, and task instructions are generated according to the perception vector, including: The fused features are input into a perceptual network based on the Transformer architecture; The fused features are processed through the multi-head self-attention mechanism of the perceptual network to generate attention-weighted features; The attention-weighted features are integrated to generate a perceptual vector; Select the decision module based on the task type; The perception vector is input into the selected decision module, which then generates task instructions.

7. The task instruction generation method based on cross-modal fusion as described in claim 1, characterized in that, After generating a perception vector based on the fused features and generating task instructions based on the perception vector, the method further includes: Collect action execution deviation data during the execution of the task instructions; Analyze the semantic accuracy data of the language responses to the task instructions; Record data on the accuracy of video scene comprehension. Collect performance metrics data for the video preprocessing module, feature extraction module, and fusion module; The action execution deviation data, language response semantic accuracy data, video scene understanding correctness index data, and performance index data are integrated into multi-dimensional feedback information; The multi-dimensional feedback information is combined with the input video, input text, and motion sensor signals to form training samples; The network parameters of the video preprocessing module, feature extraction module, and fusion module are updated through the backpropagation module; The keyframe extraction threshold, cross-modal fusion attention parameters, and perceptual network weight parameters are dynamically adjusted based on the multi-dimensional feedback information.

8. A task instruction generation device based on cross-modal fusion, characterized in that, The task instruction generation device based on cross-modal fusion includes: The video preprocessing module is used to perform decoding and noise reduction on the input video to generate a frame sequence, and to identify multiple key frames based on the similarity relationship between adjacent frames in the frame sequence. The spatiotemporal feature extraction module is used to obtain the spatial features of the multiple key frames to form a spatial feature sequence, extract temporal features based on the spatial feature sequence, and fuse the spatial feature sequence with the temporal features to form video spatiotemporal features; The multimodal feature acquisition module is used to perform semantic preprocessing on the input text to generate text semantic features and to collect motion sensor signals to obtain motion features; The cross-modal fusion module is used to perform cross-modal fusion on the video spatiotemporal features, the text semantic features, and the action features to generate fused features; The perception decision module is used to generate a perception vector based on the fused features and to generate task instructions based on the perception vector.

9. A computer device, characterized in that, The computer device includes a memory, a processor, and a task instruction generation program based on cross-modal fusion stored in the memory and executable on the processor. When the task instruction generation program based on cross-modal fusion is executed by the processor, it implements the steps of the task instruction generation method based on cross-modal fusion as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The storage medium stores a task instruction generation program based on cross-modal fusion, which, when executed by a processor, implements the steps of the task instruction generation method based on cross-modal fusion as described in any one of claims 1-7.

Citation Information

Cited By

  • Micro-action recognition method, electronic equipment, storage medium and program product

    CN121121873A

  • Customer service method and system supporting multi-modal understanding and cross-platform execution

    CN121329437A

  • Spatial intelligent dynamic video frame sampling method and system based on question semantics

    CN121436196A

  • Abnormality detection method and system fusing large and small models and cooperating with cross-time domain perception

    CN121482682A

  • Video compression method and video decoding method

    CN122027825A