Multi-modal hierarchical feature fusion and decision-making method, device, equipment and medium

By employing hierarchical feature extraction, filtering, dimensionality reduction, semantic enhancement, and cross-modal attention fusion, the problem of redundant information interference in multimodal data is solved, achieving efficient multimodal decision processing and improving the accuracy and robustness of decision-making.

CN120951246APending Publication Date: 2025-11-14PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 10 Cited by

Patent Information

Application Number
CN202511059898.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing technologies lack sufficient multimodal fusion processing capabilities, failing to effectively distinguish and filter the importance of features at different levels in multimodal data. This leads to redundant information interference, reducing the accuracy and robustness of decision-making.

Method used

By acquiring visual, linguistic, and action data, hierarchical feature extraction is performed. The importance of features at each level is analyzed, and filtering and dimensionality reduction are carried out. Semantic enhancement is performed, and cross-modal attention fusion is conducted. Finally, the data is input into a semantic reasoning network to generate decision results.

Benefits of technology

It improves the accuracy and robustness of multimodal data processing, reduces interference from irrelevant information, and enhances decision-making capabilities in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120951246A_ABST
    Figure CN120951246A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, can be applied to business scenes such as agent autonomous decision making, financial science and technology and medical health, and discloses a multi-modal hierarchical feature fusion and decision making method, device, equipment and medium, and the method comprises the steps: obtaining vision, language and motion data, and carrying out the hierarchical feature extraction to generate a multi-modal initial feature set; feature importance is analyzed, screening and dimension reduction are carried out, and screened multi-modal features are obtained; performing semantic enhancement on the screened multi-modal features to generate multi-modal semantic enhancement features; executing cross-modal attention fusion on the multi-modal semantic enhancement features to obtain cross-modal fusion features; and inputting the cross-modal fusion features into a semantic reasoning network to generate a decision result. According to the method, through multi-level screening dimension reduction, semantic enhancement and cross-modal attention fusion, the model can accurately utilize key feature relationships among multi-modal data, redundant interference is reduced, and semantic reasoning accuracy and decision-making efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a multimodal hierarchical feature fusion and decision-making method, apparatus, device, and storage medium. Background Technology

[0002] The rapid development of artificial intelligence technology has driven the widespread application of integrated visual, language, and motion data processing models in various fields. Visual-Language-Motion (VLA) models, as a core technology for realizing human-computer interaction and autonomous decision-making by intelligent agents, have gradually become the foundation for applications such as intelligent robots, autonomous driving, and intelligent customer service. However, the insufficient multimodal fusion processing capabilities in existing technologies have become a key bottleneck restricting their further development.

[0003] In the fintech sector, intelligent assistant systems are widely used for user consultation and risk analysis. They typically require a combination of user verbal input, visual authentication data, and behavioral data to provide real-time services and make decisions. However, current VLA models in financial scenarios often simply concatenate or weight speech-to-text data with image data captured by a camera when processing customer input. This fails to distinguish between background elements and key information in transaction document images or filter redundant modifiers in customer descriptions, leading to decreased model processing efficiency and reduced decision accuracy. This is particularly problematic in complex business scenarios where key risk factors are easily overlooked.

[0004] In the healthcare field, intelligent medical robots and assisted diagnostic systems also require comprehensive analysis and decision-making using visual, verbal, and motor data. For example, while a doctor verbally describes a patient's symptoms, the robot needs to acquire visual medical images and combine them with the patient's motor data. However, current technologies often use simple weighting or stacking methods in multimodal fusion, which cannot accurately filter out irrelevant background information in medical images or redundant expressions in the doctor's statements. This increases interference from redundant information, affecting the accuracy of diagnostic results and processing speed. Furthermore, existing models have poor robustness for medical records with missing or noisy multimodal data, making it difficult to maintain stable diagnostic performance.

[0005] In the field of autonomous decision-making for intelligent agents, such as intelligent shopping guides and service robots, current VLA models still have limited ability to mine complex semantic relationships between visual, linguistic, and action data. When faced with user commands that contain multi-dimensional spatial relationships and attribute descriptions, traditional models struggle to understand the logical relationships between various elements, leading to problems such as incorrect target selection and location misunderstanding, thus limiting their effectiveness in complex and dynamic environments.

[0006] In summary, existing technologies generally suffer from insufficient filtering of redundant information in multimodal data fusion, difficulty in highlighting key features, insufficient ability to mine complex semantic relationships between different modalities, and poor robustness in the case of missing or noisy multimodal data. These shortcomings significantly affect the processing efficiency and decision-making accuracy of the system in the fields of fintech, healthcare, and autonomous decision-making by intelligent agents. Summary of the Invention

[0007] The main objective of this invention is to provide a multimodal hierarchical feature fusion and decision-making method, apparatus, device, and storage medium, aiming to solve the technical problem that the prior art fails to effectively distinguish and screen the importance of features at each level in multimodal data, resulting in redundant information interfering with the multimodal fusion effect and reducing the accuracy and robustness of decision-making.

[0008] To achieve the above objectives, this invention provides a multimodal hierarchical feature fusion and decision-making method, comprising:

[0009] Acquire visual data, language data, and motion data, and perform hierarchical feature extraction on the visual data, language data, and motion data to generate a multimodal hierarchical initial feature set;

[0010] The importance of each level feature in the initial multimodal hierarchical feature set is analyzed, and the initial multimodal hierarchical feature set is filtered and dimensionality reduced based on the importance to obtain the filtered and dimensionality-reduced multimodal hierarchical features;

[0011] The selected and dimensionality-reduced multimodal hierarchical features are semantically enhanced to generate multimodal semantically enhanced features;

[0012] Cross-modal attention fusion is performed on the multimodal semantic enhancement features to obtain cross-modal fused features;

[0013] The cross-modal fusion features are input into a semantic reasoning network for processing to generate decision results.

[0014] Furthermore, to achieve the above objectives, the present invention provides a multimodal hierarchical feature fusion and decision-making apparatus, comprising:

[0015] The multimodal feature extraction module is used to acquire visual data, language data, and action data, and to perform hierarchical feature extraction on the visual data, language data, and action data to generate a multimodal hierarchical initial feature set;

[0016] The multimodal feature filtering module is used to analyze the importance of each level feature in the multimodal hierarchical initial feature set, and to filter and reduce the multimodal hierarchical feature set based on the importance to obtain the filtered and reduced multimodal hierarchical features.

[0017] The multimodal semantic enhancement module is used to semantically enhance the filtered and dimensionality-reduced multimodal hierarchical features to generate multimodal semantically enhanced features.

[0018] A cross-modal fusion module is used to perform cross-modal attention fusion on the multimodal semantic enhancement features to obtain cross-modal fused features;

[0019] The semantic reasoning module is used to input the cross-modal fusion features into the semantic reasoning network for processing and to generate decision results.

[0020] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and a multimodal hierarchical feature fusion and decision program stored in the memory and executable on the processor, wherein when the multimodal hierarchical feature fusion and decision program is executed by the processor, it implements the steps of the multimodal hierarchical feature fusion and decision method as described above.

[0021] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing a multimodal hierarchical feature fusion and decision program, wherein when the multimodal hierarchical feature fusion and decision program is executed by a processor, it implements the steps of the multimodal hierarchical feature fusion and decision method as described above.

[0022] Beneficial Effects: This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as autonomous decision-making by intelligent agents, fintech, and healthcare. It discloses a multimodal hierarchical feature fusion and decision-making method, apparatus, device, and medium, comprising: acquiring visual data, language data, and action data and extracting hierarchical features to generate an initial multimodal hierarchical feature set; analyzing the importance of features at each level in the initial multimodal hierarchical feature set, and obtaining filtered and dimensionality-reduced multimodal hierarchical features based on importance screening and dimensionality reduction; semantically enhancing the filtered and dimensionality-reduced multimodal hierarchical features to generate multimodal semantically enhanced features; performing cross-modal attention fusion on the multimodal semantically enhanced features to obtain cross-modal fused features; and inputting the cross-modal fused features into a semantic inference network for processing to generate a decision result. This invention analyzes and filters the importance of features at each level of multimodal data, reduces interference from irrelevant information, and enhances the expressive power of key features by combining semantic enhancement and cross-modal attention mechanisms, achieving efficient fusion processing and ultimately improving the accuracy and robustness of decision-making in complex scenarios. Attached Figure Description

[0023] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings:

[0024] Figure 1 This is a schematic diagram of an application environment for a multimodal hierarchical feature fusion and decision-making method according to an embodiment of the present invention;

[0025] Figure 2 This is a flowchart illustrating an embodiment of the multimodal hierarchical feature fusion and decision-making method of the present invention;

[0026] Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the multimodal hierarchical feature fusion and decision-making device of the present invention;

[0027] Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention;

[0028] Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0029] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0030] The multimodal hierarchical feature fusion and decision-making method provided in this invention can be applied to, for example... Figure 1 In this application environment, the user terminal communicates with the server via a network. The server can acquire visual, linguistic, and action data from the user terminal and perform hierarchical feature extraction to generate a multimodal hierarchical initial feature set. It analyzes the importance of features at each level in the initial feature set, and based on importance filtering and dimensionality reduction, obtains filtered and dimensionality-reduced multimodal hierarchical features. Semantic enhancement is applied to these filtered and dimensionality-reduced multimodal hierarchical features to generate multimodal semantically enhanced features. Cross-modal attention fusion is performed on these multimodal semantically enhanced features to obtain cross-modal fused features. The cross-modal fused features are then input into a semantic inference network for processing to generate decision results. This invention analyzes and filters the importance of features at each level of multimodal data, reducing interference from irrelevant information. It combines semantic enhancement and cross-modal attention mechanisms to improve the expressive power of key features, achieving efficient fusion processing and ultimately improving the accuracy and robustness of decision-making in complex scenarios. The user terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The following detailed description of specific embodiments further illustrates this invention.

[0031] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the multimodal hierarchical feature fusion and decision-making method provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0032] like Figure 2As shown, the multimodal hierarchical feature fusion and decision-making method proposed in this invention includes the following steps:

[0033] S10, acquire visual data, language data and action data, and perform hierarchical feature extraction on the visual data, language data and action data to generate a multimodal hierarchical initial feature set;

[0034] In this embodiment, the processing of visual data, language data, and motion data is the foundation for achieving multimodal information fusion. Visual data can be acquired through image acquisition devices, such as cameras capturing continuous frames or still images, to reflect information about the external environment and the target object. Language data can be acquired through voice recognition devices to collect voice input, or through text interfaces to collect text commands input by users, as an expression of human intentions or needs. Motion data can be acquired through sensors or interfaces that interact with the control system, including the position state of the robotic arm, joint angle sequences, motion trajectories, etc., reflecting the motion state of the target object in space and time.

[0035] Building upon this foundation, visual, linguistic, and motion data are input into a unified data processing pathway for hierarchical feature extraction. The implementation of hierarchical feature extraction relies on the multi-layered structure of deep learning networks. In visual data processing, different levels of visual features can be extracted using convolutional neural network layers; for example, low-level layers extract edges and texture information, mid-level layers extract local shapes and geometric structures, and high-level layers capture abstract semantic concepts. Hierarchical feature extraction of linguistic data can be achieved through various network structures, such as word embedding layers mapping lexical information, convolutional or recurrent networks extracting syntactic or contextual structures, and semantic encoding layers establishing sentence-level or document-level expressions. Hierarchical feature extraction of motion data can be achieved by combining temporal convolutional networks and recurrent neural networks to capture short-term rapid dynamics, medium-term motion trends, and long-term motion intentions. All hierarchical features from visual, linguistic, and motion data are integrated into a multimodal hierarchical initial feature set after feature concatenation or alignment. This set achieves consistent alignment in time and space, enabling subsequent processing to analyze these features within the same reference frame.

[0036] In practical applications, multi-camera systems can be used to acquire image information from different angles to improve the coverage of visual data. Language data input can be transcribed in real-time using a high-precision speech recognition model, and motion data can be captured with high precision across multiple degrees of freedom using sensor arrays with inertial measurement units. During hierarchical feature extraction, different network architectures can be selected for visual data, such as shallow convolutions for close-range detail extraction and deep networks for global context abstraction. Language data processing can combine context-aware pre-trained models with adaptive tuning modules to handle the complexity of different contexts and spoken input. Motion data can be modeled and aggregated using parallel multi-scale networks to represent both short-term dynamics and long-term patterns. During feature integration, a dynamic weighting mechanism can be used to adjust the contribution of different modalities to adapt to different task requirements, such as visual information dominance or language command dominance, in practical applications.

[0037] Example explanation: In the healthcare business field, visual data can correspond to patient medical imaging information, such as CT or MRI images; language data can come from doctors' diagnostic descriptions or patients' complaints; and motion data can be the patient's motor function performance. By simultaneously collecting the above data and extracting hierarchical features, we can help locate lesion areas, understand the semantic content of the patient's statements and analyze their movement patterns, and support doctors in forming accurate diagnostic suggestions.

[0038] In the fintech business, visual data can include customer behavior videos recorded by counter cameras, language data can include customer business requests or customer service interaction voices, and motion data can include user interaction actions on self-service devices. Through the above-mentioned hierarchical feature extraction, intelligent customer service can accurately understand customer intentions, identify abnormal behavior, and support risk assessment and business process optimization.

[0039] In scenarios where intelligent agents make autonomous decisions, visual data can come from environmental perception cameras of service robots or autonomous mobile devices, language data can come from voice commands or text instructions from human users, and motion data can come from sensor feedback of autonomous robot movement. Through the above-mentioned hierarchical feature extraction, the intelligent agent can simultaneously understand the position of the target in the environment, the complex intentions expressed by the user's language, and its own motion state, so as to achieve adaptive perception of the environment and dynamic planning of complex tasks. For example, in a warehousing and logistics environment, it can support robots to accurately locate goods, understand operation instructions, and plan the optimal movement path.

[0040] This embodiment, through the synchronous acquisition and hierarchical feature extraction of the aforementioned multimodal data, can establish a unified representation framework for visual, linguistic, and action information in the initial stage. This representation not only maintains the hierarchical structure and rich information within each modality, but also provides sufficient preparation for subsequent fusion analysis. This enables the accurate identification of key features and the filtering of irrelevant content in the task, thereby improving the processing efficiency and decision-making accuracy of the multimodal interaction system.

[0041] S20, Analyze the importance of each level feature in the initial multimodal hierarchical feature set, and perform screening and dimensionality reduction processing on the initial multimodal hierarchical feature set based on the importance to obtain the screened and dimensionality-reduced multimodal hierarchical features;

[0042] In this embodiment, firstly, to analyze the importance of features at each level in the initial feature set of a multimodal hierarchical dataset, it is necessary to understand that "multimodal" encompasses three types of data: visual, linguistic, and action data. "Hierarchical" refers to the representation of these three types of data at different levels of abstraction. For example, visual data includes low-level texture features, mid-level structural features, and high-level semantic features; linguistic data includes word, phrase, and sentence-level expressions; and action data includes short-term changes, medium-term trends, and long-term patterns. Therefore, analyzing the importance of features at each level means measuring the value of each type of feature in the current task scenario. Specifically, this can be done by calculating the correlation index with the preset task objective, statistically analyzing its discriminative ability on the training set, or using statistical indicators such as entropy and variance to measure the differences in features. Sources for importance analysis include statistical tests (such as t-tests and chi-square tests) proposed in feature selection research, information theory indicators (such as mutual information), or normalization methods for feature weights in machine learning (such as the normalization of L1 regularization coefficients). These can all serve as the theoretical basis and practical methods for the importance analysis of the current task.

[0043] Based on the aforementioned importance, the initial feature set for multimodal hierarchical analysis undergoes screening and dimensionality reduction. Specifically, the results of importance analysis are used to filter features that contribute less to the task objective, for example, by setting a threshold to select features with importance scores higher than the threshold. Simultaneously, dimensionality reduction aims to reduce the dimensionality of the feature space to decrease computational complexity and avoid overfitting. Common implementation methods include Principal Component Analysis (PCA), Factor Analysis, Matrix Factorization, and Linear Discriminant Analysis (LDA). In particular, Matrix Factorization can be performed using Singular Value Decomposition (SVD) or Non-negative Matrix Factorization (NMF), which not only reduce dimensionality but also preserve the main information of the original feature set. Screening and dimensionality reduction are closely related; screening can be seen as removing low-importance features, while dimensionality reduction further compresses the dimensionality of high-importance features. The final output is the screened and dimensionality-reduced multimodal hierarchical features, forming a high-quality, compact feature representation for subsequent processing steps, ensuring the efficiency and accuracy of subsequent operations.

[0044] Importance analysis can be achieved through statistical significance tests, such as using t-tests or chi-square tests to calculate the correlation score between each feature and the target variable, selecting features with scores higher than a set threshold as the screening results. Alternatively, support vector machines (SVM) or random forest models can be trained, utilizing the feature importance weights from the model training process, normalized, and used as the source of importance scores. For the screened feature set, dimensionality reduction can be achieved using PCA, projecting high-dimensional features onto the first few directions with the largest variance, or using SVD decomposition, truncating the first k largest singular values ​​and their corresponding singular vectors to define a low-dimensional space, preserving the main information of the feature set. LDA algorithms can also be used, mapping features to a low-dimensional space with the strongest class separability for multi-class tasks. For large-scale feature sets, screening and dimensionality reduction operations can be combined, for example, using L1-regularized weighted sparse matrix factorization to simultaneously achieve importance constraints and dimensionality compression.

[0045] Example: In the healthcare business, visual data can be patient images, language data can be electronic medical records, and motion data can be patient motion capture signals. By analyzing the importance of features at different levels and filtering and reducing dimensions, such as prioritizing the retention of features related to lesion areas or key diagnostic terms, and removing redundant image textures or repetitive language expressions, the amount of computation can be reduced and the accuracy of diagnostic models can be improved.

[0046] In the fintech business, visual data can be counter monitoring images, language data can be customer service dialogues, and action data can be terminal operation records. By using importance analysis to filter high-risk behavioral features in transaction scenarios and reduce dimensionality, such as reducing irrelevant background video pixels or lengthy embellishing language input, the operating efficiency of risk identification models can be accelerated and the accuracy of risk prediction can be improved.

[0047] In autonomous decision-making scenarios for intelligent agents, visual data can be environmental images from the mobile robot's camera, language data can be user voice commands, and motion data can be records of the robot's motion trajectory. By analyzing and filtering to prioritize the retention of features related to the target area or key command words, removing irrelevant scene features and reducing dimensionality, we can improve the robot's perception efficiency of high-value information in complex environments, reduce computational latency, and enhance real-time autonomous decision-making capabilities.

[0048] This embodiment analyzes the importance of features at each level in the initial feature set of multimodal hierarchical structure, and performs importance-based screening and dimensionality reduction. This allows for the removal of low-relevance or redundant features without losing key information, compressing the feature space, reducing the computational complexity of subsequent processing, and improving the model's ability to focus on key features, thereby enhancing the model's generalization ability and accuracy.

[0049] S30, perform semantic enhancement on the selected and dimensionality-reduced multimodal hierarchical features to generate multimodal semantically enhanced features;

[0050] In this embodiment, semantic enhancement of the filtered and dimensionality-reduced multimodal hierarchical features aims to further mine their potential high-order semantic information based on the efficient and compact features after dimensionality reduction. The filtered and dimensionality-reduced multimodal hierarchical features contain the main feature information retained from visual, linguistic, and action data. These features have undergone importance analysis and dimensionality compression, removing redundant and low-value data. However, the deep semantic connections between modalities and between different levels within the same modality still need to be strengthened. Semantic enhancement operations can be achieved through external knowledge resources or contextual modeling. For example, matching high-level visual features with concept nodes defined in a domain knowledge graph, comparing and supplementing lexical-level linguistic features with definitions in a semantic dictionary or specialized thesaurus, and aligning long-term action pattern features with a predefined behavior database.

[0051] In practical implementations, knowledge graph mapping can use embedding representation techniques, such as TransE or GraphSAGE, to embed graph nodes into the same vector space as the filtered and dimensionality-reduced multimodal hierarchical features. Then, matching nodes are selected using similarity metrics, and their semantic information is integrated to supplement visual features. For language data, semantic enhancement can be achieved through pre-trained language model encoding combined with domain-specific dictionaries for synonym expansion, hyponymy / hypernymy concept supplementation, or disambiguation. For example, for medical language data, word vector similarity combined with a medical lexicon can be used to identify and enhance the semantic relationships between drug names, symptom words, and disease names. For action features, semantic enhancement can be achieved through predefined action category models or action template databases to realize pattern completion and anomaly recognition. For example, it can identify standard pick-and-place action sequences and abnormal posture patterns in robot operations, and achieve feature semantic enhancement by aligning with standard templates.

[0052] The resulting multimodal semantic enhancement features not only contain the main expressions after dimensionality reduction and compression of each modality, but also embed higher-order semantic information supplemented by external knowledge and contextual associations. These serve as inputs for subsequent cross-modal fusion and semantic reasoning, enabling the model to better understand and interpret multimodal information in complex environments.

[0053] Semantic enhancement of visual features can be achieved by constructing knowledge graphs. For example, for objects involved in the environment of intelligent robots, a knowledge graph can be defined using object categories, attributes, and uses, and graph neural networks can be used to calculate the correlation between graph nodes and visual features, thereby enhancing visual representation. Semantic enhancement of action features can also be achieved by introducing domain-specific vocabulary to language features. For example, using the BERT model to encode vocabulary-level features and combining them with relevant definitions retrieved from a specialized thesaurus for feature fusion, thereby enhancing the expression of language features. Semantic enhancement of action features can be achieved through comparison with standard template databases, such as action trajectory template matching. The similarity of action patterns can be calculated using the Dynamic Time Warping (DTW) algorithm, and missing keyframes in the standard template can be supplemented. Various enhancement methods can be used individually or in combination as needed to adapt to the requirements of different task environments.

[0054] Example Explanation: In the healthcare business domain, the linguistic features after dimensionality reduction can be enhanced by entity concepts such as diseases, drugs, and symptoms in the medical knowledge base, enabling the model to more accurately understand the relationships between professional terms involved in the patient's electronic medical record. Visual features can be enhanced by the correspondence between organs, lesion sites, and pathological morphologies defined in the medical imaging knowledge graph, making the image feature expression more consistent with medical professional semantics. Action features can be supplemented by standard medical operation action templates to improve the professionalism of action recognition.

[0055] In the field of fintech business, visual features can be enhanced by matching the category labels of common items in the counter scene with visual features; linguistic features can be expanded by financial service terminology and supplemented by financial regulations; and action features can be compared with standard business operation sequences and abnormal behavior templates to strengthen the semantic expression of risk identification.

[0056] In autonomous decision-making scenarios for intelligent agents, visual features can be enhanced by knowledge graphs of object attributes and uses in the task environment, language features can be expanded by the hierarchical conceptual relationships of command words, and action features can be completed by comparing standard operation templates with task target templates. Overall, this improves the semantic integrity and contextual consistency of the autonomous decision-making process and enhances the perception, understanding, and execution capabilities in complex environments.

[0057] This embodiment enhances the semantics of the selected, dimensionality-reduced multimodal hierarchical features, compensating for any contextual or domain knowledge lost during the dimensionality reduction process. This makes the features not only dimensionally concise but also semantically rich in expression. This processing not only improves the ability of features to express complex semantic relationships implicit in the current task but also enhances the accuracy and robustness of subsequent multimodal information integration and semantic reasoning, ultimately improving the overall decision-making system's application performance in complex scenarios.

[0058] S40, perform cross-modal attention fusion on the multimodal semantic enhancement features to obtain cross-modal fused features;

[0059] In this embodiment, cross-modal attention fusion of multimodal semantic enhancement features aims to address the issues of inconsistent relationships and difficulty in uniformly quantifying the importance of features across different modalities. Multimodal semantic enhancement features are a set of high-value features retained after hierarchical feature extraction, importance filtering, and semantic enhancement, including visual enhancement features, language enhancement features, and action enhancement features. The core idea of ​​cross-modal attention fusion is to dynamically adjust the interrelationships between different modalities using attention mechanisms, highlighting the contributions of key modal features under different contextual conditions.

[0060] In practical implementation, visual enhancement features from the multimodal semantic enhancement features are first used as query input, linguistic enhancement features as the first key-value pair, and action enhancement features as the second key-value pair. By calculating the attention weights between visual and linguistic enhancement features, the correlation between visual and linguistic information is quantified, and these weights are used to weighted convergence of the linguistic enhancement features to generate visual-linguistic fusion features. Next, by calculating the attention weights between visual and action enhancement features, visual-action fusion features are generated. Finally, by calculating the attention weights between linguistic and action enhancement features, linguistic-action fusion features are generated. These fusion operations are all implemented using matrix multiplication and normalization functions (such as softmax), ensuring the interpretability of attention weights while enhancing the consistent expression of key semantics across modalities.

[0061] Finally, the visual-language fusion features, visual-action fusion features, and language-action fusion features are integrated through a weighted summation. The weights can be dynamically learned during model training, thus obtaining cross-modal fusion features. These cross-modal fusion features not only include enhanced representations of each modality but also strengthen the dynamic semantic interaction relationships between different modalities, providing a more compact and consistent representation for subsequent semantic inference network inputs.

[0062] In the fusion of visual and linguistic augmentation features, a multi-head attention mechanism can be employed to achieve multi-perspective information interaction through parallel multiple sets of queries and key-value mappings, enabling independent modeling of the correspondence between visual and linguistic features in different subspaces. The fusion of visual and action augmentation features can be achieved through residual connections after layer normalization to stabilize training and enhance semantic alignment. The fusion of linguistic and action augmentation features can combine bidirectional attention computation, with each serving as a query and key-value pair, allowing action features to focus on important cues in the language regarding the intent of the action. Finally, during weighted summation, learnable dynamic fusion weights or a simple average pooling strategy can be used, with the specific choice adaptable to task requirements.

[0063] Example: In the healthcare business, cross-modal attention fusion can be applied to dynamically combine visual enhancement features of suspicious areas in medical images with linguistic enhancement features of doctors' text descriptions. For example, attention mechanisms can be used to highlight lesion areas in images that are related to text descriptions, thereby improving the multimodal consistency of diagnostic support systems.

[0064] In the fintech business, cross-modal attention fusion can be applied to counter service scenarios, combining the motion enhancement features of customer gestures with the language enhancement features of dialogue content and visual identity verification features. For example, it can dynamically highlight semantic and action signals related to customer risk warnings to improve the accuracy of risk warnings.

[0065] In autonomous decision-making scenarios for intelligent agents, cross-modal attention fusion can be applied to the task environment. The robot perceives objects through visual enhancement features, parses task commands through language enhancement features, and understands the operational context through action enhancement features. Attention fusion enables the robot to adjust its behavioral strategies based on the most relevant modal information in the environment and instructions, thereby improving the accuracy and robustness of autonomous decision-making.

[0066] This embodiment achieves dynamic capture of key semantic interactions between different modalities by performing cross-modal attention fusion on multimodal semantic enhancement features, thereby improving the alignment of cross-modal features in spatial, semantic, and temporal dimensions. This dynamic weight adjustment mechanism effectively reduces the interference of information redundancy and modal conflicts on model performance, enhances the discriminative power of fused features in complex environments, and provides a highly consistent and expressive multimodal unified representation for subsequent decision generation.

[0067] S50, the cross-modal fusion features are input into the semantic reasoning network for processing to generate decision results.

[0068] In this embodiment, cross-modal fusion features are input into a semantic reasoning network for processing and to generate decision results. The aim is to utilize the fused multimodal information to output specific task results that meet the needs of the scenario. Cross-modal fusion features, as a unified representation after dynamic attention weighting across different modalities, already possess strong semantic consistency and spatiotemporal context alignment capabilities. The purpose of inputting them into the semantic reasoning network is to further model the deep semantic relationships, temporal logic, and contextual dependencies within the fusion features, and to derive results that can be used for intelligent execution or decision support.

[0069] In practice, the cross-modal fused features are first mapped to the encoder module of the semantic inference network through the input layer. The encoder typically employs multi-layer self-attention computation units or a stacked Transformer structure, capturing global semantic dependencies in the fused features layer by layer and performing contextual semantic enhancement. The output encoded features are then input into the decoder module, which dynamically adjusts the decoding target according to the task type. If the task is action control, the decoder further maps the encoded features to the action parameter space, for example, mapping them to joint angle sequences or motion vectors through linear layers. If the task is language generation, the decoder maps the encoded features to the vocabulary distribution space, combining the word sequence prediction distribution output by the language model to generate semantically coherent text content.

[0070] The entire processing flow can also optimize network weights through multi-task joint training, making the semantic reasoning network adaptable to different task scenarios and ensuring the accuracy and consistency of action control or language generation results.

[0071] Multi-head self-attention computation units can be used in the encoder of semantic reasoning networks to enhance the dependency modeling capability of cross-modal fusion features within the global context. In action control tasks, the decoder can use residual convolutional units superimposed with linear mappings to output multi-dimensional action prediction results with temporal consistency. In language generation tasks, the decoder can combine conditional language modeling strategies, using cross-modal fusion features as conditional context to guide the distribution of word generation, supporting multi-turn dialogues or paragraph-level responses. Furthermore, to meet different scenario requirements, task conditional identifiers can be introduced to dynamically switch the target decoder for different tasks, and the number of layers, hidden dimensions, and regularization parameters of each module can be adjusted to adapt to application environments with different complexity and real-time requirements.

[0072] Example: In the healthcare business domain, semantic reasoning networks receive cross-modal fusion features that integrate visual, linguistic, and action data as input, and decode the clinical context and temporal information implicit in the fusion features into treatment suggestions. For example, in surgical assistance systems, it automatically infers the relationship between the patient's current position and intraoperative instructions and generates precise surgical operation instructions.

[0073] In the fintech business, semantic reasoning networks receive cross-modal fusion features and then decode comprehensive risk indicators of user interaction behavior. Based on a comprehensive analysis of customer speech, facial expressions and actions in financial counter interactions, they infer customer risk levels and generate compliance recommendations.

[0074] In the scenario of autonomous decision-making by intelligent agents, the cross-modal fusion features are decoded through a semantic reasoning network. Visual cues collected in the environment, semantics of user commands and motion trend information are uniformly modeled and generated to generate specific action parameters or continuous dialogue content that the robot can directly execute. This enables the autonomous task decision-making and execution of intelligent agents under multimodal input conditions.

[0075] This embodiment processes cross-modal fused features into a semantic reasoning network, further uncovering deep semantic dependencies and temporal logical relationships within the fused features, thus improving the accuracy of mapping cross-modal information to specific task results. This processing method effectively solves the contextual misalignment problem of multimodal data in semantic understanding and task decision-making, ensuring highly consistent and accurate decision results under diverse input conditions.

[0076] This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as autonomous decision-making by intelligent agents, fintech, and healthcare. It discloses a method, apparatus, device, and medium for multimodal hierarchical feature fusion and decision-making, comprising: acquiring visual data, language data, and action data and extracting hierarchical features to generate an initial multimodal hierarchical feature set; analyzing the importance of features at each level in the initial multimodal hierarchical feature set, and obtaining filtered and dimensionality-reduced multimodal hierarchical features based on importance screening and dimensionality reduction; semantically enhancing the filtered and dimensionality-reduced multimodal hierarchical features to generate multimodal semantically enhanced features; performing cross-modal attention fusion on the multimodal semantically enhanced features to obtain cross-modal fused features; and inputting the cross-modal fused features into a semantic inference network for processing to generate a decision result. This invention analyzes and filters the importance of features at each level of multimodal data, reduces interference from irrelevant information, and enhances the expressive power of key features by combining semantic enhancement and cross-modal attention mechanisms, achieving efficient fusion processing and ultimately improving the accuracy and robustness of decision-making in complex scenarios.

[0077] In one embodiment, step S10 includes:

[0078] S101, acquires visual data, language data, and motion data;

[0079] S102, extract the low-level edge features of the visual data through a convolutional neural network, extract the mid-level shape features of the visual data through a spatial attention mechanism, and extract the high-level semantic features of the visual data through a residual network;

[0080] S103, extract lexical features of the language data through a pre-trained language model, extract phrase features of the language data through a convolutional neural network, and extract sentence features of the language data through a bidirectional gated recurrent unit;

[0081] S104, extract the short-term dynamic features and medium-term trend features of the action data through a temporal convolutional network;

[0082] S105, extract the long-term pattern features of the action data through the gating loop unit;

[0083] S106, the bottom edge features, the middle shape features, the high semantic features, the vocabulary features, the phrase features, the sentence features, the short-term dynamic features, the medium-term trend features, and the long-term pattern features are integrated into the multimodal hierarchical initial feature set.

[0084] In this embodiment, acquiring visual data, language data, and motion data are fundamental steps in multimodal intelligent processing. Visual data can be acquired by obtaining continuous frame sequences through image sensors or video cameras. Language data can be acquired by obtaining speech streams through sound pickup devices and transcribing them into text using a speech recognition module. Motion data can be acquired by obtaining dynamic trajectory and attitude parameters based on inertial measurement units, position sensors, or external visual observations. This data acquisition process needs to ensure the synchronization and integrity of the data to guarantee the accuracy of subsequent feature extraction.

[0085] Layered feature extraction from visual data includes the extraction of low-level edge features, mid-level shape features, and high-level semantic features. Low-level edge features are extracted directly from the gradient and texture information within the pixel neighborhood using shallow convolutional kernels of a convolutional neural network, capturing local attributes such as object boundaries and texture details. Mid-level shape features are extracted by introducing a spatial attention mechanism, allowing the network to focus on structurally strong regions in the visual data, representing the object's contour, geometric structure, and spatial arrangement. High-level semantic features utilize stacked deep convolutional units in a residual network, combined with residual connections to improve information transfer capability and gradient stability, extracting abstract concepts such as category labels, attribute information, and contextual information contained in the image.

[0086] The hierarchical feature extraction process for language data first uses a pre-trained language model (e.g., BERT, GPT) to perform context-sensitive embedding representations of each word in the text data, forming lexical-level features that reflect the semantic meaning and contextual relationships of words. Then, a convolutional neural network is used to extract phrase-level features, characterizing word combination patterns and local grammatical structures through local convolution operations. Finally, bidirectional gated recurrent units are used to model the complete sentence, combining forward and backward contextual information to generate sentence-level features that capture the global semantics and logical relationships of the sentence.

[0087] Hierarchical feature extraction of motion data is categorized into three types: short-term dynamic features, medium-term trend features, and long-term pattern features. Short-term dynamic features are obtained by using temporal convolutional networks to perform convolution operations on motion sequences within a local time window, revealing detailed changes in motion. Medium-term trend features analyze motion trajectories of moderate duration using wider convolutional kernels or longer windows in temporal convolutional networks, reflecting trends and patterns during the motion process. Long-term pattern features employ gated recurrent units, relying on the recurrent unit memory mechanism to process long-term motion data and uncover long-term dependencies within motion sequences.

[0088] Finally, these features extracted through different structures and models—including low-level edge features, mid-level shape features, and high-level semantic features of visual data; lexical, phrase, and sentence-level features of language data; and short-term dynamic features, mid-term trend features, and long-term pattern features of action data—are integrated according to their source and hierarchical attributes to form a multimodal hierarchical initial feature set. The structure of this set reflects the completeness of information at different levels of abstraction for each modality of data, providing fine-grained and complementary foundational feature support for subsequent importance analysis, filtering, dimensionality reduction, and semantic enhancement.

[0089] This embodiment extracts and integrates fine-grained, hierarchical multimodal features into a unified multimodal hierarchical initial feature set. This effectively ensures that key information from each modality of visual, linguistic, and action data at different levels of abstraction is fully characterized and preserved. This processing method provides complete and reliable input for subsequent feature selection and semantic enhancement, enabling the subsequent fusion and decision-making stages to maximize the retention of high-value features while reducing redundant information. This improves the accuracy and robustness of overall multimodal information processing, ultimately supporting deep understanding and efficient decision-making across modal information in complex scenarios.

[0090] In one embodiment, step S20 above includes:

[0091] S201, perform edge response intensity analysis on the bottom edge features, middle shape features and high semantic features in the multimodal hierarchical initial feature set, and obtain the importance scores of the bottom edge features, the middle shape features and the high semantic features respectively;

[0092] S202, perform semantic combination rationality analysis on the lexical features, phrase features and sentence features in the multimodal hierarchical initial feature set, and obtain the importance scores of lexical features, phrase features and sentence features respectively;

[0093] S203, Perform task objective relevance analysis on the short-term dynamic features, medium-term trend features and long-term pattern features in the initial feature set of the multimodal hierarchical structure, and obtain the importance scores of the short-term dynamic features, medium-term trend features and long-term pattern features respectively;

[0094] S204, the bottom edge features, middle shape features, high semantic features, word-level features, phrase-level features, sentence-level features, short-term dynamic features, medium-term trend features, long-term pattern features and corresponding feature importance scores are concatenated to form importance enhancement features;

[0095] S205, The importance enhancement features are input into the filtering network and the filtering features are obtained through a self-attention mechanism;

[0096] S206, Perform matrix decomposition and dimensionality reduction on the selected features to obtain the multimodal hierarchical features after dimensionality reduction.

[0097] In this embodiment, the initial feature set of the multimodal hierarchical model contains features from multiple modalities and at various levels. Therefore, it is necessary to first analyze the importance of features at different levels within the set. Bottom-level edge features, mid-level shape features, and high-level semantic features correspond to information content at different levels of abstraction in the visual data. The edge response intensity analysis of bottom-level edge features assesses the prominence and stability of edge features in the image by calculating the average value, variance, or global statistics of the Laplacian operator response of pixel gradient magnitudes. The importance analysis of mid-level shape features focuses on evaluating the continuity and closure of structural contours, for example, by using the integral of curvature or spatial consistency scoring of the shape contour. The importance of high-level semantic features is measured by comparing their semantic similarity with predefined target labels or categories, using indices such as cosine similarity or KL divergence, thus obtaining an importance score for the high-level semantic features.

[0098] Lexical, phrase, and sentence-level features correspond to multi-granular semantic representations of language data. The importance analysis of lexical features can combine the term frequency-inverse document frequency (TF-IDF) metric or context-based attention weight distribution to measure the contribution of words to the overall semantics. The importance of phrase-level features can be scored through lexical integrity, dependency tightness, or frequency of occurrence within a sentence. The importance analysis of sentence-level features typically relies on global semantic plausibility evaluation, such as using the perplexity of a language model or the matching accuracy of alignment-annotated data.

[0099] Short-term dynamic features, medium-term trend features, and long-term pattern features are used to characterize behavioral information at different time scales in motion data. The importance of short-term dynamic features is analyzed through the relevance of rapid motion details in the task objective, such as the degree of alignment with predefined motion segment templates. The importance of medium-term trend features is measured through the velocity change patterns of continuous sequences and the smoothness of motion curves. The importance of long-term pattern features is based on the behavioral logic matching of the overall task, such as the correctness of the action sequence or the temporal consistency score.

[0100] The importance analysis results yielded importance scores for bottom-level edge features, mid-level shape features, high-level semantic features, lexical features, phrase-level features, sentence-level features, short-term dynamic features, medium-term trend features, and long-term pattern features. These scores, along with their corresponding feature data, were concatenated to form importance-enhanced features, enabling subsequent processing to utilize both the original feature information and their importance indices.

[0101] Importance-enhancing features are input into a filtering network for processing. This filtering network can be implemented using a deep neural network structure, such as a multi-layer fully connected network combined with a self-attention mechanism. It utilizes the weighted relationships of the self-attention mechanism to identify key parts of the global features, adaptively strengthening the focus on highly important features while weakening low-importance redundant features. The filtered features then enter the dimensionality reduction stage as the final filtering features.

[0102] Matrix factorization dimensionality reduction employs linear or nonlinear dimensionality reduction algorithms such as Singular Value Decomposition (SVD), Principal Component Analysis (PCA), or Nonnegative Matrix Factorization (NMF) to compress the selected feature matrix into a low-dimensional space while preserving maximum information. The dimensionality reduction result is the selected multimodal hierarchical feature, providing concise and effective input data for subsequent semantic enhancement and cross-modal fusion.

[0103] This embodiment performs independent importance analysis on features of different modalities and levels in the initial feature set of multimodal hierarchies, and introduces importance scores as additional input. By using a filtering network and a self-attention mechanism to enhance the focus on key information, and combining matrix factorization to reduce feature dimensions, it can significantly reduce redundant information, improve the compactness and discriminative power of feature expression, and provide more valuable input for subsequent semantic understanding and cross-modal alignment. This helps to improve the efficiency and accuracy of the overall system under complex multimodal tasks.

[0104] In one embodiment, step S30 above includes:

[0105] S301, Knowledge graph mapping is performed on the high-level semantic features in the multimodal hierarchical features after dimensionality reduction and screening to obtain enhanced high-level semantic features;

[0106] S302, the lexical features in the filtered and dimensionality-reduced multimodal hierarchical features are enhanced by knowledge base retrieval to obtain enhanced lexical features;

[0107] S303, Perform task scenario encoding enhancement on the long-term pattern features in the filtered and dimensionality-reduced multimodal hierarchical features to obtain task scenario enhanced features;

[0108] S304, perform kinematic constraint enhancement on the task scene enhancement features to obtain enhanced long-term pattern features;

[0109] S305, integrate the enhanced high-level semantic features, the enhanced lexical-level features, the enhanced long-term pattern features, and the bottom-level edge features, middle-level shape features, phrase-level features, sentence-level features, short-term dynamic features, and medium-term trend features from the filtered and reduced multimodal hierarchical features into multimodal semantic enhancement features.

[0110] In this embodiment, the multimodal hierarchical features after dimensionality reduction contain rich multi-level semantic and structural information. First, deep enhancement processing is required for the high-level semantic features, lexical features, and long-term pattern features. The high-level semantic features after dimensionality reduction are enhanced through knowledge graph mapping. By establishing mapping relationships with entities and relation nodes in a predefined semantic network, external knowledge background is introduced in the spatial dimension. This process calculates the similarity score between the high-level semantic features and nodes in the knowledge graph, selecting the upstream and downstream relationships of the most relevant nodes to expand features, enabling the enhanced high-level semantic features to capture information such as concept hierarchy, category attributes, and associated entities.

[0111] The lexical features after dimensionality reduction are enhanced through knowledge base retrieval. By utilizing a specially constructed domain lexical database, auxiliary semantic information such as definitions, attributes, and common collocations are added to the original lexical features through keyword matching, semantic nearest neighbor search, and contextual expansion. This enhances the semantic interpretation ability of words in context and is particularly suitable for scenarios with a wealth of professional terms and terminology.

[0112] The long-term pattern features, after dimensionality reduction, are enhanced through task scenario encoding. Task scenario encoding, based on a pre-defined behavioral template library, aligns long-term action patterns with specific task contexts, such as using sequence alignment algorithms or temporal template mapping mechanisms, enabling the long-term pattern features to express possible temporal evolution paths within the task. These enhanced long-term pattern features are further enhanced through kinematic constraints, introducing physical constraints on action execution, such as joint range of motion, velocity limits, and path continuity. Enhanced long-term pattern features are generated through rule mapping and condition filtering, ensuring that the implicit motion intentions within the features conform to realistic feasibility and physical consistency.

[0113] All enhanced features, including enhanced high-level semantic features, enhanced lexical features, and enhanced long-term pattern features, as well as low-level edge features, mid-level shape features, phrase-level features, sentence-level features, short-term dynamic features, and medium-term trend features that were not semantically enhanced but were retained after dimensionality reduction, are combined into multimodal semantic enhancement features through integration operations. This forms a complete feature set containing rich semantic context, task constraints, and multi-level information, laying the foundation for subsequent cross-modal fusion.

[0114] This embodiment employs differentiated semantic enhancement processing on different types of features within the multimodal hierarchical features after dimensionality reduction and selection. By embedding external knowledge graphs, domain knowledge bases, task behavior templates, and kinematic rules, each feature retains its original expressive power while gaining stronger semantic interpretability and task relevance. Through multi-dimensional supplementation and enhancement, the semantic alignment accuracy and cross-modal consistency of subsequent feature fusion can be improved, reducing ambiguity and semantic shift, thereby enhancing the accuracy and robustness of the overall decision-making results.

[0115] In one embodiment, step S40 above includes:

[0116] S401, extract enhanced high-level semantic features as visual semantic enhancement features, extract enhanced lexical-level features as language semantic enhancement features, and extract enhanced long-term pattern features as action semantic enhancement features from the multimodal semantic enhancement features;

[0117] S402, the visual semantic enhancement feature is used as the query vector, the language semantic enhancement feature is used as the first key vector and the first value vector, and the action semantic enhancement feature is used as the second key vector and the second value vector;

[0118] S403, determine the visual language attention weights based on the query vector and the first key vector;

[0119] S404, determine visual language fusion features based on the visual language attention weights and the first value vector;

[0120] S405, determine the visual action attention weights based on the query vector and the second key vector;

[0121] S406, determine visual action fusion features based on the visual action attention weights and the second value vector;

[0122] S407, Determine language action attention weights based on the language semantic enhancement features and the action semantic enhancement features;

[0123] S408, determine the language action fusion feature based on the language action attention weight and the second value vector;

[0124] S409, the visual-language fusion feature, the visual-action fusion feature, and the language-action fusion feature are weighted and summed to obtain the cross-modal fusion feature.

[0125] In this embodiment, cross-modal attention fusion, when processing multimodal semantic enhancement features, first requires rigorous extraction of various key features. Multimodal semantic enhancement features contain multi-level information from different sources. It is necessary to clearly distinguish and extract features that enhance high-level semantics, lexical-level features, and long-term pattern features, which are respectively designated as visual semantic enhancement features, linguistic semantic enhancement features, and action semantic enhancement features. This process requires feature channel parsing of the original input tensor to ensure that features from different sources are not confused and that dimensional alignment is maintained. The visual semantic enhancement feature serves as the query vector, which will guide the dominant attention direction in subsequent operations. The linguistic semantic enhancement feature serves as the first key vector and first value vector, and the action semantic enhancement feature serves as the second key vector and second value vector. It is crucial to ensure that they are strictly aligned in the tensor representation to satisfy the conditions for matrix multiplication.

[0126] Subsequently, the query vector and the first key vector are subjected to a dot product similarity calculation, followed by scaling and softmax normalization to form a visual-language attention weight matrix. This matrix measures the similarity distribution between visual and linguistic features at specific locations. The weighting operation of this matrix on the first value vector directly outputs the visual-language fusion feature. In the visual-action attention path, the query vector and the second key vector are similarly subjected to a dot product, scaling, and softmax normalization to obtain a visual-action attention weight matrix. This matrix is ​​then used to weight the second value vector to generate the visual-action fusion feature. In the linguistic-action path, the attention weights between the linguistic semantic enhancement features and the action semantic enhancement features are calculated to form a linguistic-action attention weight matrix. This matrix is ​​then used to weight the second value vector to obtain the linguistic-action fusion feature.

[0127] Finally, the visual-language fusion features, visual-action fusion features, and language-action fusion features are weighted and summed according to weight factors to form cross-modal fusion features. The weight factors can be determined based on the model parameter training results or dynamically adjusted according to external rules in a specific task. This summation operation must ensure strict dimensionality consistency among the three fusion features to avoid summation failures caused by inconsistent tensor shapes, ensuring that the cross-modal fusion features can fully express the comprehensive information of the three fusion paths. The entire processing flow should be implemented within an efficient matrix operation framework to guarantee the real-time performance and computational stability of multimodal high-dimensional data during the fusion process.

[0128] In the process of cross-modal attention fusion of multimodal semantic enhancement features, a multi-head attention mechanism can be used instead of single-head attention computation. This involves setting multiple attention heads in the visual-language, visual-action, and language-action paths, with each attention head processing similarity representations in different subspaces in parallel. The results from these attention heads are then concatenated and integrated through a linear transformation to form a richer cross-modal feature representation. This approach significantly improves the fine-grained representation of the associations between visual details, action patterns, and linguistic descriptions, and is particularly suitable for complex correspondences between local visual regions and phrases, action descriptions, etc.

[0129] Furthermore, cross-attention gating units can be introduced during cross-modal attention fusion, allowing the fusion strength of visual language, visual actions, and language action paths to be dynamically adjusted through learnable gating factors. This enables adaptive adjustment of the dependence on each modal input based on different task contexts during the feature fusion stage, effectively suppressing cross-modal noise interference and improving the robustness of fused features. This is particularly beneficial for processing local low-quality regions and redundant embellished language in medical image scenes, reducing decision bias caused by unnecessary feature interference.

[0130] Alternatively, a shared query vector mechanism can be employed, using visual semantic enhancement features as a shared global query vector, and performing attention matching with language semantic enhancement features and action semantic enhancement features respectively. This approach allows for focused enhancement of the visually dominant semantic components within different modalities during the fusion phase, unifying the alignment goals across different modalities. This is particularly suitable for applications requiring a vision-driven approach, such as multimodal fusion when agents perform autonomous decision-making in vision-driven tasks.

[0131] This embodiment employs attention-based cross-modal fusion to establish dynamic, fine-grained semantic alignment channels among vision, language, and action. It utilizes a weighted summation method to eliminate information inconsistencies and redundant interference between individual modalities. It adaptively focuses on highly correlated local regions among multimodal features, highlighting the most valuable feature combinations for decision-making tasks. This improves the model's accuracy and robustness in handling complex multimodal scenarios and helps address the insufficient intermodal interaction representation capabilities of traditional models.

[0132] In one embodiment, step S50 above includes:

[0133] S501, The cross-modal fusion features are input into the encoder for processing to obtain encoded features;

[0134] S502, The encoded features are input into the decoder for processing to obtain the decoded features;

[0135] S503, For motion control tasks, the decoded features are mapped to the joint angle space to generate joint angle prediction values;

[0136] S504, For a multimodal question answering task, the decoded features are mapped to a word probability distribution to generate a word probability distribution;

[0137] S505, generate motion control commands based on the predicted joint angle values, or generate natural language responses based on the word probability distribution.

[0138] In this embodiment, the process of processing cross-modal fusion features into a semantic reasoning network first requires constructing an encoder structure that can efficiently accept multi-dimensional fusion features. This encoder can parse the unified vector representation formed after the fusion of different modalities. The encoder employs a multi-layer self-attention module, with each layer consisting of a self-attention sub-layer and a feedforward sub-layer. The former is used to calculate the dependencies between elements in the cross-modal fusion features and capture the contextual relevance of information from different modalities, while the latter is used to enhance the abstract expressive power of the fusion features through nonlinear mapping. The input dimension of the encoder needs to be strictly aligned with the output dimension of the cross-modal fusion features to ensure dimensional consistency and avoid feature distortion. The multi-head attention units of each encoder layer independently calculate the attention weight matrix, and the results are recombined through concatenation and linear transformation to enhance the model's ability to represent the global semantics and fine-grained information of the fusion features. Through multi-layer stacking, the output encoded features can possess strong contextual expressive power and abstract semantic representation ability while maintaining multi-modal semantic consistency.

[0139] When encoded features are passed to the decoder, they first serve as the key and value inputs to the decoder's cross-attention module. The query vector is typically generated from task-related cues or historical decision context encoding representations, guiding the decoder to dynamically filter encoded features for the current task objective. The decoder's internal cross-attention module calculates the relevance weights between the current task objective and the encoded features, and aggregates the encoded features using these weights to generate a feature representation closely related to the task context. The cross-attention module works in conjunction with the decoder's self-attention module, enabling the decoder to comprehensively utilize current task context information and global encoded features during the decoding process, ensuring global semantic consistency between the generated results and the input features.

[0140] For motion control tasks, the decoded features output by the decoder are mapped to the joint angle space through a linear transformation. The dimension of the mapping matrix strictly corresponds to the number of joints, ensuring that each dimension of the output vector represents the predicted angle value of the corresponding joint of the agent. The mapping result is adjusted for range using a non-linear activation function to limit the predicted angle values ​​to the physically executable range. During training, the model is trained by backpropagation using the error between the predicted angle and the actual executed angle as the loss function, thereby gradually optimizing the weight parameters of the mapping matrix and improving the accuracy of motion prediction.

[0141] For multimodal question answering tasks, the decoded features output by the decoder are multiplied by the word embedding matrix through the output layer, mapping the features to a vector representation in the vocabulary space. The output is then converted into a probability distribution using a softmax function, where each dimension of the probability distribution represents the probability value of the corresponding word in the vocabulary as part of the output. This process ensures that the decoded features can be correctly mapped to the language output sequence under multimodal context constraints, while also supporting decoding strategies based on maximum probability or probability sampling.

[0142] Ultimately, based on the different branches of the current task, the predicted joint angles of the output joints in the action control task are parsed to form action control instructions, guiding the execution of physical actions by the agent; the probability distribution of the output words in the multimodal question answering task is gradually sampled through a sequence generation algorithm to generate complete natural language responses, ensuring semantic coherence and contextual consistency. The entire process emphasizes operational logic such as input-output consistency, cross-module dimensional alignment, feature space mapping, and dynamic context adaptation in its technical implementation, ensuring a balance between multi-task sharing and high-precision adaptation in the processing chain.

[0143] This embodiment achieves hierarchical processing and task adaptation mapping of cross-modal fusion features through encoder and decoder modules. It can make full use of the context and temporal dependencies of multimodal semantic information, ensuring that action control tasks and multimodal question answering tasks can share the same feature input and be independently mapped to their respective task spaces. While improving task robustness and accuracy, it reduces the model's dedicated design requirements for different tasks, and achieves effective adaptation of unified input and multi-task output.

[0144] In one embodiment, after step S50 above, the method further includes:

[0145] S601 collects joint angle execution error data based on the execution results of motion control commands;

[0146] S602, based on the execution results of natural language responses, collect user satisfaction evaluation data;

[0147] S603, collects task completion index data based on task completion status;

[0148] S604, Determine the loss function based on the joint angle execution error data, user satisfaction evaluation data, task completion index data, and the multimodal semantic enhancement features;

[0149] S605, determine the parameter gradients of the feature extraction network, the dimensionality reduction network, the semantic enhancement network, the cross-modal fusion network, and the semantic reasoning network through the loss function and the backpropagation algorithm;

[0150] S606, Update the parameters of the feature extraction network, the filtering and dimensionality reduction network, the semantic enhancement network, the cross-modal fusion network, and the semantic reasoning network based on the parameter gradient.

[0151] In this embodiment, after the cross-modal fusion feature input semantic reasoning network processes the data and generates decision results, it is necessary to immediately collect execution feedback on the decision results to optimize system performance. First, when the agent performs an action control task, the action execution monitoring module acquires real-time feedback information on the physical execution of the action control commands. It measures the actual joint angle values ​​and compares them joint-by-joint with the predicted joint angle values, calculates the angle deviation of each joint, and summarizes the data to form joint angle execution error data. This data serves as a direct quantitative feedback on the accuracy of action control, has a clear numerical definition, and is convenient for subsequent use in constructing the loss function.

[0152] For multimodal question-answering tasks, the system collects user satisfaction evaluation data after receiving natural language responses through the user interface or question-answering conversation logs. Satisfaction can be comprehensively derived through user ratings, positive and negative feedback tags, or implicit sentiment analysis models based on the dialogue history context, serving as an important indicator of the output quality of the question-answering task. User satisfaction evaluation data reflects the subjective acceptance of the language output results from the user's perspective.

[0153] Meanwhile, the system also collects task completion status data for all task types. Task completion index data can be obtained by judging the achievement conditions of predefined task objectives, such as whether the action of the action control task has reached the specified position, whether the question and answer task has accurately covered all expected information points, etc. The system comprehensively calculates and generates task completion index to quantify the overall execution effect of the task.

[0154] After collecting the above three types of feedback data, multimodal semantic enhancement features need to be introduced into the feedback optimization process as an important contextual basis. The multimodal semantic enhancement features obtained here are all the hierarchical features included in the decision generation, including bottom edge features, mid-level shape features, phrase-level features, sentence-level features, short-term dynamic features, and mid-term trend features, to ensure that the contextual relationships of the complete feature space are considered in the feedback optimization.

[0155] The loss function is constructed based on the combined effects of joint angle execution error data, user satisfaction evaluation data, task completion index data, and multimodal semantic enhancement features. The specific process includes:

[0156] First, joint angle execution error data is defined as the motion accuracy loss term, measuring the average absolute error between the actual and expected angles of each joint during robot motion execution, reflecting the precision of motion control. This loss is achieved by collecting global joint sequences, expanding them by time frames, and using a weighted average to reduce the impact of abnormal frames on the overall error. Second, user satisfaction evaluation data is mapped as a question-and-answer quality loss term. Standardized scores (e.g., normalized between 0 and 1) are used as a supervisory signal for language interaction performance, constraining the language generation submodule to optimize its response accuracy and user intent relevance. Third, task completion index data is transformed into a task achievement loss term, representing the degree to which the agent achieves its goal in a specific task context. By defining binary or multi-level labels for task completion (e.g., whether a specified position was reached, whether a correct response was generated), this is transformed into cross-entropy loss, ensuring direct constraints on the overall task goal. When constructing the loss function, adjustable weight coefficients λ1, λ2, and λ3 are introduced for these three types of loss, respectively. Their relative importance is set through empirical adjustment or task importance analysis to ensure the training balance among multiple tasks and prevent one task from having too strong a dominance over the overall optimization process and affecting the performance of other tasks.

[0157] While integrating the losses from the three task categories mentioned above, a multimodal semantic enhancement feature is introduced as a regularization term into the loss function. Specifically, the distribution distance between the multimodal semantic enhancement feature and its fused output feature is recorded at each training iteration. For example, Kullback-Leibler divergence or mean squared error is used as a distance metric to ensure that the model retains complete multimodal contextual information during parameter updates. This allows the model to maintain global consistency and rationality of its output even when task feedback is incomplete, utilizing the existing multimodal context. Finally, all loss terms are combined through weighted summation to form a single optimization objective function, which is formally represented as a linear superposition of action accuracy loss, question-answering quality loss, task achievement loss, and the multimodal regularization term. This optimization objective serves as the gradient source for backpropagation after each training round, driving the joint parameter updates of the feature extraction network, dimensionality reduction network, semantic enhancement network, cross-modal fusion network, and semantic reasoning network. This achieves overall optimization of the three objectives (action, language, and task) and the preservation of multimodal context, ensuring that the trained agent can make accurate, fast, and robust multimodal-driven autonomous decisions in complex environments.

[0158] By employing a loss function and backpropagation algorithm, the system can automatically calculate the parameter gradients of the feature extraction network, the dimensionality reduction network, the semantic enhancement network, the cross-modal fusion network, and the semantic reasoning network. Specifically, during the calculation process, the weight parameters of each network and the partial derivatives of the loss function with respect to the parameters are propagated back layer by layer, ensuring that the parameter update direction and magnitude of each layer are correct and reasonable, thus gradually reducing the value of the loss function.

[0159] Finally, based on the calculated parameter gradients, the system synchronously updates all parameters of the feature extraction network, the dimensionality reduction network, the semantic enhancement network, the cross-modal fusion network, and the semantic reasoning network. This enables each network to possess higher-precision multimodal representation and reasoning capabilities under the guidance of the latest round of feedback data, continuously improving the accuracy of action execution, the precision of language output, and the reliability of overall task completion. The entire process demonstrates a strict logical closed loop, fully utilizing feedback information and combining it with the global multimodal context to ensure the adaptive dynamic optimization of network parameters.

[0160] The feature extraction network plays a crucial role in the overall process, performing multi-dimensional and multi-level decomposition and encoding of raw visual, linguistic, and action data. It serves as the starting point for the entire multimodal information processing. Through convolutional and recurrent structures of varying depths, this network transforms raw data into low-level edge features, mid-level shape features, high-level semantic features, and lexical, phrase, sentence, and multi-granular action features suitable for subsequent processing. This directly determines the quality and completeness of the initial feature set for multimodal hierarchical processing. Updating the network's parameters can improve the encoding accuracy of raw inputs of different data types, reduce redundancy and noise in the feature extraction stage, and make subsequent feature importance analysis and selection more targeted and separable.

[0161] The filtering and dimensionality reduction network operates in the importance analysis and filtering stage of the initial feature set for multimodal hierarchical transformation. Its main function is to perform effective feature filtering and compression based on importance scores obtained from edge response strength analysis, semantic combination rationality analysis, and task-objective relevance analysis, generating information-condensed, dimensionality-reduced multimodal hierarchical features. Updating the network's parameters can adjust the filtering sensitivity and dimensionality reduction accuracy during the self-attention mechanism and matrix factorization dimensionality reduction process, further ensuring that the output dimensionality-reduced multimodal hierarchical features maintain semantic integrity while possessing optimal information density, providing efficient input for subsequent semantic enhancement.

[0162] Semantic augmentation networks are responsible for introducing external knowledge or contextual encoding at a high level of multimodal hierarchical features, enabling the augmented high-level semantic features, lexical features, and long-term pattern features to possess richer semantic relevance and ensuring that hidden relationships between data are enhanced. This network performs multi-dimensional augmentation through knowledge graph mapping, knowledge base retrieval, and task scenario encoding, combined with kinematic constraints. Updating the network's parameters can improve the relevance and adaptability of the knowledge embedding process, optimize the contextual consistency and semantic expressive power of multimodal semantic augmentation features, and strengthen the foundation for subsequent cross-modal alignment.

[0163] The cross-modal fusion network, after outputting multimodal semantic enhancement features, handles the interaction and coupling between visual, linguistic, and action-based information. By designing attention weights for visual-linguistic, visual-action, and linguistic-action modes, fusion features are extracted separately and then weighted and integrated into a single cross-modal fusion feature. This provides a unified and tightly coupled input representation for the semantic reasoning network. Updating the network's parameters helps to more accurately adjust the attention weight distribution, enabling the network to better align the primary and secondary relationships of different modalities within the context, ensuring optimal interaction between data from different sources within the cross-modal fusion feature.

[0164] As the final stage of decision output, the semantic reasoning network receives cross-modal fusion features and performs comprehensive analysis and reasoning within a unified semantic space, adapting to both action control and multimodal question answering tasks. Through an encoder-decoder structure, it progressively abstracts the fusion features and maps them to joint angle predictions or lexical probability distributions. The semantic reasoning network's role is to transform multimodal fusion features into executable and interactive results. Updating the network's parameters improves the rigor and task adaptability of the reasoning path, reduces action control bias, enhances the fluency and accuracy of question answer generation, and makes the decision results more aligned with user needs and task objectives. Each network update is interconnected with the preceding steps, ensuring that the overall system's input, processing, and output continuously adapt and optimize through multiple rounds of feedback.

[0165] Example Description: In an intelligent service robot scenario, the robot performs indoor object handling tasks, and its autonomous decision-making ability relies on multimodal perception and understanding. First, the robot acquires visual data of the current environment through a high-definition camera, uses a microphone array to acquire language commands from the human user, and obtains its own motion data through inertial and joint sensors. All acquired data enters a multi-channel parallel processing flow. A convolutional neural network extracts low-level edge features from the visual data to perceive the contours of objects in the scene; a spatial attention mechanism extracts mid-level shape features to identify object categories and geometric shapes; and a residual network extracts high-level semantic features to capture complex visual semantics. For language data, the robot uses a pre-trained language model to extract lexical-level features, a convolutional neural network to extract phrase-level features, and a bidirectional gated recurrent unit to extract sentence-level features, comprehensively analyzing the multi-layered semantics of the language input. For motion data, the robot uses a temporal convolutional network to extract short-term dynamic features and mid-term trend features to capture the continuity and stages of the current action, while a gated recurrent unit extracts long-term pattern features to understand task intent and global plan. These features are integrated to form a multimodal hierarchical initial feature set.

[0166] Next, the robot performs importance analysis on features at each level. In the visual part, edge response intensity analysis calculates the importance scores of bottom-level edges, mid-level shapes, and high-level semantic features, emphasizing the perception priority of objects related to the current task target. In the linguistic part, semantic combination rationality analysis ensures the accuracy of key descriptive words and phrases within the context. In the action part, task target relevance analysis highlights movement patterns describing the relationship between the robot and the task target. By concatenating each feature with its importance score, importance-enhanced features are generated and input into a filtering network. After further weight adjustment via a self-attention mechanism, matrix factorization is performed for dimensionality reduction, resulting in the filtered and dimensionality-reduced multimodal hierarchical features.

[0167] Building upon this foundation, the robot undergoes semantic enhancement. High-level semantic features are mapped and associated with indoor objects and task rules through a knowledge graph, improving the understanding of complex object relationships. Lexical-level features enrich the contextual meaning of command words through knowledge base retrieval, improving language understanding accuracy. Long-term pattern features, combined with task scenario encoding and kinematic constraints, calibrate the robot's movement intentions and physical feasibility. Finally, these enhanced key features are integrated with unenhanced but useful low-level edge, mid-level shape, phrase-level, sentence-level, short-term dynamic, and mid-term trend features to generate multimodal semantic enhancement features.

[0168] The robot performs cross-modal attention fusion on multimodal semantic enhancement features. Visual semantic enhancement features are extracted from the high-level semantic enhancement features as the query vector, lexical enhancement features are used as the first key-value pair, and long-term pattern enhancement features are used as the second key-value pair. The robot first calculates the attention weights for vision and language to obtain visual-language fusion features; then it calculates the attention weights for vision and action to obtain visual-action fusion features; finally, it calculates the attention weights for language and action to obtain language-action fusion features. The three fusion features are weighted and summed to form the cross-modal fusion feature, which is used for subsequent decision-making.

[0169] The robot inputs cross-modal fused features into a semantic reasoning network for processing. An encoder generates coded features, and a decoder generates decoded features. If the current task is a motion control task (e.g., moving an object to a designated location), the robot maps the decoded features to joint angle space, generates predicted joint angles, and uses these to form motion control commands to drive itself to complete the moving action. If the task is a multimodal question-answering task (e.g., a user asking for the location of an item indoors), the robot maps the decoded features to a lexical probability distribution, generates a natural language response, and accurately answers the user's question.

[0170] After decision-making and execution, the robot records feedback data: for example, it records joint angle execution errors after the action is completed, receives user satisfaction evaluations after a verbal response, and collects task completion metrics after the overall task is completed. The robot acquires multimodal semantic enhancement features as contextual references and uses the feedback data and multimodal semantic enhancement features together to calculate a multi-task loss function. Using this loss function, the robot calculates the parameter gradients of the feature extraction network, the dimensionality reduction network, the semantic enhancement network, the cross-modal fusion network, and the semantic reasoning network through the backpropagation algorithm, and then updates the parameters of these networks one by one. Through multiple rounds of updates, the robot continuously optimizes the entire chain from multimodal perception to autonomous decision-making, gradually improving perception accuracy, understanding ability, and task decision-making execution performance.

[0171] In complex and dynamic environments, intelligent agents, through a complete process of multimodal data acquisition, hierarchical processing, filtering and optimization, semantic enhancement, cross-modal fusion, semantic reasoning and decision-making, and adaptive optimization based on execution feedback, enable robots to exhibit high intelligence and robustness in real-world applications.

[0172] In the healthcare field, intelligent assisted diagnostic systems need to acquire and integrate information from multiple data sources to assist doctors in making comprehensive assessments and decisions regarding patients. First, the system collects the patient's visual data, including images of skin lesions and imaging scans (such as CT or MRI), as well as verbal interaction data between the patient and medical staff, such as recorded complaints and medical record Q&A content. Simultaneously, it collects the patient's movement data during the examination, including gait, posture, and range of motion information. Convolutional neural networks are used to extract features from the visual data, separating edge texture features to detect lesion contours, mid-level shape features to extract lesion area structure, and high-level semantic features to characterize complex pathological patterns. For the verbal data, a pre-trained language model is used to extract lexical features of key medical terms, convolutional neural networks are used to extract phrase-level features to analyze the logic of symptom descriptions, and bidirectional gated recurrent units are used to extract sentence-level features to understand complete complaints. Motion data is processed through temporal convolutional networks to obtain short-term dynamic features and mid-term trend features for analyzing movement disorder manifestations, while gated recurrent units are used to extract long-term movement pattern features to identify chronic motor function deterioration trends. All these features are integrated to form a multimodal hierarchical initial feature set.

[0173] Next, the system performs importance analysis on features at each level of the initial feature set. For visual features, importance scores are calculated for bottom-level edge features, mid-level shape features, and high-level semantic features through edge response intensity analysis, emphasizing the criticality of features such as tumor boundary clarity, shape regularity, and multi-organ relationships in images. For linguistic features, the relevance of medical expressions between words, phrases, and sentences is assessed based on the rationality of semantic combination. For action features, task-goal relevance analysis is used to determine the connection between patient movement abnormalities and the diagnostic task. All features and their importance scores are concatenated to form importance-enhanced features, which are then input into the screening network. The screening network uses a self-attention mechanism to assign dynamic weights to different features to strengthen key dimensions. Subsequently, matrix factorization is performed on the screening results to reduce noise and improve the efficiency of subsequent processing, resulting in multimodal hierarchical features after dimensionality reduction.

[0174] Then, the system performs semantic enhancement on these filtered and dimensionality-reduced features. For visual features, knowledge graph mapping is used to associate high-level semantic features of imaging with existing standardized pathological patterns in the medical imaging database, such as the standard morphology of skin lesions. For linguistic features, knowledge base retrieval enhancement is used to verify the correspondence between patient descriptions and standard medical terminology; for example, mapping "shortness of breath" to respiratory disease entries. For action features, task scenario encoding is used to parse movement patterns in the context of rehabilitation therapy, and kinematic constraints further enhance the rationality of long-term pattern features. Finally, the enhanced high-level visual, lexical-level linguistic, and long-term action features are integrated with other unenhanced hierarchical features (bottom-level edge features, mid-level shape features, phrase-level features, sentence-level features, short-term dynamic features, and medium-term trend features) into multimodal semantic enhancement features.

[0175] Next, the system performs cross-modal attention fusion processing, extracting enhanced high-level semantic features as visual semantic enhancement features, enhanced lexical-level features as linguistic semantic enhancement features, and enhanced long-term pattern features as action semantic enhancement features. The visual semantic enhancement features are used as query vectors, forming first and second key-value pairs with the linguistic and action semantic enhancement features, respectively. Visual-linguistic attention weights and visual-action attention weights are calculated, and visual-linguistic and visual-action fusion features are calculated based on the weights and corresponding value vectors. Simultaneously, linguistic-action attention weights are calculated based on the linguistic and action semantic enhancement features to derive the linguistic-action fusion features. The system then weights and sums the three sets of fusion features to obtain the cross-modal fusion features.

[0176] Cross-modal fusion features are input into a semantic reasoning network for decision processing. The network first performs contextual semantic compression on the fusion features and extracts latent representations through an encoder, then a decoder transforms the encoded features into a decision output adapted to the specific task. For different task scenarios, if the system determines it's a rehabilitation movement guidance task, it maps the decoded features to joint angle space, generating predicted joint angle values ​​for the patient's rehabilitation movements, such as suggestions for knee flexion and extension range. If the system determines it's an intelligent question-and-answer scenario, it maps the decoded features to a lexical probability distribution, generating natural language responses to the patient's questions, such as "Your pain symptoms suggest further MRI examination." Finally, it generates movement control commands or natural language responses as intelligent auxiliary outputs.

[0177] After the decision results are output, the system initiates a feedback collection and model adaptive update mechanism. It tracks and analyzes joint angle execution errors by monitoring patient rehabilitation movements, collects patient or doctor satisfaction scores for natural language responses, and statistically analyzes task completion status (e.g., whether examinations were completed as suggested). Combining this data with multimodal semantic enhancement features, a weighted multi-task loss function is constructed. This function incorporates action accuracy error, question-answer quality and user satisfaction, and task completion as multi-objective optimization metrics. Simultaneously, the distribution of multimodal semantic enhancement features is used as a regularization term to restrict the global consistency among text, image, and action data during training. This loss function, in conjunction with the backpropagation algorithm, calculates parameter gradients to update the parameters of the feature extraction network, dimensionality reduction network, semantic enhancement network, cross-modal fusion network, and semantic reasoning network. This improves the overall assisted diagnostic system's adaptability and decision accuracy to complex multimodal data in medical scenarios.

[0178] The medical and health intelligent agent can efficiently process multimodal information such as patients' vision, language, and movement, dynamically adjust the diagnostic model to adapt to patient feedback, and ultimately achieve a higher level of autonomous decision support in various medical tasks such as assisted diagnosis, rehabilitation guidance, and personalized question answering.

[0179] In the fintech field, intelligent customer service systems need to understand customer intent in real time, analyze customer behavior, and output diversified decision-making suggestions to improve service quality and business efficiency. First, the system collects visual data (such as images of invoices and ID documents uploaded by customers), linguistic data (such as voice or text conversations with customer service personnel), and motion data (such as gestures, swipes, or clicks on the customer's mobile application). For visual data, convolutional neural networks extract edge features from invoices to identify layout, mid-level shape features to parse security features such as seals and watermarks, and high-level semantic features to extract key fields such as amount and bank name. For linguistic data, a pre-trained language model extracts lexical features from customer statements, convolutional neural networks extract phrase-level features to identify phrases such as "apply for a credit card" or "account frozen," and bidirectional gated recurrent units extract sentence-level features to understand complex complaints or business requests. For motion data, temporal convolutional networks extract short-term dynamic features and mid-term trend features to identify customer operating habits and abnormal behaviors, and gated recurrent units extract long-term pattern features to analyze historical operation paths. All these features are integrated to form a multimodal hierarchical initial feature set.

[0180] Next, the system analyzes the importance of features at each level within the set. Visual features are calculated using edge response intensity analysis to determine the key parts of customer-submitted images, such as the clarity of ID photos or the completeness of seals on documents. Linguistic features are analyzed for semantic combination rationality to determine the fit between user expressions and standard business terminology. Action features are analyzed for task objective relevance to determine the degree of association between click trajectories, swipe behaviors, and the current business scenario (such as high-risk operations or sensitive transactions). All features are concatenated with importance scores to form importance-enhanced features, which are then input into the filtering network. The filtering network dynamically assigns weights to features of each dimension using a self-attention mechanism, reinforcing key elements relevant to the current financial business. After matrix factorization and dimensionality reduction, the filtered multimodal hierarchical features are obtained, reducing noise and improving processing efficiency.

[0181] Then, the system performs semantic enhancement on the filtered and dimensionality-reduced features. Visual high-level semantic features are enhanced by mapping to a standardized knowledge graph of financial industry bills and certificates, improving the ability to identify abnormal bills. Language-level features are enhanced through knowledge base retrieval, quickly locating sensitive business terms such as "transfer limit" or "frozen funds." Long-term action pattern features are enhanced through task scenario encoding, analyzing the operational sequence in financial risk control scenarios. Kinematic constraint enhancement ensures that action patterns conform to application operation specifications. Finally, the system integrates enhanced visual, language, and action semantic features with other unenhanced features to form multimodal semantically enhanced features.

[0182] Next, the system performs cross-modal attention fusion. Enhanced high-level semantic features are used as visual semantic enhancement features, enhanced lexical-level features as linguistic semantic enhancement features, and enhanced long-term pattern features as action semantic enhancement features. The visual semantic enhancement features are used as query vectors to form key-value pairs with the linguistic and action semantic enhancement features, respectively. Visual-linguistic attention weights and visual-action attention weights are calculated sequentially to generate corresponding fused features. Simultaneously, attention weights are also calculated between the linguistic and action semantic enhancement features to form a linguistic-action fused feature. The three sets of fused features are weighted and summed to output the cross-modal fused feature.

[0183] Subsequently, the cross-modal fusion feature input semantic reasoning network performs decision processing. The encoder compresses and models the input features, while the decoder performs reasoning output based on the business type. For action control tasks, such as automatically filling out and submitting customer authorization instructions, the decoder outputs the operation sequence corresponding to key information; for multimodal question-answering tasks, such as a customer inquiring about "details of the last five transactions," the decoder outputs a targeted natural language response. Finally, the system generates action control instructions (such as form auto-completion) or natural language responses (such as account query results).

[0184] After the decision results are output, the system collects execution feedback and dynamically updates the model. By analyzing the execution results of customer authorization instructions, the system collects operation success rate and error rate as analogies to joint angle execution error data, analyzes customer ratings of question-and-answer results as user satisfaction evaluation data, and statistically analyzes whether the task is completed in a closed loop as task completion index data. Combined with multimodal semantic enhancement features, these data are used to construct a multi-task loss function. This loss function weights and superimposes operation accuracy, question-and-answer quality, and task closure rate to form a unified optimization objective, and uses multimodal semantic enhancement features as a regularization term to constrain the training process to maintain global consistency of visual, linguistic, and action data, preventing overfitting to a single financial business scenario. By calculating parameter gradients through the loss function and backpropagation algorithm, the system dynamically updates the parameters of the feature extraction network, dimensionality reduction network, semantic enhancement network, cross-modal fusion network, and semantic reasoning network, continuously improving the system's multimodal understanding, interaction, and decision-making capabilities in the financial service process.

[0185] Intelligent customer service systems can efficiently process multimodal data such as invoice images, voice requests, and interactive behaviors submitted by customers, dynamically adapt to customer behavior and feedback, improve the quality of autonomous decision-making in multiple business scenarios such as financial consultation, risk control, automatic form filling, and account verification, and enhance customer experience and the security of financial business.

[0186] This embodiment significantly improves the overall performance of multimodal data processing and decision-making by having each of the aforementioned networks assume its own functional role and adaptively update through multiple rounds of feedback. First, the feature extraction network update refines and denoises the encoding of visual, linguistic, and action data, helping the filtering and dimensionality reduction network to more efficiently remove redundant and low-relevance information. Second, the filtering and dimensionality reduction network update further improves the balance between information density and expressiveness of input features, reducing unnecessary computational burden. The semantic enhancement network update improves the relevance and robustness of semantic context embedding, making the semantic information expression between different modalities more consistent. The cross-modal fusion network update improves the attention weight distribution during multimodal data interaction, effectively mitigating alignment bias between modalities. The semantic reasoning network update optimizes the reasoning path based on fused features, making the generation of action control commands and natural language responses more accurate and in line with task objectives. Through end-to-end parameter updates and adaptive adjustments, the system can ultimately generate decision results that meet semantic expectations efficiently and accurately even when faced with complex, diverse, and potentially incomplete or noisy multimodal inputs.

[0187] In one embodiment, a multimodal hierarchical feature fusion and decision-making apparatus is provided, which corresponds one-to-one with the multimodal hierarchical feature fusion and decision-making method described in the above embodiments. (Refer to...) Figure 3 , Figure 3This is a schematic diagram of the functional modules of a preferred embodiment of the multimodal hierarchical feature fusion and decision-making device of the present invention. The modules include a multimodal feature extraction module 10, a multimodal feature filtering module 20, a multimodal semantic enhancement module 30, a cross-modal fusion module 40, and a semantic reasoning module 50. Detailed descriptions of each functional module are as follows:

[0188] The multimodal feature extraction module 10 is used to acquire visual data, language data, and action data, and to perform hierarchical feature extraction on the visual data, language data, and action data to generate a multimodal hierarchical initial feature set.

[0189] The multimodal feature filtering module 20 is used to analyze the importance of each level feature in the multimodal hierarchical initial feature set, and to filter and reduce the multimodal hierarchical initial feature set based on the importance to obtain the filtered and reduced multimodal hierarchical features.

[0190] The multimodal semantic enhancement module 30 is used to perform semantic enhancement on the filtered and dimensionality-reduced multimodal hierarchical features to generate multimodal semantic enhancement features;

[0191] The cross-modal fusion module 40 is used to perform cross-modal attention fusion on the multimodal semantic enhancement features to obtain cross-modal fused features;

[0192] The semantic reasoning module 50 is used to input the cross-modal fusion features into the semantic reasoning network for processing and to generate decision results.

[0193] In one embodiment, the multimodal feature extraction module 10 is specifically used for:

[0194] Acquire visual data, language data, and motion data;

[0195] The visual data is extracted using a convolutional neural network to extract the low-level edge features, the visual data is extracted using a spatial attention mechanism to extract the mid-level shape features, and the visual data is extracted using a residual network to extract the high-level semantic features.

[0196] Lexical features of the language data are extracted by a pre-trained language model, phrase-level features of the language data are extracted by a convolutional neural network, and sentence-level features of the language data are extracted by a bidirectional gated recurrent unit.

[0197] Short-term dynamic features and medium-term trend features of the action data are extracted using a temporal convolutional network.

[0198] Long-term pattern features of the action data are extracted using a gating loop unit;

[0199] The bottom-level edge features, the middle-level shape features, the high-level semantic features, the lexical features, the phrase-level features, the sentence-level features, the short-term dynamic features, the medium-term trend features, and the long-term pattern features are integrated into the multimodal hierarchical initial feature set.

[0200] In one embodiment, the multimodal feature filtering module 20 is specifically used for:

[0201] Edge response intensity analysis is performed on the bottom edge features, middle shape features, and high semantic features in the multimodal hierarchical initial feature set to obtain the importance scores of the bottom edge features, the middle shape features, and the high semantic features, respectively.

[0202] A semantic combination rationality analysis is performed on the lexical features, phrase features, and sentence features in the multimodal hierarchical initial feature set to obtain the importance scores of lexical features, phrase features, and sentence features, respectively.

[0203] Task objective relevance analysis was performed on the short-term dynamic features, medium-term trend features, and long-term pattern features in the initial feature set of the multimodal hierarchical structure to obtain the importance scores of the short-term dynamic features, medium-term trend features, and long-term pattern features, respectively.

[0204] The bottom-level edge features, mid-level shape features, high-level semantic features, lexical features, phrase-level features, sentence-level features, short-term dynamic features, medium-term trend features, and long-term pattern features are concatenated with their corresponding feature importance scores to form an importance-enhanced feature.

[0205] The importance enhancement features are input into the filtering network and processed through a self-attention mechanism to obtain the filtering features;

[0206] The selected features are subjected to matrix factorization to reduce dimensionality, resulting in multimodal hierarchical features after dimensionality reduction.

[0207] In one embodiment, the multimodal semantic enhancement module 30 is specifically used for:

[0208] Knowledge graph mapping is applied to the high-level semantic features in the filtered and dimensionality-reduced multimodal hierarchical features to enhance them, resulting in enhanced high-level semantic features.

[0209] The lexical features in the filtered and dimensionality-reduced multimodal hierarchical features are enhanced by knowledge base retrieval to obtain enhanced lexical features;

[0210] Task scenario-enhanced features are obtained by performing task scenario encoding enhancement on the long-term pattern features in the filtered and dimensionality-reduced multimodal hierarchical features.

[0211] Kinematic constraints are applied to the enhanced features of the task scenario to obtain enhanced long-term pattern features;

[0212] The enhanced high-level semantic features, the enhanced lexical features, the enhanced long-term pattern features, and the bottom edge features, middle shape features, phrase-level features, sentence-level features, short-term dynamic features, and medium-term trend features from the filtered and dimensionality-reduced multimodal hierarchical features are integrated into multimodal semantic enhancement features.

[0213] In one embodiment, the cross-modal fusion module 40 is specifically used for:

[0214] From the multimodal semantic enhancement features, we extract enhanced high-level semantic features as visual semantic enhancement features, enhanced lexical-level features as language semantic enhancement features, and enhanced long-term pattern features as action semantic enhancement features;

[0215] The visual semantic enhancement features are used as the query vector, the language semantic enhancement features are used as the first key vector and the first value vector, and the action semantic enhancement features are used as the second key vector and the second value vector.

[0216] Visual language attention weights are determined based on the query vector and the first key vector;

[0217] Visual language fusion features are determined based on the visual language attention weights and the first value vector;

[0218] Visual action attention weights are determined based on the query vector and the second key vector;

[0219] Visual action fusion features are determined based on the visual action attention weights and the second value vector;

[0220] Based on the language semantic enhancement features and the action semantic enhancement features, determine the language action attention weights;

[0221] The language action fusion features are determined based on the language action attention weights and the second value vector;

[0222] The visual-language fusion features, the visual-action fusion features, and the language-action fusion features are weighted and summed to obtain cross-modal fusion features.

[0223] In one embodiment, the semantic reasoning module 50 is specifically used for:

[0224] The cross-modal fusion features are input into the encoder for processing to obtain coded features;

[0225] The encoded features are input into the decoder for processing to obtain the decoded features;

[0226] For motion control tasks, the decoded features are mapped to the joint angle space to generate joint angle prediction values;

[0227] For multimodal question answering tasks, the decoded features are mapped to a word probability distribution to generate a word probability distribution;

[0228] Action control commands are generated based on the predicted joint angle values, or natural language responses are generated based on the vocabulary probability distribution.

[0229] In one embodiment, the semantic reasoning module 50 is specifically used for:

[0230] Based on the execution results of motion control commands, collect joint angle execution error data;

[0231] Based on the execution results of natural language responses, collect user satisfaction evaluation data;

[0232] Based on the task completion status, collect task completion index data;

[0233] The loss function is determined based on the joint angle execution error data, user satisfaction evaluation data, task completion index data, and the multimodal semantic enhancement features.

[0234] The parameter gradients of the feature extraction network, the dimensionality reduction network, the semantic enhancement network, the cross-modal fusion network, and the semantic reasoning network are determined using the loss function and the backpropagation algorithm.

[0235] The parameters of the feature extraction network, the filtering and dimensionality reduction network, the semantic enhancement network, the cross-modal fusion network, and the semantic reasoning network are updated based on the parameter gradient.

[0236] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used for communication with external user terminals via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a multimodal hierarchical feature fusion and decision-making method on the server side.

[0237] In one embodiment, a computer device is provided, which may be a user terminal, and its internal structure diagram may be as follows: Figure 5As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the user-side functions or steps of a multimodal hierarchical feature fusion and decision-making method.

[0238] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:

[0239] Acquire visual data, language data, and motion data, and perform hierarchical feature extraction on the visual data, language data, and motion data to generate a multimodal hierarchical initial feature set;

[0240] The importance of each level feature in the initial multimodal hierarchical feature set is analyzed, and the initial multimodal hierarchical feature set is filtered and dimensionality reduced based on the importance to obtain the filtered and dimensionality-reduced multimodal hierarchical features;

[0241] The selected and dimensionality-reduced multimodal hierarchical features are semantically enhanced to generate multimodal semantically enhanced features;

[0242] Cross-modal attention fusion is performed on the multimodal semantic enhancement features to obtain cross-modal fused features;

[0243] The cross-modal fusion features are input into a semantic reasoning network for processing to generate decision results.

[0244] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0245] Acquire visual data, language data, and motion data, and perform hierarchical feature extraction on the visual data, language data, and motion data to generate a multimodal hierarchical initial feature set;

[0246] The importance of each level feature in the initial multimodal hierarchical feature set is analyzed, and the initial multimodal hierarchical feature set is filtered and dimensionality reduced based on the importance to obtain the filtered and dimensionality-reduced multimodal hierarchical features;

[0247] The selected and dimensionality-reduced multimodal hierarchical features are semantically enhanced to generate multimodal semantically enhanced features;

[0248] Cross-modal attention fusion is performed on the multimodal semantic enhancement features to obtain cross-modal fused features;

[0249] The cross-modal fusion features are input into a semantic reasoning network for processing to generate decision results.

[0250] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0251] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0252] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0253] It should be noted that if any software tools or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The embodiments described above are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A multimodal hierarchical feature fusion and decision-making method, characterized in that, Includes the following steps: Acquire visual data, language data, and motion data, and perform hierarchical feature extraction on the visual data, language data, and motion data to generate a multimodal hierarchical initial feature set; The importance of each level feature in the initial multimodal hierarchical feature set is analyzed, and the initial multimodal hierarchical feature set is filtered and dimensionality reduced based on the importance to obtain the filtered and dimensionality-reduced multimodal hierarchical features; The selected and dimensionality-reduced multimodal hierarchical features are semantically enhanced to generate multimodal semantically enhanced features; Cross-modal attention fusion is performed on the multimodal semantic enhancement features to obtain cross-modal fused features; The cross-modal fusion features are input into a semantic reasoning network for processing to generate decision results.

2. The multimodal hierarchical feature fusion and decision-making method as described in claim 1, characterized in that, Acquire visual data, language data, and motion data, and perform hierarchical feature extraction on the visual data, language data, and motion data to generate a multimodal hierarchical initial feature set, including: Acquire visual data, language data, and motion data; The visual data is extracted using a convolutional neural network to extract the low-level edge features, the visual data is extracted using a spatial attention mechanism to extract the mid-level shape features, and the visual data is extracted using a residual network to extract the high-level semantic features. Lexical features of the language data are extracted by a pre-trained language model, phrase-level features of the language data are extracted by a convolutional neural network, and sentence-level features of the language data are extracted by a bidirectional gated recurrent unit. Short-term dynamic features and medium-term trend features of the action data are extracted using a temporal convolutional network. Long-term pattern features of the action data are extracted using a gating loop unit; The bottom-level edge features, the middle-level shape features, the high-level semantic features, the lexical features, the phrase-level features, the sentence-level features, the short-term dynamic features, the medium-term trend features, and the long-term pattern features are integrated into the multimodal hierarchical initial feature set.

3. The multimodal hierarchical feature fusion and decision-making method as described in claim 1, characterized in that, The importance of features at each level in the initial multimodal hierarchical feature set is analyzed, and the initial multimodal hierarchical feature set is filtered and dimensionality reduced based on the importance to obtain the filtered and dimensionality-reduced multimodal hierarchical features, including: Edge response intensity analysis is performed on the bottom edge features, middle shape features, and high semantic features in the multimodal hierarchical initial feature set to obtain the importance scores of the bottom edge features, the middle shape features, and the high semantic features, respectively. A semantic combination rationality analysis is performed on the lexical features, phrase features, and sentence features in the multimodal hierarchical initial feature set to obtain the importance scores of lexical features, phrase features, and sentence features, respectively. Task objective relevance analysis was performed on the short-term dynamic features, medium-term trend features, and long-term pattern features in the initial feature set of the multimodal hierarchical structure to obtain the importance scores of the short-term dynamic features, medium-term trend features, and long-term pattern features, respectively. The bottom-level edge features, mid-level shape features, high-level semantic features, lexical features, phrase-level features, sentence-level features, short-term dynamic features, medium-term trend features, and long-term pattern features are concatenated with their corresponding feature importance scores to form an importance-enhanced feature. The importance enhancement features are input into the filtering network and processed through a self-attention mechanism to obtain the filtering features; The selected features are subjected to matrix factorization to reduce dimensionality, resulting in multimodal hierarchical features after dimensionality reduction.

4. The multimodal hierarchical feature fusion and decision-making method as described in claim 1, characterized in that, The selected and dimensionality-reduced multimodal hierarchical features are semantically enhanced to generate multimodal semantically enhanced features, including: Knowledge graph mapping is applied to the high-level semantic features in the filtered and dimensionality-reduced multimodal hierarchical features to enhance them, resulting in enhanced high-level semantic features. The lexical features in the filtered and dimensionality-reduced multimodal hierarchical features are enhanced by knowledge base retrieval to obtain enhanced lexical features; Task scenario-enhanced features are obtained by performing task scenario encoding enhancement on the long-term pattern features in the filtered and dimensionality-reduced multimodal hierarchical features. Kinematic constraints are applied to the enhanced features of the task scenario to obtain enhanced long-term pattern features; The enhanced high-level semantic features, the enhanced lexical features, the enhanced long-term pattern features, and the bottom edge features, middle shape features, phrase-level features, sentence-level features, short-term dynamic features, and medium-term trend features from the filtered and dimensionality-reduced multimodal hierarchical features are integrated into multimodal semantic enhancement features.

5. The multimodal hierarchical feature fusion and decision-making method as described in claim 1, characterized in that, Cross-modal attention fusion is performed on the multimodal semantic enhancement features to obtain cross-modal fused features, including: From the multimodal semantic enhancement features, we extract enhanced high-level semantic features as visual semantic enhancement features, enhanced lexical-level features as language semantic enhancement features, and enhanced long-term pattern features as action semantic enhancement features; The visual semantic enhancement features are used as the query vector, the language semantic enhancement features are used as the first key vector and the first value vector, and the action semantic enhancement features are used as the second key vector and the second value vector. Visual language attention weights are determined based on the query vector and the first key vector; Visual language fusion features are determined based on the visual language attention weights and the first value vector; Visual action attention weights are determined based on the query vector and the second key vector; Visual action fusion features are determined based on the visual action attention weights and the second value vector; Based on the language semantic enhancement features and the action semantic enhancement features, determine the language action attention weights; The language action fusion features are determined based on the language action attention weights and the second value vector; The visual-language fusion features, the visual-action fusion features, and the language-action fusion features are weighted and summed to obtain cross-modal fusion features.

6. The multimodal hierarchical feature fusion and decision-making method as described in claim 1, characterized in that, The cross-modal fusion features are input into a semantic reasoning network for processing to generate decision results, including: The cross-modal fusion features are input into the encoder for processing to obtain coded features; The encoded features are input into the decoder for processing to obtain the decoded features; For motion control tasks, the decoded features are mapped to the joint angle space to generate joint angle prediction values; For multimodal question answering tasks, the decoded features are mapped to a word probability distribution to generate a word probability distribution; Action control commands are generated based on the predicted joint angle values, or natural language responses are generated based on the vocabulary probability distribution.

7. The multimodal hierarchical feature fusion and decision-making method as described in claim 1, characterized in that, After the cross-modal fusion features are input into the semantic reasoning network for processing and a decision result is generated, the process further includes: Based on the execution results of motion control commands, collect joint angle execution error data; Based on the execution results of natural language responses, collect user satisfaction evaluation data; Based on the task completion status, collect task completion index data; The loss function is determined based on the joint angle execution error data, user satisfaction evaluation data, task completion index data, and the multimodal semantic enhancement features. The parameter gradients of the feature extraction network, the dimensionality reduction network, the semantic enhancement network, the cross-modal fusion network, and the semantic reasoning network are determined using the loss function and the backpropagation algorithm. The parameters of the feature extraction network, the filtering and dimensionality reduction network, the semantic enhancement network, the cross-modal fusion network, and the semantic reasoning network are updated based on the parameter gradient.

8. A multimodal hierarchical feature fusion and decision-making device, characterized in that, The multimodal hierarchical feature fusion and decision-making device includes: The multimodal feature extraction module is used to acquire visual data, language data, and action data, and to perform hierarchical feature extraction on the visual data, language data, and action data to generate a multimodal hierarchical initial feature set; The multimodal feature filtering module is used to analyze the importance of each level feature in the multimodal hierarchical initial feature set, and to filter and reduce the multimodal hierarchical feature set based on the importance to obtain the filtered and reduced multimodal hierarchical features. The multimodal semantic enhancement module is used to semantically enhance the filtered and dimensionality-reduced multimodal hierarchical features to generate multimodal semantically enhanced features. A cross-modal fusion module is used to perform cross-modal attention fusion on the multimodal semantic enhancement features to obtain cross-modal fused features; The semantic reasoning module is used to input the cross-modal fusion features into the semantic reasoning network for processing and to generate decision results.

9. A computer device, characterized in that, The computer device includes a memory, a processor, and a multimodal hierarchical feature fusion and decision program stored in the memory and executable on the processor. When executed by the processor, the multimodal hierarchical feature fusion and decision program implements the steps of the multimodal hierarchical feature fusion and decision method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The storage medium stores a multimodal hierarchical feature fusion and decision program, which, when executed by a processor, implements the steps of the multimodal hierarchical feature fusion and decision method as described in any one of claims 1-7.

Citation Information

Cited By

  • Children brain health intelligent question and answer method and system fused with multi-modal data

    CN121525814A

  • Remote sensing visual question and answer method based on large language model and multi-level attention mechanism

    CN121542456A

  • Remote sensing visual question answering method based on large language model and multi-level attention mechanism

    CN121542456B

  • Ai follow-up visit data processing method and system based on cloud platform

    CN121562632A

  • Intelligent processing method and system for aviation logistics

    CN121563353A