Decision-making method and device based on multi-modal information, equipment and medium

By acquiring historical visual, linguistic, and action data, action decisions are generated and updated, solving the problem of inflexible decision-making in multimodal information interaction in traditional action decoders and achieving efficient and accurate action decisions in complex environments.

CN120952168APending Publication Date: 2025-11-14PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511060286.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Traditional action decoders struggle to dynamically combine visual features, linguistic features, and action history features in multimodal information interaction, resulting in inflexible and inaccurate action decisions in complex environments, impacting the automated service experience and operational security.

Method used

By acquiring visual data, language commands, and action history data, visual features, language features, and action history features are extracted respectively. These features are then fused to generate preprocessed multimodal features, which are then processed by a hierarchical action decoder to generate action sequence features. Action decisions are then updated in conjunction with environmental feedback information.

Benefits of technology

It improves the adaptability and accuracy of action decisions, enhances the model's ability to respond to changing scenarios, and ensures that action sequences can be adjusted and optimized in real time in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120952168A_ABST
    Figure CN120952168A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, can be applied to business scenes such as agent autonomous decision making, financial science and technology and medical health, and discloses a decision making method and device based on multi-modal information, equipment and a medium. Comprising the steps of obtaining visual data, language instructions and action historical data, processing the visual data, the language instructions and the action historical data into visual features, language features and action historical features, fusing the features to generate preprocessed multi-modal features, generating action sequence features by using a layered action decoder, mapping the action sequence features into control parameters to generate action decisions, and outputting the action decisions. Environment feedback information is collected, new action sequence features are generated based on the environment feedback information, and action decisions are updated. Through multi-modal information processing and hierarchical action decoding, vision, language and action historical information are dynamically fused, action sequence generation and decision updating are optimized in combination with environment feedback, the adaptability and accuracy of action decision in a complex environment are effectively improved, and the response ability of a model to variable scenes is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a decision-making method, apparatus, device, and storage medium based on multimodal information. Background Technology

[0002] In existing Vision-Language-Action (VLA) models, the action decoder is a key module responsible for transforming multimodal information into specific action decisions. However, traditional action decoders often employ simple fully connected layers or basic recurrent neural network structures, which generally suffer from simplistic decision-making logic and difficulty adapting to complex task requirements. When faced with multi-source heterogeneous data, traditional action decoders cannot deeply fuse and dynamically model visual, linguistic, and action history information, making them ill-equipped to handle the high uncertainty of the task environment.

[0003] In the fintech business, existing action decoders struggle to handle the complex relationships between multimodal data, including visual credentials, text instructions, and historical operation records. For example, in the business process of smart teller machines, traditional decoders cannot flexibly match standard operation sequences based on the user's natural language requests and real-time camera images, resulting in insufficient accuracy and intelligence in business decisions, impacting the automated service experience and business risk management capabilities.

[0004] In the healthcare field, existing action decoders also face similar bottlenecks. Faced with multimodal data input including medical images, clinical instructions, and patient history, traditional decoders lack the ability to deeply model the dynamic relationships between data, making it difficult to provide accurate and efficient action decision support for surgical robots or remote diagnostic systems. This deficiency makes it difficult for medical robots to adjust their action strategies in real time when performing complex tasks, affecting operational safety and medical quality.

[0005] In applications such as robotics and autonomous driving, the limitations of action decoders are even more pronounced. Existing technologies cannot fully utilize dynamic environmental changes and semantic instructions within visual scenes, resulting in inflexible action decisions in dynamic environments. For example, in multi-robot collaborative tasks, traditional action decoders cannot rationally plan collaborative paths between robots, easily leading to conflicts and inefficient action sequences, thus reducing overall task execution efficiency. Summary of the Invention

[0006] The main objective of this invention is to provide a decision-making method, apparatus, device, and storage medium based on multimodal information, aiming to solve the technical problem that traditional action decoders cannot dynamically combine visual features, language features, and action history features in multimodal information interaction to generate accurate action decisions that adapt to complex environmental changes.

[0007] To achieve the above objectives, the present invention provides a decision-making method based on multimodal information, comprising:

[0008] Acquire visual data, language commands, and action history data, and process the visual data, language commands, and action history data to obtain visual features, language features, and action history features, respectively;

[0009] By fusing the aforementioned visual features, linguistic features, and action history features, preprocessed multimodal features are obtained;

[0010] The preprocessed multimodal features are processed by a hierarchical action decoder to generate action sequence features;

[0011] The action sequence features are mapped to control parameters, and action decisions are generated based on the control parameters;

[0012] Collect environmental feedback information, generate new action sequence features based on the environmental feedback information, and update the action decision based on the new action sequence features.

[0013] Furthermore, to achieve the above objectives, the present invention provides a decision-making device based on multimodal information, comprising:

[0014] A multimodal perception module is used to acquire visual data, language commands, and action history data, and process the visual data, language commands, and action history data to obtain visual features, language features, and action history features, respectively.

[0015] A multimodal fusion module is used to fuse the visual features, language features, and action history features to obtain preprocessed multimodal features;

[0016] The action decoding module is used to process the preprocessed multimodal features through a hierarchical action decoder to generate action sequence features;

[0017] An action decision module is used to map the action sequence features to control parameters and generate action decisions based on the control parameters.

[0018] The feedback adjustment module is used to collect environmental feedback information, generate new action sequence features based on the environmental feedback information, and update the action decision based on the new action sequence features.

[0019] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and a decision-making program based on multimodal information stored in the memory and executable on the processor, wherein when the decision-making program based on multimodal information is executed by the processor, it implements the steps of the decision-making method based on multimodal information as described above.

[0020] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing a decision-making program based on multimodal information, wherein when the decision-making program based on multimodal information is executed by a processor, it implements the steps of the decision-making method based on multimodal information as described above.

[0021] Beneficial Effects: This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as intelligent agent autonomous decision-making, fintech, and healthcare. It discloses a decision-making method, apparatus, device, and medium based on multimodal information, including: acquiring visual data, language commands, and historical action data and processing them into visual features, language features, and historical action features; fusing visual features, language features, and historical action features to generate preprocessed multimodal features; using a hierarchical action decoder to process the preprocessed multimodal features to generate action sequence features; mapping the action sequence features to control parameters and generating action decisions; collecting environmental feedback information and generating new action sequence features based on the environmental feedback information; and updating the action decisions. This invention dynamically fuses visual, language, and historical action information through multimodal information processing and hierarchical action decoding, and optimizes action sequence generation and decision updates by combining environmental feedback. This effectively improves the adaptability and accuracy of action decisions in complex environments and enhances the model's responsiveness to changing scenarios. Attached Figure Description

[0022] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings:

[0023] Figure 1 This is a schematic diagram of an application environment for a decision-making method based on multimodal information according to an embodiment of the present invention;

[0024] Figure 2 This is a flowchart illustrating an embodiment of the decision-making method based on multimodal information of the present invention;

[0025] Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the decision-making device based on multimodal information of the present invention;

[0026] Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention;

[0027] Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0028] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0029] The decision-making method based on multimodal information provided in this invention can be applied to, for example... Figure 1 In this application environment, the user terminal communicates with the server via a network. The server can acquire visual data, language commands, and action history data from the user terminal and process them into visual features, language features, and action history features. It then fuses these features to generate preprocessed multimodal features, uses a hierarchical action decoder to process these features to generate action sequence features, maps these features to control parameters to generate action decisions, collects environmental feedback information, generates new action sequence features based on this feedback, and updates the action decisions. This invention dynamically fuses visual, language, and action history information through multimodal information processing and hierarchical action decoding, combining environmental feedback to optimize action sequence generation and decision updates. This effectively improves the adaptability and accuracy of action decisions in complex environments and enhances the model's responsiveness to changing scenarios. The user terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.

[0030] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the decision-making method based on multimodal information provided by the present invention. It should be noted that although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0031] like Figure 2 As shown, the decision-making method based on multimodal information proposed in this invention includes the following steps:

[0032] S10, acquire visual data, language commands and action history data, and process the visual data, language commands and action history data to obtain visual features, language features and action history features respectively;

[0033] In this embodiment, the processing of visual data, language commands, and action history data is a fundamental operation for multimodal information input. Firstly, acquiring visual data involves collecting image data from an image acquisition device. This image data can originate from an RGB camera, depth camera, or multispectral imaging equipment. The acquired data is stored in the form of a digital image matrix, typically exhibiting a multi-channel pixel intensity value distribution. Acquiring language commands involves obtaining language content from the user interface through speech recognition or text input. The input format can be a voice signal or a text string, requiring the use of a speech-to-text module or direct text input from a predefined interface. Acquiring action history data involves reading logs of previously executed action sequences, typically saved in a time-series format, containing action control commands from previous moments.

[0034] For visual data processing, spatial texture features and shape contour information are first extracted layer by layer through a multi-layer convolutional neural network structure. This process involves operations such as spatial filtering, nonlinear activation, and pooling. In specific implementation, convolutional kernels with different receptive fields can be used to capture local features at different scales. After obtaining preliminary spatial features, multi-scale fusion processing is then performed to form visual features. At this stage, the visual features are tensor representations after feature compression and normalization, possessing spatial hierarchical distribution information. Language instruction processing includes first segmenting the entire text into word or phrase sequences using word segmentation algorithms. Word segmentation algorithms come from natural language processing toolkits such as NLTK and spaCy. Subsequently, part-of-speech tagging is performed on the words or phrases, and the tagging results serve as additional information for subsequent semantic encoding. Part-of-speech tagging is implemented using statistical models or deep learning models, such as conditional random fields or bidirectional LSTM models. The word segmentation sequences and part-of-speech tagging information are encoded into word vectors, which can be obtained through pre-trained word vector models such as GloVe or Word2Vec. The dimension of the word vectors is determined according to the model requirements, such as 300-dimensional or 768-dimensional. The semantic information supplementation of word vectors needs to be combined with external knowledge bases such as WordNet and ConceptNet. Semantic association features are added during the encoding process to enrich the contextual understanding ability of each word or phrase. Finally, the multi-head attention mechanism is used to model the dependency relationship between different words in a global scope. The multi-head attention mechanism comes from the Transformer structure and its role is to improve the representation ability of long texts and complex semantic relationships.

[0035] Processing historical motion data requires extracting sequence change patterns from time-series data. First, the historical motion data is standardized to enhance the consistency of data distribution across all dimensions and avoid bias introduced by scale differences. Then, time-series dependency modeling networks such as Long Short-Term Memory (LSTM) and Gated Recurrent Units (GRUs) are used to encode temporal dependencies. Time-step unrolling operations capture long-term dependencies in the sequence, obtaining feature representations in the time dimension. These historical motion features are output as sequence vectors or hidden-state tensors to express the dynamic change patterns of past motion sequences.

[0036] The order in which the three features are obtained should ensure that spatial feature extraction of visual data is completed first, followed by word segmentation, part-of-speech tagging, and encoding of language instructions. Finally, temporal feature extraction of action history data is performed using a time-series modeling network, ensuring that the features are fully expressed in terms of time, space, and semantics. Strict type and dimension alignment requirements must be maintained between the inputs and outputs of each step to ensure that the features can be used for subsequent fusion processing.

[0037] For visual data processing, convolutional neural network structures of varying depths can be selected, such as ResNet, DenseNet, or a lightweight MobileNet. The kernel size can be adjusted to adapt to the resolution and detail of different visual data. For example, in medical scenarios, deeper networks are used to extract fine-grained structures when processing high-resolution medical images, while lightweight networks can be used to reduce computation when processing invoices or scanned documents in financial scenarios. For language instruction processing, different word segmentation and part-of-speech tagging models can be selected. For example, a BERT-based Chinese word segmenter can be used in Chinese scenarios, while an NLTK-based segmenter can be used in English scenarios. Word vectors can be static or context-driven, such as ELMo or BERT word vectors, depending on the task requirements. When supplementing with semantic information, the source and scope of the knowledge base can be adjusted. In the healthcare field, a medical terminology knowledge base can be used, while in the financial field, a financial terminology ontology can be used. For action history data processing, the number of hidden units and the time step length of the time series modeling network can be adjusted to adapt to the length and complexity of different historical action sequences. For example, high-frequency, long-time-series data can be adapted for medical equipment operation history analysis, while low-frequency, short-series data can be adapted for financial transaction operation history analysis.

[0038] Example: In the healthcare business, high-resolution real-time images of the surgical site are collected from visual sensors in the operating room as visual data, surgical step instructions are input by the doctor's voice as language commands, and the previous operation sequences of the surgical robot are combined as action history data. After the above processing, a multimodal feature input that can be used by intelligent assisted surgical robots is formed.

[0039] In the fintech business, customer identity documents and images are collected as visual data during counter or online services, customer service requests or interaction logs are used as verbal instructions, and historical transaction operation sequences are used as action history data. After processing, these data are provided to financial service robots to improve the accuracy and security of automated service decisions.

[0040] In the field of autonomous decision-making by intelligent agents, home service robots collect indoor environmental images as visual data, user service requests input by voice or terminal as language commands, and the robot's own historical movement paths and task completion records as action history data. After processing, multimodal features are formed to guide the robot to autonomously plan movement paths and perform task actions in complex home environments.

[0041] This embodiment, through the joint collection and targeted processing of visual data, language commands, and historical action data, enables multimodal input to express external environment, user needs, and historical behavior information more completely and precisely, forming a complementary and highly correlated feature set. This provides high-quality data support for subsequent fusion and decision-making, thereby effectively improving the accuracy and stability of action decisions generated in a changing environment.

[0042] S20, the visual features, language features and action history features are fused to obtain preprocessed multimodal features;

[0043] In this embodiment, the process of fusing visual features, linguistic features, and action history features aims to integrate heterogeneous data from different modalities to form a unified representation that can be used for subsequent processing. In specific implementation, the dimensions of visual features, linguistic features, and action history features first need to be standardized and aligned to ensure they can be combined in the same computational space. Visual features are typically high-dimensional spatial tensors representing the spatial semantic structure of an image, derived from the intermediate or terminal outputs of convolutional neural networks. Linguistic features are context embedding vectors, derived from the multi-head attention output of a language encoding model, containing semantic dependencies. Action history features are time-series encoded hidden states, derived from the modeling output of historical action sequences by recurrent neural networks. Specific methods for dimension alignment include linear mapping using fully connected layers or extending single-dimensional vectors to the target tensor dimension through dimension expansion operations, or compressing the high-dimensional tensor to make it consistent with other modal features.

[0044] After dimensional alignment, weighted enhancement operations are performed on each modal feature. A spatial attention weight matrix is ​​applied to visual features, which adaptively calculates the saliency of different spatial locations within the visual features using convolutional kernels, enhancing the representation of key visual regions. Semantic association weighting is applied to linguistic features, using a relevance matrix calculated from the global context to adjust the semantic contribution of linguistic vectors. Temporal association weighting is applied to action history features, using dynamic importance weights at each time step in the sequence to adjust the expressive power of the hidden states. The three weighted features are concatenated along either the channel or feature dimension to form a high-dimensional fused feature representation. To control feature dimensionality, the fused feature can be compressed using linear projection or 1x1 convolution to obtain the final preprocessed multimodal features. This preprocessed multimodal feature provides a unified, structured data representation containing spatial, semantic, and temporal information, providing sufficient contextual information support for subsequent model processing.

[0045] In the healthcare field, the dimension alignment stage can use fully connected layers of a specific size to map endoscopic image features, surgical voice command embeddings, and the hidden states of robot action sequences to a unified dimension. In the weighted enhancement operation, saliency maps of key anatomical regions in medical images can be used to guide spatial attention, and the importance weights of specialized terms in language features can be calculated using a medical terminology database. Furthermore, the weights of historical action sequences can be adjusted based on the surgical task context.

[0046] In the fintech business field, the dimensional consistency of the convolutional output of the bill image, the embedding of user transaction voice requests, and the embedding of historical transaction behavior sequences can be adjusted through a fully connected layer, and weighted enhancement can be achieved by combining the salience of bill field positions, the financial semantic relevance matrix, and the time weight of high-frequency trading windows.

[0047] In scenarios where intelligent agents make autonomous decisions, service robots map environmental image features acquired by visual sensors, embedded service language requests input by users, and hidden states of historical task execution records to a unified space. Then, they weight and concatenate these features through an adaptive attention mechanism to form a multimodal unified feature representation.

[0048] This embodiment integrates multi-modal preprocessed features formed by fusing visual features, linguistic features, and action history features. This integrates multi-source heterogeneous data and significantly improves the model's ability to jointly express environmental states, task requirements, and historical behaviors. This enables the model to more accurately perceive the current environment and combine it with historical context information, thereby improving the accuracy and robustness of subsequent action decisions. In particular, it has stronger adaptability and generalization ability in complex environments, multi-task constraints, or highly dynamic changing scenarios.

[0049] S30, the preprocessed multimodal features are processed by a hierarchical action decoder to generate action sequence features;

[0050] In this embodiment, the process of generating action sequence features by processing preprocessed multimodal features through a hierarchical action decoder aims to parse the fused multimodal data layer by layer and transform it into a high-dimensional time-series representation that can be used for subsequent action decisions. The preprocessed multimodal features contain three types of information: visual, linguistic, and action history. Their spatial, semantic, and temporal characteristics need to be effectively preserved and encoded in an orderly manner. First, the semantic understanding layer of the hierarchical action decoder receives the preprocessed multimodal feature input and maps it as a high-dimensional tensor to the semantic embedding space. This mapping can be achieved through a multi-head self-attention mechanism combined with global contextual relationship calculation. The self-attention mechanism obtains global dependencies by calculating the autocorrelation matrix of the input tensor and weights different semantic units, outputting a serialized semantic representation. The semantic embedding output by the semantic understanding layer not only includes the joint multimodal representation but also clarifies the contextual relationships and semantic orientations between different modalities.

[0051] Subsequently, the dynamic programming layer constructs a graph structure representation of the output of the semantic understanding layer. Graph nodes represent time-step states or context units in a multimodal sequence, and node attributes are initialized by semantic embedding vectors. The connections between nodes can be determined by the initial environment model, domain knowledge rules, or data-driven learning results. The dynamic programming layer operates on the graph structure, dynamically updating node states and edge connections by combining current environment state data and task objective descriptions. The dynamic update process includes updating node embedding vectors (e.g., aggregating neighbor node information through a graph convolutional neural network) and edge weights (e.g., adjusting edge weights based on similarity or importance between nodes), ensuring that the graph structure optimizes as the environment changes and the objective is adjusted. The updated graph structure is used to run optimal path planning algorithms, such as dynamic programming or graph search algorithms, within this layer to generate action planning features. Action planning features are not merely single action outputs, but rather a planned trajectory representation that includes temporal sequence and action state sequence.

[0052] Finally, the action generation layer takes action planning features and original action history features as input, fusing contextual information from historical trajectories and the current planned path. The action generation layer can employ sequence models such as Gated Recurrent Units (GRU) or Long Short-Term Memory (LSTM) networks to capture temporal dependencies and predict the next action. Through recursive computation with time-step expansion, it outputs action sequence features, which serve as an encoded representation of the complete temporal context, containing rich spatial, semantic, and temporal interaction information, providing high-quality input for downstream modules to map control parameters.

[0053] In the field of healthcare, a multi-head self-attention mechanism can be used as the main implementation method of semantic understanding layer in surgical robot systems. By combining the embedded features of intraoperative visual input, doctor's language instructions and historical action trajectories, a graph structure model including surgical target, key anatomical sites and intraoperative instrument status can be constructed. The dynamic programming layer updates graph nodes by combining surgical stage and current tissue status, and the action generation layer predicts operation sequence through LSTM network to control the path of surgical instruments.

[0054] In the fintech business, the Transformer encoder can be used as a semantic understanding layer to perform global semantic context parsing of multimodal features of bill image features, customer voice request embeddings, and historical transaction trajectories. Combined with the customer's real-time transaction context and predefined risk control rules, a transaction graph structure can be constructed. The dynamic planning layer performs risk path planning and node adjustment, and the action generation layer predicts and recommends transaction behavior sequences through time series models.

[0055] In the field of autonomous decision-making for intelligent agents, a lightweight multi-head self-attention mechanism can be used as the semantic understanding layer to analyze the visual images of the user's home environment, the embedding of voice commands, and the features of historical movement trajectories. Combined with the real-time home layout map, a spatial task planning graph structure can be constructed. The dynamic planning layer dynamically adjusts the node state and path, and the action generation layer uses a temporal convolutional network (TCN) to predict the movement path and interaction behavior sequence of the home robot.

[0056] This embodiment utilizes a hierarchical action decoder to process preprocessed multimodal features, enabling deep semantic parsing, dynamic environment modeling, and temporal correlation fusion of visual, linguistic, and action history features. This allows for the full encoding and optimization of multimodal information across spatial, semantic, and temporal dimensions, improving the accuracy and adaptability of action sequence features. Compared to traditional decoders that rely solely on a single fully connected network or basic recurrent neural network, the hierarchical action decoder effectively extracts key action sequence information under dynamic environments and diverse instruction inputs. This generates high-quality action sequence features that adapt to complex environmental changes, providing accurate and robust input for subsequent generation of control parameters and action decisions.

[0057] S40, map the action sequence features to control parameters, and generate action decisions based on the control parameters;

[0058] In this embodiment, mapping action sequence features to control parameters involves linearly mapping the generated action sequence feature vector from a high-dimensional space to a low-dimensional space. This is typically achieved through a fully connected mapping network, converting the multidimensional feature distribution into a continuous set of parameters usable for physical control. During this mapping process, the input action sequence features are sequentially input into the linear mapping matrix, weighted and summed with matrix weights and a bias term added, resulting in a continuous numerical vector. The mapped numerical vector is then processed by a nonlinear activation function, such as a sigmoid or tanh function with range constraints, to ensure that the output range conforms to the parameter range of the control execution unit, such as joint angles, movement speeds, or steering angles. The obtained initial control parameters need further verification of their physical feasibility. Based on hardware execution constraints and safety rules, the values ​​are checked for exceeding limits or being unreasonable, forming a control parameter verification result. When the verification result shows that the parameters violate regulations, a correction vector based on a preset physical boundary is calculated. This correction vector is used to adjust the control parameters exceeding the boundary to the executable range. The adjustment method can be simple pruning, smooth mapping, or incremental gradient projection. The corrected intermediate control parameters are verified by the dynamics simulation module. Simulation analysis examines whether the expected motion corresponding to the parameters poses a collision risk, exhibits abnormal motion stability, or is unreachable. When the simulation verification passes, the parameters are directly output as the final adjusted control parameters. When the simulation verification fails, a parameter correction increment is generated based on the simulation feedback and superimposed on the initial correction vector to regenerate new intermediate control parameters. This process repeats until the parameters meet the executable conditions. Finally, based on the adjusted control parameters and according to a preset mapping relationship, continuous motion control commands are generated. These commands may include actuator drive commands, joint command sequences, steering commands, and speed setpoints, completing the output of the motion decision. This process ensures that the motion sequence features generated from multimodal information undergo rigorous physical constraint checks and dynamic feasibility verification, generating motion decision commands that meet the execution requirements.

[0059] Different parameter mapping techniques can be used in different implementations. For example, deep fully connected neural networks can be used instead of single-layer fully connected layers to improve the nonlinear fitting capability of feature-to-control parameter mapping through multi-layer structures. Different types of nonlinear activation functions can also be used for range constraints. For instance, in medical robotic arm control, the tanh function is used to adapt the joint angle range; in the execution unit of a financial robot, the sigmoid function is used to adapt the control execution signal ratio. Physical feasibility verification rules can also be adjusted according to the scenario. In intelligent rehabilitation robots, verification rules can be set based on the range of human motion; in autonomous driving, verification rules can be adjusted based on the vehicle dynamics model. The dynamics simulation module can also select different simulation engines according to specific needs. For example, rigid body dynamics simulation can be used for robot joint control, and vehicle kinematics simulation can be used for autonomous driving system control.

[0060] This embodiment maps action sequence features to control parameters and combines physical constraint verification with dynamic simulation verification. This ensures that the output control parameters not only meet the task objectives but also satisfy hardware executability and dynamic environment adaptability, effectively avoiding execution failures or safety risks caused by abnormal parameters.

[0061] S50: Collect environmental feedback information, generate new action sequence features based on the environmental feedback information, and update the action decision based on the new action sequence features.

[0062] In this embodiment, the operation of collecting environmental feedback information involves multi-dimensional perception of the real-time state of the environment, mainly executed through external sensor modules. These sensor modules may include LiDAR, depth cameras, inertial measurement units (IMUs), environmental audio pickup devices, etc., to acquire environmental state change data. The data may be in the form of spatial position coordinates, velocity changes, attitude information, or external constraints. The collected environmental state data is input into the hierarchical action decoder structure through the data access module, specifically into the dynamic planning layer, for dynamically adjusting the current graph structure representation. Within the dynamic planning layer, the node attributes of the graph structure are updated in conjunction with the environmental feedback information. For example, the current spatial coordinates of position nodes and the real-time reachability of state nodes are updated, and edge connection weights are corrected, such as adjusting path cost, traversability, or flow priority. The updated graph structure representation reflects the dynamic constraints and spatial relationships of the current environment, becoming the basis for new action planning. Subsequently, the action planning layer performs optimized path planning based on the updated graph structure, generating new action planning features. These new features reflect the adaptive adjustment of the action path after environmental changes. The action generation layer takes new action planning features as input, performs time-series decoding operations, and outputs new action sequence features. These features encode a temporal action template that adapts the corrected action path to the current environment. Finally, the new action sequence features are converted into adjusted continuous action control commands through a control parameter mapping process consistent with the previous one, updating the action decision results and ensuring that the action decisions reflect changes in the current environment in real time. This entire process guarantees that action decisions are always synchronized with the environmental state, dynamically optimized, and possess environmental adaptability and real-time adjustment capabilities.

[0063] In different implementations, different types of sensor combinations can be configured according to the complexity of the environment. For example, in the operating environment of a medical robot, a 3D vision sensor and a near-field tactile sensor are used together to collect environmental change data to capture changes in patient position or dynamic constraints. In an autonomous driving financial service terminal, LiDAR and visual camera data are combined to collect pedestrian position and traffic sign information. The dynamic programming layer can select different graph optimization algorithms to update node states and edge weights. For example, Dijkstra's algorithm can be used to update the shortest feasible path, or the A* algorithm can be used to introduce heuristic acceleration. In the action generation layer, different sequence generation networks can be used according to the real-time requirements of the task. For example, in a financial robot terminal with strict latency requirements, a lightweight recurrent neural network is used to quickly generate action sequences, while in a high-precision medical robotic arm, a bidirectional recurrent neural network combined with a multi-head attention mechanism is used to generate fine-grained action sequences. Action decision updates can support frame-by-frame updates or batch updates, and can be flexibly adjusted according to the needs of different application scenarios.

[0064] This embodiment collects environmental feedback information and updates node states and edge connection weights in real time at the dynamic programming layer. This enables the action sequence features to adapt to environmental changes in real time, thereby updating action decisions. This allows the system to maintain stable and reliable task execution in dynamic and complex environments, improves the environmental adaptability and decision accuracy of multimodal decision-making in complex applications, and enhances the system's operational flexibility, efficiency, and security in dynamic environments in fields such as healthcare, financial services, and autonomous decision-making by intelligent agents.

[0065] This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as autonomous decision-making by intelligent agents, fintech, and healthcare. It discloses a decision-making method, apparatus, device, and medium based on multimodal information, comprising: acquiring visual data, language commands, and historical action data and processing them into visual features, language features, and historical action features; fusing the visual features, language features, and historical action features to generate preprocessed multimodal features; using a hierarchical action decoder to process the preprocessed multimodal features to generate action sequence features; mapping the action sequence features to control parameters and generating action decisions; collecting environmental feedback information and generating new action sequence features based on the environmental feedback information; and updating the action decisions. This invention dynamically fuses visual, language, and historical action information through multimodal information processing and hierarchical action decoding, and optimizes action sequence generation and decision updates by combining environmental feedback, effectively improving the adaptability and accuracy of action decisions in complex environments and enhancing the model's responsiveness to changing scenarios.

[0066] In one embodiment, step S10 above includes:

[0067] S101, acquire visual data and extract multi-level spatial features from the visual data;

[0068] S102, Perform multi-scale pooling operation on the multi-level spatial features to obtain spatial features of uniform size;

[0069] S103, apply a channel spatial attention mechanism to the spatial features of the uniform size to obtain visual features;

[0070] S104, Obtain language instructions, perform word segmentation and part-of-speech tagging on the language instructions, and generate word segmentation sequences and part-of-speech tagging information respectively;

[0071] S105, the word segmentation sequence and part-of-speech tagging information are encoded into word vectors;

[0072] S106, use a knowledge base to supplement the semantic information of the word vectors to obtain enhanced word vectors;

[0073] S107, Apply a multi-head attention mechanism to the enhanced word vectors to obtain language features;

[0074] S108, acquire historical motion data, and extract the temporal variation characteristics of the historical motion data;

[0075] S109, Apply a temporal dependency modeling network to process the temporal change features to establish long-term dependencies and obtain action history features.

[0076] In this embodiment, acquiring visual data, language commands, and historical action data requires multimodal input acquisition capabilities. Visual data can be acquired through high-resolution RGB cameras, depth cameras, structured light sensors, or multispectral imaging devices to ensure rich spatial and textural information coverage. After visual data input, a feature extraction network is used to extract multi-level spatial features from the input image or frame sequence. Typically, multi-layer convolutional operations of convolutional neural networks are used to construct pyramid-shaped feature maps to obtain local details and global context. After multi-level spatial features are generated, they are processed through multi-scale pooling operations, using a combination of max pooling and average pooling to reduce the dimensionality of the feature maps within different scale windows. This aligns the feature maps at each level in spatial dimensions, outputting spatial features of uniform size, ensuring that features of different resolutions can be directly aligned and integrated.

[0077] The uniformly sized spatial features are then input into the channel spatial attention module. This module calculates weights based on the response values ​​of different channels and spatial locations, dynamically adjusting the feature importance of each region and channel in the space. This enhances the feature representation ability of task-related visual regions, suppresses irrelevant interference information, and outputs visual features. Language command acquisition can be achieved through user input, text-to-speech transcription, or an external task scheduling system. The input language data is first processed by a word segmentation tool, dividing long texts into word units and using grammatical analysis tools to label parts of speech, generating word segments and part-of-speech tagging information. Subsequently, the word segments and part-of-speech tagging information are mapped to a high-dimensional continuous vector space through a word vector encoding network, forming a semantic embedding representation.

[0078] The word vector processing further integrates with an external knowledge base. By matching the association between word entities and domain knowledge, it supplements contextual semantic information and prior knowledge to form enhanced word vectors. These enhanced word vectors are then processed by a multi-head attention mechanism, utilizing parallel computation of attention distribution weights across multiple subspaces to capture long-distance dependencies, improving the comprehensive expressive ability of semantic details and global structure, and outputting linguistic features. When acquiring historical action data, temporal change features need to be extracted from historical task execution records. These records may include action sequence, timestamps, speed, acceleration, and external environmental states. For the extracted temporal change data, a temporal dependency modeling network is used for sequence modeling. For example, models such as Long Short-Term Memory (LSTM) and Gated Recurrent Unit (GRU) are used to capture short-term and long-term dependencies in the time series, enhancing the learning and expression of complex historical motion patterns, and ultimately outputting historical action features. Each process maintains strict correlation in data flow and logic, ensuring high consistency and usability in the generation of visual features, linguistic features, and historical action features, laying a solid foundation for subsequent feature fusion.

[0079] This embodiment enhances the expressive power of visual information for key areas of the spatial environment by extracting multi-level spatial features from visual data and performing multi-scale pooling and attention weighting. Combined with word segmentation, part-of-speech tagging, word vector encoding, and knowledge base supplementation of language instructions, it improves the semantic understanding accuracy of language data. Furthermore, it utilizes a multi-head attention mechanism to improve long-distance semantic relationship modeling and strengthen the accuracy of language features. For action history data, through temporal change feature extraction and temporal dependency network processing, it captures the dynamic patterns and long-term dependency information of historical motion behavior, comprehensively improving the feature quality and consistency of multimodal input data. This enables subsequent action decision-making processes to fully integrate rich visual, linguistic, and temporal information, achieving accurate adaptation and efficient response to task scenarios in dynamic and complex environments.

[0080] In one embodiment, step S20 above includes:

[0081] S201, perform feature alignment processing on the visual features, language features and action history features to obtain dimension-aligned visual features, dimension-aligned language features and dimension-aligned action history features;

[0082] S202, Apply spatial attention weighting to the dimension-aligned visual features to obtain weighted visual features;

[0083] S203, apply semantic association weighting to the dimension-aligned language features to obtain weighted language features;

[0084] S204, apply temporal correlation weighting to the dimension-aligned action history features to obtain weighted action history features;

[0085] S205, the weighted visual features, weighted language features, and weighted action history features are spliced ​​together to obtain fused features;

[0086] S206, Perform feature compression processing on the fused features to obtain preprocessed multimodal features.

[0087] In this embodiment, visual features, linguistic features, and action history features are aligned to address the significant differences in dimensionality and scale among the various modalities, enabling the integration of these multi-source features into a single embedding space. Visual features are extracted using a convolutional neural network, linguistic features using a semantic encoding network, and action history features using a sequence modeling network; these features exhibit diverse dimensions and data distributions. Feature alignment can be based on fully connected mapping, mapping each modal feature to the same target dimensionality space through different linear transformation matrices, ensuring strict consistency in the tensor dimensions of the three features. After feature alignment, dimension-aligned visual features, dimension-aligned linguistic features, and dimension-aligned action history features are obtained, providing a consistent tensor input structure for subsequent weighted fusion.

[0088] Building upon this foundation, spatial attention weighting is applied to dimension-aligned visual features. This involves calculating weights across the spatial dimension using a spatial attention mechanism to highlight image region features relevant to the current task. For example, it can emphasize the edges of key objects or interactive areas in the input scene image. Spatial attention weighting is calculated by summing the spatial positions of visual features to output weighted visual features. Semantic association weighting is applied to dimension-aligned linguistic features. This weighting calculates the association strength between different words or phrases in the language using self-attention or a contextual association matrix, emphasizing task-relevant semantic parts. This makes the model's expression of instruction intent more accurate in multimodal fusion, outputting weighted linguistic features. Temporal association weighting is applied to dimension-aligned action history features. This weighting is based on a time dependency matrix or the temporal weight distribution of a recurrent neural network, highlighting key nodes in the historical action sequence and reducing attention to irrelevant historical states, outputting weighted action history features.

[0089] Weighted visual features, weighted linguistic features, and weighted action history features are concatenated. This concatenation method, rather than simple superposition, ensures faithful representation of the three types of features while maintaining their independence. The resulting fused features often have high dimensionality; therefore, feature compression is performed using dimensionality reduction networks or linear projection modules to reduce redundant information and computational complexity, while retaining the feature components that best represent multimodal correlations. The final result is a preprocessed multimodal feature.

[0090] This embodiment addresses the differences in the representational dimensions and scales of multimodal data by unifying the dimensional alignment of visual features, linguistic features, and action history features, enabling computation within the same space. Spatial attention weighting, semantic association weighting, and temporal association weighting are applied to visual features, linguistic features, and action history features respectively, enhancing the expressive ability of each modality for task-related information. By concatenating and compressing the multimodal weighted features, the dimensionality and redundancy of the feature space are reduced, allowing the preprocessed multimodal features output to more accurately reflect the overall relationship between visual, linguistic, and historical information. This improves the model's ability to express and integrate multimodal inputs, providing a rich and consistent feature foundation for subsequent decision-making steps, effectively enhancing decision accuracy and adaptability to dynamic environments.

[0091] In one embodiment, step S30 above includes:

[0092] S301, the preprocessed multimodal features are encoded through the semantic understanding layer of the hierarchical action decoder to obtain semantic understanding features;

[0093] S302, a graph structure representation based on the semantic understanding features is constructed through the dynamic programming layer of the hierarchical action decoder;

[0094] S303, In the dynamic programming layer, the node states and edge connections of the graph structure representation are updated in combination with the current environment state and task objectives to generate an updated graph structure representation;

[0095] S304, the dynamic programming layer plans the optimal action path based on the updated graph structure representation to obtain action planning features;

[0096] S305, the action planning features and action history features are fused through the action generation layer of the hierarchical action decoder to generate action sequence features.

[0097] In this embodiment, during the processing of preprocessed multimodal features by the hierarchical action decoder, the preprocessed multimodal features are first encoded through the semantic understanding layer of the hierarchical action decoder. The encoding operation of the semantic understanding layer can be based on a multilayer perceptron, a graph neural network embedding encoder, or other deep feature encoding modules. Its function is to perform semantic abstraction of the input fused features in a global dimension, extracting mid-to-high-order feature expressions related to task semantics and context, thereby forming semantic understanding features. Semantic understanding features not only retain the multimodal information in the original input, but also enhance semantic consistency through encoding methods, providing a unified feature foundation for subsequent structured planning.

[0098] After obtaining semantic understanding features, a graph structure representation based on these features is constructed through the dynamic programming layer of the hierarchical action decoder. This step maps the input semantic understanding features to node representations and defines the edge connections between nodes, expressing the action decision space in graph form. Node representations can use feature embedding vectors, and edge connections can be initialized based on prior task knowledge or spatial and semantic similarity to form an initial graph structure. Next, the graph structure representation is updated in the dynamic programming layer, taking into account the current environmental state and task objective. The update process includes updating the state of each node in the graph and readjusting the edge connections. For example, node state updates can incorporate weighted corrections to node feature vectors based on environmental sensor data, and edge connection updates can consider task constraints to adjust or prune edge weights, thereby generating an updated graph structure representation that matches the environment and task objective.

[0099] After obtaining the updated graph structure representation, the dynamic programming layer plans the optimal action path based on this structure, thus obtaining action planning features. Optimal action path planning can employ dynamic programming algorithms, graph search algorithms (such as A* algorithm and Dijkstra's algorithm), or reinforcement learning algorithms to find the optimal node sequence or path under the current environment and task objective, and integrate the node states in this path to form action planning features, which are used to guide the generation of subsequent action sequences.

[0100] Finally, the action planning features and action history features are fused through the action generation layer of the hierarchical action decoder. The fusion operation can be performed by concatenation, weighted averaging, attention mechanism weighting, etc., to fully combine the currently planned optimal action path information and historical action execution information, ensuring that the output action sequence features not only meet the requirements of the current context and environment, but also have continuity and consistency, thereby generating the final action sequence features that can be used for action control.

[0101] In practical implementation, the semantic understanding layer encoding can use a multi-layer fully connected network, with the ReLU activation function enhancing nonlinear expressive power. During the dynamic programming layer's construction of the graph representation, key states defined in the task can be used as nodes, and edge connections can be initialized based on the spatial proximity of objects in the visual scene. Node state updates can be performed by comparing with dynamically detected target states in the current environment, calculating differences, and adjusting node weights. Edge connection updates can dynamically adjust edge weights based on task priority or path accessibility. An improved A-algorithm can be used for action path planning, incorporating a task objective reward function to optimize path solving. The fusion operation of the action generation layer can employ a multi-head self-attention mechanism, assigning attention weights to action planning features and action history features separately before weighted summation to enhance adaptability to complex contextual relationships. Different graph search algorithms can be selected for path planning in different scenarios. Dijkstra's algorithm is used to guarantee a globally optimal solution in constrained environments, while a heuristic A-algorithm is used to reduce computational complexity in scenarios with high real-time requirements.

[0102] This embodiment addresses the potential semantic inconsistency between multimodal inputs by semantically encoding preprocessed multimodal features, enabling subsequent graph structure representations to be constructed within a unified semantic space. A dynamic programming layer dynamically updates the graph structure based on the current environmental state and task objectives, improving the adaptability of action planning to environmental changes. By planning the optimal action path based on the updated graph structure and integrating historical action execution information, the continuity and accuracy of the action sequence are enhanced. The final generated action sequence features fully integrate environmental dynamics, task objectives, and historical behavioral experience, improving the accuracy and stability of decision-making. Especially in complex dynamic scenarios, it can flexibly adjust action plans, effectively improving the system's task completion efficiency and execution reliability.

[0103] In one embodiment, step S40 above includes:

[0104] S401, Input the action sequence features into a fully connected layer for high-dimensional linear transformation to obtain a fully connected output;

[0105] S402, Apply a nonlinear activation function with range constraints to the fully connected output to obtain initial control parameters;

[0106] S403, Verify the initial control parameters based on physical constraints and generate verification results;

[0107] S404, when the verification result indicates that the initial control parameters are in violation, determine the parameter correction vector based on the physical constraint boundary;

[0108] S405, merge the parameter correction vector with the initial control parameters to generate intermediate control parameters;

[0109] S406, The feasibility of the intermediate control parameters is verified through dynamic simulation;

[0110] S407, when the intermediate control parameter passes verification, the intermediate control parameter is used as the adjusted control parameter;

[0111] S408, when the intermediate control parameter fails verification, obtain the collision state and stability index output by the dynamic simulation, determine the correction increment based on the collision state and stability index, superimpose the correction increment with the parameter correction vector to generate an updated parameter correction vector, generate a new intermediate control parameter based on the updated parameter correction vector, and use the new intermediate control parameter as the adjusted control parameter.

[0112] S409, Based on the adjusted control parameters, a continuous motion control command is generated as the motion decision.

[0113] In this embodiment, the action sequence features are first input into a fully connected layer to achieve a high-dimensional linear transformation. This high-dimensional linear transformation involves matrix multiplication of the input feature matrix and the weight matrix, followed by the addition of a bias term. This maps the original action sequence features to another linear space, giving the features stronger expressive power and adapting them to the input requirements of subsequent processing modules. The linear transformation result is the fully connected output. Subsequently, a nonlinear activation function with range constraints is applied to the fully connected output. Range constraints are achieved by selecting an appropriate activation function (e.g., sigmoid or tanh) and scaling and shifting its output to limit its numerical range within a predefined interval, preventing abnormal values ​​from being input into subsequent modules. This processing generates initial control parameters, serving as a set of physically controllable continuous numerical variables.

[0114] Then, the initial control parameters are physically constrained. Physical constraint verification includes boundary checks on the control parameters based on the physical model or task context requirements, such as whether the control joint angles, velocities, and accelerations are within the mechanically permissible range. The verification result is the verification outcome. If the verification result indicates that the initial control parameters meet the constraints, the process can directly proceed to motion instruction generation. If the verification result indicates that the initial control parameters violate the constraints, a parameter correction vector is determined based on the physical constraint boundaries. The parameter correction vector is determined by calculating the difference between the initial control parameters and the physical boundary extreme values, and determining the correction direction and magnitude to adjust the initial parameters to near the physically feasible range.

[0115] The parameter correction vector is fused with the initial control parameters to obtain intermediate control parameters. This fusion operation is generally completed through linear weighted summation or vector superposition, and the adjustment range can be adaptively determined to reduce the impact of adjustment on control accuracy. Subsequently, the intermediate control parameters are verified through dynamic simulation. Dynamic simulation verification involves inputting the intermediate control parameters into the robot's dynamics or kinematics model to simulate the physical response after execution, determining whether collisions, instability, or other infeasible states exist. If the intermediate control parameters pass verification, they are directly used as the adjusted control parameters for subsequent motion command generation modules.

[0116] If the intermediate control parameters fail to pass the dynamic simulation verification, further adjustments are required. Specifically, this involves obtaining the collision state and stability index from the dynamic simulation output. The collision state can be calculated based on the spatial distance between each node and obstacle in the simulation, while the stability index can be measured using parameters such as the centroid position, support polygons, and torque balance. Based on the collision state and stability index, a correction increment is determined to further adjust the parameter correction vector. The correction increment can be determined using a weighted function, considering the collision severity and stability margin. The correction increment is then superimposed onto the parameter correction vector to obtain the updated parameter correction vector. Based on the updated parameter correction vector, the initial control parameters are re-fused to generate new intermediate control parameters, which are then used as the adjusted control parameters to ensure safe and feasible control output even under complex conditions. Finally, continuous motion control commands are generated based on the adjusted control parameters as the action decision. The motion control commands are calculated based on the correspondence between the control parameters and the actuator model, ensuring that the control commands meet the requirements of continuity, smoothness, and dynamic feasibility.

[0117] This embodiment maps action sequence features to control parameters and progressively incorporates physical constraint verification, parameter correction, dynamic simulation verification, and iterative optimization. This ensures that the generated control parameters strictly meet physical feasibility requirements, and introduces simulation feedback for targeted correction when multiple non-compliance occurs, avoiding the direct output of unexecutable action decisions. Compared to directly outputting control parameters based solely on network prediction, this approach effectively improves the feasibility of control parameters, the accuracy and stability of action control, and enhances the overall system robustness in complex, dynamic, and constraint-rich environments, reducing task failures or safety risks caused by abnormal control parameters.

[0118] In one embodiment, step S50 above includes:

[0119] S501, acquires real-time environmental state change information collected by the sensor;

[0120] S502, the environmental state change information is input into the dynamic programming layer of the hierarchical action decoder, and the node states and edge connections of the graph structure representation are optimized in the dynamic programming layer to generate an optimized graph structure representation;

[0121] S503, Generate new action planning features based on the optimized graph structure representation;

[0122] S504, The new action planning features are processed by the action generation layer of the hierarchical action decoder to generate new action sequence features;

[0123] S505, the new action sequence features are mapped to new control parameters, and an updated action decision is generated based on the new control parameters.

[0124] In this embodiment, firstly, environmental state change information is collected in real time by sensors. This information may include the distribution of obstacles in the space, ambient lighting conditions, the motion trajectory of dynamic targets, and other external conditions that affect task execution. The collection method can employ a multimodal sensor array, such as a combination of an RGB-D camera, LiDAR, and inertial measurement unit, acquiring data in real time via a high-frequency, low-latency data stream. The collected environmental state change information is then passed as input data to the dynamic planning layer of the hierarchical action decoder.

[0125] The dynamic programming layer combines environmental state change information with preprocessed multimodal features and internal state estimates. It optimizes the internally maintained graph structure representation, a directed graph consisting of nodes and edges. Nodes represent potential state-space locations during action execution, and edges represent transition relationships and their costs between states. Node state optimization updates the attached state estimation vectors to nodes, such as spatial location estimates, risk indicators, and constraint states. Edge connection optimization adjusts edge weights to reflect the feasibility and cost of state transitions under the current environmental conditions, such as increased obstacle avoidance path costs or adjusted dynamic target interaction costs. The optimized graph structure representation is context-adaptable to the current environment and accurately describes the reachable state space and potential paths for action execution.

[0126] New action planning features are generated based on the optimized graph structure representation. These new features are calculated by performing shortest path search, minimum risk path search, or multi-objective trade-off path search within the optimized graph structure. The action planning features are encoded as vectors, containing information such as the node sequence of the planned path, the total path cost, and the distribution of risk indicators. Subsequently, the new action planning features are processed by an action generation layer. This layer can employ a sequence modeling network (such as a multilayer long short-term memory network LSTM or a gated recurrent unit GRU) to perform sequence modeling on the action planning features, generating new action sequence features. This processing ensures that the action sequence features are highly correlated with the latest environmental state and possess predictive adaptability to future dynamic changes.

[0127] The new action sequence features are input into the action decision mapping module for mapping. The mapping process uses a fully connected network or other nonlinear mapping model to convert the new action sequence features into a continuous set of control parameters. The control parameter set represents the specific executable command parameters in the continuous action space, such as angle, position, velocity, and acceleration. Finally, an updated action decision is generated based on the new control parameters. The updated action decision directly replaces the old action decision, enabling the overall system to adapt to changes in the current environment and achieve dynamic closed-loop control.

[0128] This embodiment acquires real-time environmental state change information and dynamically interacts with the graph structure representation in the hierarchical action decoder. It can quickly and adaptively optimize node states and edge connections after environmental changes occur, generating action planning features tailored to the current environmental conditions. These features are then transformed into new action sequence features by the action generation layer and mapped to new control parameters, ultimately updating the action decisions. Compared to static planning or fixed action sequences, this approach offers rapid adaptability and path correction capabilities to complex dynamic environments, reducing execution deviations caused by sudden environmental changes and improving the overall dynamic robustness and task completion rate of the system.

[0129] In one embodiment, after step S50 above, the method further includes:

[0130] S601, construct a multi-dimensional reward function composed of task completion, execution efficiency and energy consumption;

[0131] S602, execute the action decision and record the execution result;

[0132] S603, determine the instantaneous reward value based on the execution result and the multi-dimensional reward function;

[0133] S604, Construct a time-series decision dataset containing the preprocessed multimodal features, action decisions, execution results, and environmental feedback information;

[0134] S605, Based on the instantaneous reward value and the time-series decision dataset, optimize the parameters of the hierarchical action decoder.

[0135] In this embodiment, a multi-dimensional reward function is first constructed, comprising task completion, execution efficiency, and energy consumption. Task completion reward measures whether the action decision achieves the expected goal; for example, in a robot task, this could be based on whether the target location was successfully reached or the specified action was correctly completed. Execution efficiency reward measures the time consumption performance of the action decision; for example, it can be based on the time taken to complete the task or the progress within a time window. Energy consumption reward measures the degree of optimization of the action decision in terms of energy resource consumption; for example, it can be based on measuring the energy usage of the drive device or the power integral value during execution. This reward function, by assigning different weights to the multi-dimensional indicators, constitutes a comprehensive real-time reward function.

[0136] After executing an action decision, the execution results are recorded. These results include whether the action was successfully completed, the actual path, execution time, energy consumed, and dynamic changes and task status reflected in environmental feedback. An immediate reward value is determined based on the execution results and a multi-dimensional reward function. Specifically, the metric values ​​collected from the execution results are used as input to the reward function, and the immediate reward value is output. This value directly reflects the performance of the current action decision across multiple dimensions.

[0137] Subsequently, a time-series decision dataset is constructed, which includes preprocessed multimodal features, action decisions, execution results, and environmental feedback information. The time-series format means that the dataset records continuous decision-making states arranged chronologically, establishing causal relationships along the timeline, making it suitable for dynamic modeling and learning. Preprocessed multimodal features describe the comprehensive perceptual state before a decision is made, action decisions record the decision output, and execution results and environmental feedback reflect the action execution and the environment's response to the decision, respectively. The entire dataset is used for subsequent training and updates.

[0138] Based on real-time reward values ​​and a time-series decision dataset, the parameters of a hierarchical action decoder are optimized. The optimization process calculates the cumulative reward performance of action decisions and performs gradient backpropagation updates using time-series data. The optimization objective is to maximize the cumulative reward or minimize the reward loss function. The optimization targets include the semantic understanding layer parameters, dynamic programming layer parameters, and action generation layer parameters of the hierarchical action decoder. By simultaneously adjusting the parameters of these three layers, the overall performance of the hierarchical action decoder is improved, resulting in more accurate understanding and decision outputs of multimodal inputs and better adaptability to environmental changes and dynamic constraints.

[0139] Example: In the field of autonomous decision-making for intelligent agents, the entire decision-making process can be applied to complex task scenarios, such as the dynamic task execution of multi-task collaborative robots in a semi-structured environment.

[0140] First, the agent acquires visual data, language commands, and action history data. Visual data includes environmental image sequences, obstacle distributions, and target object locations. Language commands originate from the operator's natural language input. Action history data describes historical trajectories, posture changes, and related execution context. Visual data undergoes multi-level spatial feature extraction to form spatial perception information at different scales. Multi-scale pooling is then used to standardize the dimensions and unify scale representation. The standardized spatial features are weighted using a channel spatial attention mechanism to highlight salient visual regions relevant to the task and environment. Language commands are segmented and part-of-speech tagging to generate semantic units, encoded as word vectors, and semantically enhanced based on a domain knowledge base to enrich contextual understanding. Subsequently, the enhanced word vectors aggregate semantic relationships through a multi-head attention mechanism to capture important task intentions within the commands. Action history data is processed by extracting temporal trends and combining them with a temporal dependency modeling network to characterize long-term dynamic patterns, supporting the agent's retrospective use of previous experiences. Finally, visual features, language features, and action history features serve as the initial multimodal inputs.

[0141] Next, the agent aligns these three types of multimodal features to ensure consistent dimensionality and meet the requirements of subsequent parallel computation. Visual features are weighted with spatial attention to enhance spatial contextual information; linguistic features are weighted with semantic association to highlight key task semantics; and action history features are weighted with temporal association to preserve action continuity and causality. The weighted multimodal features are then concatenated to form fused features. These fused features are further compressed to remove redundant dimensions and improve computational efficiency, resulting in a comprehensive state description that serves as input to the agent from the preprocessed multimodal features.

[0142] Preprocessed multimodal features are input into a hierarchical action decoder. The semantic understanding layer is responsible for integrating perceived state and semantic information to generate a compact semantic understanding representation. The dynamic programming layer constructs a graph structure representation based on the semantic understanding features, where node states and edge connections characterize the current environment layout and task dependencies. This graph structure is iteratively optimized within the dynamic programming layer by combining environmental states and task objectives, updating node attributes and edge weights to ensure that the graph structure maintains optimal reachability and path rationality under environmental changes or objective adjustments. The optimized graph structure is used to generate action planning features. These planning features are input into the action generation layer and fused with historical action features to obtain action sequence features. These action sequence features provide the data foundation for the next step of action parameter generation.

[0143] Subsequently, the action sequence features are mapped to control parameters. First, a high-dimensional linear transformation is performed through a fully connected layer to standardize the feature representation of the action space. The output of the fully connected layer is then passed through a nonlinear activation function with range constraints to ensure the usability of the initial control parameters within the action domain. The agent verifies the initial control parameters against physical constraints, such as the reachability of the robotic arm, joint limits, and safety boundaries. When the verification indicates a parameter violation, a parameter correction vector is determined based on the physical boundaries and fused with the initial control parameters to generate intermediate control parameters. The feasibility of the intermediate control parameters is verified through dynamic simulation. If the simulation passes, the parameters are adopted; otherwise, the agent obtains the collision state and stability index from the simulation output, calculates the correction increment, adds it to the parameter correction vector, and regenerates the intermediate control parameters until verification is successful. The adjusted control parameters are used to generate continuous action control commands, ultimately forming the action decision.

[0144] After executing an action decision, the agent collects environmental feedback information through sensors, such as dynamic changes in obstacles, target movement, and external interference signals. This environmental feedback information is input into the dynamic programming layer of the hierarchical action decoder, which optimizes the node states and edge connections in the graph representation in real time, dynamically adjusting the planning graph to adapt to environmental updates. The optimized graph structure generates new action planning features, which are processed by the action generation layer to obtain new action sequence features. These new action sequence features are then mapped to new control parameters to generate updated action decisions, ensuring that the agent dynamically adjusts its strategy during continuous operation.

[0145] The immediate reward value is calculated by combining the behavioral results and environmental feedback after execution with a multi-dimensional reward function that considers task completion, execution efficiency, and energy consumption. The agent records preprocessed multimodal features, action decisions, execution results, and environmental feedback to form a time-series decision dataset. Based on the immediate reward value and the dataset, a parameter optimization method is used to jointly update the parameters of the semantic understanding layer, dynamic programming layer, and action generation layer of the hierarchical action decoder to improve the model's long-term decision accuracy, adaptability, and task execution efficiency in complex tasks.

[0146] Intelligent agents can continuously enhance their adaptability to environmental dynamics and complex tasks through a closed-loop process from multimodal perception, historical experience, autonomous action planning to adaptive execution, and through real-time environmental perception and feedback iterative optimization.

[0147] In the healthcare field, the intelligent agent autonomous decision-making process can be specifically applied to robot-assisted surgical systems to dynamically adjust surgical paths, action plans, and execution operations to adapt to complex surgical environments and real-time changes in physician instructions. The system first extracts multi-level spatial features from high-resolution surgical visual images, encompassing surgical instruments, tissue structures, and local lesion locations. Then, it achieves size standardization through multi-scale pooling and combines channel spatial attention weighting to highlight areas relevant to the target tissue. The physician inputs surgical intentions via natural language commands. The system performs word segmentation and part-of-speech tagging on the input statements, encoding them into word vectors, and supplementing them with a medical knowledge base to enhance understanding of medical terminology and clinical instructions. Semantic aggregation is then achieved through a multi-head attention mechanism. The system combines historical surgical data to extract time-series variation features and processes them through a temporal modeling network to form a complete surgical context. After aligning the three types of features, spatial attention, semantic association, and temporal association weighting are applied, followed by concatenation and compression to obtain preprocessed multimodal features that comprehensively describe the current surgical environment, semantic intent, and historical operations.

[0148] These preprocessed multimodal features are input into a hierarchical action decoder. The semantic understanding layer encodes the current environmental state and semantic information. The dynamic programming layer constructs and updates the graph structure representation in real time, dynamically adjusting node and edge attributes to describe the tissue relationships and spatial constraints within the surgical area. Based on this optimized graph structure, the optimal surgical path is planned. The action generation layer fuses the planned path with the surgical history to generate action sequence features. These action sequence features are mapped to continuous operational parameters, such as the position, angle, and force control commands of surgical instruments. The parameters are verified through physical constraints and dynamic simulation to ensure that the actions are feasible within the constraints of joint limits, spatial accessibility, and patient safety. Precise control is achieved through feedback-driven iterative correction.

[0149] During the surgery, the system collects real-time environmental feedback information, such as changes in tissue movement, bleeding, and instrument displacement. This information is input into the description of the nodes and edges in the dynamic programming layer to generate new action plans and action sequence features, which are then mapped to new control parameters. The surgical path is adjusted in real-time to ensure safety and efficiency. Post-execution records and surgical outcomes are used for immediate reward calculation. Combined with indicators such as task completion (e.g., completeness of target tissue resection), execution efficiency (e.g., surgical time), and energy consumption, a time-series decision dataset is constructed. Based on the reward value and the decision dataset, the system jointly updates the parameters of the semantic understanding layer, dynamic programming layer, and action generation layer through parameter optimization methods. This enables adaptive optimization for the next round of surgical tasks, improving the accuracy, flexibility, and reliability of the robotic system in complex medical scenarios.

[0150] In the financial health sector, intelligent agent autonomous decision-making processes can be applied to multimodal customer risk assessment and dynamic decision support systems, particularly in complex credit approval or financial health management service scenarios. The system first acquires visual data from customer facial recognition images and scene monitoring images, extracting multi-level spatial features to identify facial emotions, identity consistency, and behavioral characteristics. Multi-scale pooling is used to unify the spatial feature size, and a channel spatial attention mechanism is applied to enhance key features associated with the customer's state. Customer-provided text descriptions, intent statements, and consultation statements serve as linguistic inputs. These are segmented and tagged with parts of speech, encoded into word vectors, and then supplemented with semantic background information, such as risk vocabulary and policy terms, by accessing a financial health knowledge base. A multi-head attention mechanism is used to aggregate semantic associations, resulting in deeply understood linguistic features. Combining customer historical interaction records, transaction behavior, and account dynamics, temporal change features are extracted, and a temporal dependency modeling network is used to form a complete financial behavior context description.

[0151] The aforementioned visual features, linguistic features, and action history features, after feature alignment processing, are concatenated and compressed into preprocessed multimodal features through spatial attention weighting, semantic association weighting, and temporal association weighting, respectively, forming a holographic representation of the current customer state. This preprocessed multimodal feature is input into a hierarchical action decoder. The semantic understanding layer encodes it with multi-dimensional financial semantics, and the dynamic planning layer constructs and dynamically updates a graph structure representation of the customer's risk features, incorporating the current market environment and target risk control requirements into the updates of nodes and edges in the graph to plan the optimal risk control response path. The action generation layer fuses the graph structure output with the customer's historical financial behavior to generate action sequence features. These action sequence features are processed through high-dimensional spatial linear mapping and nonlinear activation functions to generate initial control parameters, representing the customer's comprehensive risk score, credit limit recommendation, and financial health recommendation. The control parameters are tested under physical constraints (e.g., credit limit range, risk threshold, compliance constraints) and verified through dynamic financial business simulations (e.g., approval process simulation, market volatility simulation) to ensure feasibility. When a non-compliant or infeasible parameter is detected, the system iteratively adjusts the correction vector based on the financial stability indicators and risk deviation status output by dynamic simulation, regenerates and updates the control parameters, and outputs continuous risk control decisions or customer service suggestions.

[0152] The system collects real-time environmental feedback information during interaction, such as changes in customer status, real-time market conditions, and public opinion data. This information is input into the dynamic programming layer to optimize the customer graph structure risk model, generating new action plans and action sequence features, and adjusting risk control decisions in real time. The execution results of each round of decisions are recorded and combined with multi-dimensional reward functions, based on factors such as task completion (risk mitigation level), execution efficiency (approval time and accuracy), and cost consumption (risk control investment and compliance costs), to construct a time-series decision dataset. Based on real-time reward values ​​and historical decision data, the system uses parameter optimization methods to jointly update the parameters of the semantic understanding layer, dynamic programming layer, and action generation layer of the hierarchical action decoder. This enables dynamic adaptation and optimization for future customer risk assessment and financial health management tasks, significantly improving the accuracy, agility, and compliance of financial services.

[0153] This embodiment optimizes and updates the parameters of the hierarchical action decoder by using a multi-dimensional reward function, combining performance metrics such as task completion, execution efficiency, and energy consumption. This improves the quality and robustness of action decisions. By using a time-series decision dataset to support parameter optimization, the decoder can not only reflect the immediate state but also learn cross-temporal dependencies from historical decisions. Overall, this enhances the adaptive capability in complex dynamic environments, shortens task completion time, reduces energy consumption, and improves the sensitivity and response speed of action decisions to real-time environmental feedback.

[0154] In one embodiment, a decision-making device based on multimodal information is provided, which corresponds one-to-one with the decision-making method based on multimodal information described in the above embodiments. (Refer to...) Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the decision-making device based on multimodal information of the present invention. The modules include a multimodal perception module 10, a multimodal fusion module 20, an action decoding module 30, an action decision-making module 40, and a feedback adjustment module 50. Detailed descriptions of each functional module are as follows:

[0155] The multimodal perception module 10 is used to acquire visual data, language commands, and action history data, and process the visual data, language commands, and action history data to obtain visual features, language features, and action history features, respectively.

[0156] The multimodal fusion module 20 is used to fuse the visual features, language features, and action history features to obtain preprocessed multimodal features;

[0157] Action decoding module 30 is used to process the preprocessed multimodal features through a hierarchical action decoder to generate action sequence features;

[0158] Action decision module 40 is used to map the action sequence features to control parameters and generate action decisions based on the control parameters;

[0159] The feedback adjustment module 50 is used to collect environmental feedback information, generate new action sequence features based on the environmental feedback information, and update the action decision based on the new action sequence features.

[0160] In one embodiment, the multimodal sensing module 10 is specifically used for:

[0161] Acquire visual data and extract multi-level spatial features from the visual data;

[0162] Multi-scale pooling operations are performed on the multi-level spatial features to obtain spatial features of uniform size;

[0163] The channel spatial attention mechanism is applied to the spatial features of the uniform size to obtain visual features;

[0164] Obtain language instructions, perform word segmentation and part-of-speech tagging on the language instructions, and generate word segmentation sequences and part-of-speech tagging information respectively;

[0165] The word segmentation sequence and part-of-speech tagging information are encoded into word vectors;

[0166] The word vectors are supplemented with semantic information using a knowledge base to obtain enhanced word vectors;

[0167] A multi-head attention mechanism is applied to the enhanced word vectors to obtain linguistic features;

[0168] Acquire historical action data and extract the temporal variation characteristics of the historical action data;

[0169] The temporal dependency modeling network is applied to process the temporal change features to establish long-term dependencies and obtain action history features.

[0170] In one embodiment, the multimodal fusion module 20 is specifically used for:

[0171] The visual features, language features, and action history features are aligned to obtain dimension-aligned visual features, dimension-aligned language features, and dimension-aligned action history features.

[0172] Spatial attention weighting is applied to the dimension-aligned visual features to obtain weighted visual features;

[0173] Semantic association weighting is applied to the dimensionally aligned language features to obtain weighted language features;

[0174] Apply temporal correlation weighting to the dimension-aligned action history features to obtain weighted action history features;

[0175] By combining the weighted visual features, weighted language features, and weighted action history features, a fused feature is obtained.

[0176] The fused features are subjected to feature compression processing to obtain preprocessed multimodal features.

[0177] In one embodiment, the action decoding module 30 is specifically used for:

[0178] The preprocessed multimodal features are encoded by the semantic understanding layer of the hierarchical action decoder to obtain semantic understanding features;

[0179] A graph structure representation based on the semantic understanding features is constructed through the dynamic programming layer of the hierarchical action decoder;

[0180] In the dynamic programming layer, the node states and edge connections of the graph structure representation are updated by combining the current environment state and task objectives, thereby generating an updated graph structure representation.

[0181] The dynamic programming layer plans the optimal action path based on the updated graph structure representation to obtain action planning features.

[0182] The action generation layer of the hierarchical action decoder fuses the action planning features and action history features to generate action sequence features.

[0183] In one embodiment, the action decision module 40 is specifically used for:

[0184] The action sequence features are input into a fully connected layer for high-dimensional linear transformation to obtain a fully connected output.

[0185] The fully connected output is processed by a nonlinear activation function with range constraints to obtain initial control parameters;

[0186] The initial control parameters are verified based on physical constraints, and verification results are generated.

[0187] When the verification result indicates that the initial control parameters are in violation, a parameter correction vector is determined based on the physical constraint boundary.

[0188] The intermediate control parameters are generated by fusing the parameter correction vector with the initial control parameters;

[0189] The feasibility of the intermediate control parameters was verified through dynamic simulation.

[0190] When the intermediate control parameter passes verification, the intermediate control parameter is used as the adjusted control parameter;

[0191] When the intermediate control parameters fail to pass verification, the collision state and stability index of the dynamic simulation output are obtained, the correction increment is determined based on the collision state and stability index, the correction increment is superimposed with the parameter correction vector to generate an updated parameter correction vector, a new intermediate control parameter is generated based on the updated parameter correction vector, and the new intermediate control parameter is used as the adjusted control parameter.

[0192] Based on the adjusted control parameters, continuous motion control commands are generated as motion decisions.

[0193] In one embodiment, the feedback adjustment module 50 is specifically used for:

[0194] Acquire real-time environmental state change information collected by sensors;

[0195] The environmental state change information is input into the dynamic programming layer of the hierarchical action decoder. In the dynamic programming layer, the node states and edge connections of the graph structure representation are optimized to generate an optimized graph structure representation.

[0196] New action planning features are generated based on the optimized graph structure representation;

[0197] The new action planning features are processed by the action generation layer of the hierarchical action decoder to generate new action sequence features;

[0198] The new action sequence features are mapped to new control parameters, and updated action decisions are generated based on the new control parameters.

[0199] In one embodiment, the feedback adjustment module 50 is specifically used for:

[0200] Construct a multi-dimensional reward function composed of task completion rate, execution efficiency, and energy consumption;

[0201] Execute the action decision and record the execution result;

[0202] The immediate reward value is determined based on the execution result and the multi-dimensional reward function.

[0203] Construct a time-series decision dataset containing the preprocessed multimodal features, action decisions, execution results, and environmental feedback information;

[0204] The parameters of the hierarchical action decoder are optimized based on the instant reward value and the time-series decision dataset.

[0205] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used for communication with external user terminals via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a multimodal information-based decision-making method on the server side.

[0206] In one embodiment, a computer device is provided, which may be a user terminal, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides decision-making and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements user-side functions or steps of a decision-making method based on multimodal information.

[0207] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:

[0208] Acquire visual data, language commands, and action history data, and process the visual data, language commands, and action history data to obtain visual features, language features, and action history features, respectively;

[0209] By fusing the aforementioned visual features, linguistic features, and action history features, preprocessed multimodal features are obtained;

[0210] The preprocessed multimodal features are processed by a hierarchical action decoder to generate action sequence features;

[0211] The action sequence features are mapped to control parameters, and action decisions are generated based on the control parameters;

[0212] Collect environmental feedback information, generate new action sequence features based on the environmental feedback information, and update the action decision based on the new action sequence features.

[0213] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0214] Acquire visual data, language commands, and action history data, and process the visual data, language commands, and action history data to obtain visual features, language features, and action history features, respectively;

[0215] By fusing the aforementioned visual features, linguistic features, and action history features, preprocessed multimodal features are obtained;

[0216] The preprocessed multimodal features are processed by a hierarchical action decoder to generate action sequence features;

[0217] The action sequence features are mapped to control parameters, and action decisions are generated based on the control parameters;

[0218] Collect environmental feedback information, generate new action sequence features based on the environmental feedback information, and update the action decision based on the new action sequence features.

[0219] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0220] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0221] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0222] It should be noted that if any software tools or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The embodiments described above are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A decision-making method based on multimodal information, characterized in that, Includes the following steps: Acquire visual data, language commands, and action history data, and process the visual data, language commands, and action history data to obtain visual features, language features, and action history features, respectively; By fusing the aforementioned visual features, linguistic features, and action history features, preprocessed multimodal features are obtained; The preprocessed multimodal features are processed by a hierarchical action decoder to generate action sequence features; The action sequence features are mapped to control parameters, and action decisions are generated based on the control parameters; Collect environmental feedback information, generate new action sequence features based on the environmental feedback information, and update the action decision based on the new action sequence features.

2. The decision-making method based on multimodal information as described in claim 1, characterized in that, Acquire visual data, language commands, and action history data, and process the visual data, language commands, and action history data to obtain visual features, language features, and action history features, respectively, including: Acquire visual data and extract multi-level spatial features from the visual data; Multi-scale pooling operations are performed on the multi-level spatial features to obtain spatial features of uniform size; The channel spatial attention mechanism is applied to the spatial features of the uniform size to obtain visual features; Obtain language instructions, perform word segmentation and part-of-speech tagging on the language instructions, and generate word segmentation sequences and part-of-speech tagging information respectively; The word segmentation sequence and part-of-speech tagging information are encoded into word vectors; The word vectors are supplemented with semantic information using a knowledge base to obtain enhanced word vectors; A multi-head attention mechanism is applied to the enhanced word vectors to obtain linguistic features; Acquire historical action data and extract the temporal variation characteristics of the historical action data; The temporal dependency modeling network is applied to process the temporal change features to establish long-term dependencies and obtain action history features.

3. The decision-making method based on multimodal information as described in claim 1, characterized in that, By fusing the aforementioned visual features, linguistic features, and action history features, preprocessed multimodal features are obtained, including: The visual features, language features, and action history features are aligned to obtain dimension-aligned visual features, dimension-aligned language features, and dimension-aligned action history features. Spatial attention weighting is applied to the dimension-aligned visual features to obtain weighted visual features; Semantic association weighting is applied to the dimensionally aligned language features to obtain weighted language features; Apply temporal correlation weighting to the dimension-aligned action history features to obtain weighted action history features; By combining the weighted visual features, weighted language features, and weighted action history features, a fused feature is obtained. The fused features are subjected to feature compression processing to obtain preprocessed multimodal features.

4. The decision-making method based on multimodal information as described in claim 1, characterized in that, The preprocessed multimodal features are processed by a hierarchical action decoder to generate action sequence features, including: The preprocessed multimodal features are encoded by the semantic understanding layer of the hierarchical action decoder to obtain semantic understanding features; A graph structure representation based on the semantic understanding features is constructed through the dynamic programming layer of the hierarchical action decoder; In the dynamic programming layer, the node states and edge connections of the graph structure representation are updated by combining the current environment state and task objectives, thereby generating an updated graph structure representation. The dynamic programming layer plans the optimal action path based on the updated graph structure representation to obtain action planning features. The action generation layer of the hierarchical action decoder fuses the action planning features and action history features to generate action sequence features.

5. The decision-making method based on multimodal information as described in claim 1, characterized in that, Mapping the action sequence features to control parameters and generating action decisions based on the control parameters includes: The action sequence features are input into a fully connected layer for high-dimensional linear transformation to obtain a fully connected output. The fully connected output is processed by a nonlinear activation function with range constraints to obtain initial control parameters; The initial control parameters are verified based on physical constraints, and verification results are generated. When the verification result indicates that the initial control parameters are in violation, a parameter correction vector is determined based on the physical constraint boundary. The intermediate control parameters are generated by fusing the parameter correction vector with the initial control parameters; The feasibility of the intermediate control parameters was verified through dynamic simulation. When the intermediate control parameter passes verification, the intermediate control parameter is used as the adjusted control parameter; When the intermediate control parameters fail to pass verification, the collision state and stability index of the dynamic simulation output are obtained, the correction increment is determined based on the collision state and stability index, the correction increment is superimposed with the parameter correction vector to generate an updated parameter correction vector, a new intermediate control parameter is generated based on the updated parameter correction vector, and the new intermediate control parameter is used as the adjusted control parameter. Based on the adjusted control parameters, continuous motion control commands are generated as motion decisions.

6. The decision-making method based on multimodal information as described in claim 1, characterized in that, Collecting environmental feedback information, generating new action sequence features based on the environmental feedback information, and updating the action decision based on the new action sequence features, including: Acquire real-time environmental state change information collected by sensors; The environmental state change information is input into the dynamic programming layer of the hierarchical action decoder. In the dynamic programming layer, the node states and edge connections of the graph structure representation are optimized to generate an optimized graph structure representation. New action planning features are generated based on the optimized graph structure representation; The new action planning features are processed by the action generation layer of the hierarchical action decoder to generate new action sequence features; The new action sequence features are mapped to new control parameters, and updated action decisions are generated based on the new control parameters.

7. The decision-making method based on multimodal information as described in claim 1, characterized in that, After collecting environmental feedback information and generating new action sequence features based on the environmental feedback information, and updating the action decision based on the new action sequence features, the method further includes: Construct a multi-dimensional reward function composed of task completion rate, execution efficiency, and energy consumption; Execute the action decision and record the execution result; The immediate reward value is determined based on the execution result and the multi-dimensional reward function. Construct a time-series decision dataset containing the preprocessed multimodal features, action decisions, execution results, and environmental feedback information; The parameters of the hierarchical action decoder are optimized based on the instant reward value and the time-series decision dataset.

8. A decision-making device based on multimodal information, characterized in that, The decision-making device based on multimodal information includes: A multimodal perception module is used to acquire visual data, language commands, and action history data, and process the visual data, language commands, and action history data to obtain visual features, language features, and action history features, respectively. A multimodal fusion module is used to fuse the visual features, language features, and action history features to obtain preprocessed multimodal features; The action decoding module is used to process the preprocessed multimodal features through a hierarchical action decoder to generate action sequence features; An action decision module is used to map the action sequence features to control parameters and generate action decisions based on the control parameters. The feedback adjustment module is used to collect environmental feedback information, generate new action sequence features based on the environmental feedback information, and update the action decision based on the new action sequence features.

9. A computer device, characterized in that, The computer device includes a memory, a processor, and a decision-making program based on multimodal information stored in the memory and executable on the processor, wherein when executed by the processor, the decision-making program based on multimodal information implements the steps of the decision-making method based on multimodal information as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The storage medium stores a decision-making program based on multimodal information, which, when executed by a processor, implements the steps of the decision-making method based on multimodal information as described in any one of claims 1-7.

Citation Information

Cited By

  • Content recommendation method and related device

    CN121144620A

  • Robot control reconstruction method and system based on PLC replacement

    CN121625172A

  • Multi-round dialogue credible correction method and device in visual language large model

    CN122264134A