Multi-modal fusion feature driven motion control method, device, equipment and medium
Patent Information
- Application Number
- CN202511051744.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2045-07-29
AI Technical Summary
[0006]本发明的主要目的在于提供一种多模态融合特征驱动的动作控制方法、装置、设备及存储介质,旨在解决现有技术中存在的决策易陷入局部最优、多模态信息融合不充分以及动作规划缺乏全局最优性与环境适应性的技术问题
[0023] Beneficial Effects: This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as embodied intelligence, fintech, and healthcare. It discloses a multimodal fusion feature-driven action control method, device, equipment, and medium, comprising: acquiring first modal input information and second modal input information; generating a first modal feature vector based on the first modal input information; generating a second modal feature vector based on the second modal input information; fusing the first and second modal feature vectors to generate a multimodal fusion feature; generating action commands based on the multimodal fusion feature; generating an initial action plan based on the action command, the current state of the device, current environmental information, and the task objective; generating a globally optimal action sequence based on the initial action plan; and controlling the device to execute the globally optimal action sequence. This invention generates an initial action plan by combining the action command generated from the multimodal fusion feature with the device state, environmental information, and task objective, and optimizes and generates a globally optimal action sequence based on the initial action plan to control device execution. This improves the decision-making performance of intelligent entities in complex environments, achieves efficient fusion of multimodal information, enhances the globality and environmental adaptability of action planning, avoids local optima problems, and improves the accuracy and flexibility of overall task execution.
Smart Images

Figure CN120952044B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a multimodal fusion feature-driven motion control method, apparatus, device, and storage medium. Background Technology
[0002] Despite the widespread application of embodied intelligence technology, existing perception and decision-making methods still have significant shortcomings in dealing with complex environments and multi-task requirements. Especially in the processing and fusion of perceived information, traditional algorithms typically rely on single-modal inputs and lack the ability to deeply understand multi-source information. This makes it difficult for embodied intelligence systems to accurately acquire comprehensive environmental information in dynamic and changing application scenarios, thus affecting the scientific validity of overall decision-making and the stability of execution.
[0003] In the field of embodied intelligence, especially in applications such as industrial production, complex assembly, and autonomous operation, robots face complex environmental structures and diverse task types. Existing perception and decision-making technologies generally suffer from the inability to escape local optima. When dealing with complex spatial layouts or multi-task collaboration requirements, robots often lack global optimality in the generated action strategies due to incomplete perception information or inaccurate information fusion, ultimately affecting task efficiency and operational accuracy.
[0004] In the fintech sector, intelligent robots and terminals are increasingly involved in business processing, physical operations, and risk inspections. However, existing technologies have significant limitations in handling multimodal information in complex financial business environments. Visual, audio, and other multi-source data are often intertwined in financial business scenarios. Traditional fusion processing methods cannot efficiently and accurately integrate this information, leading to information gaps or misjudgments during high-precision operations or dynamic decision-making by robots. This severely impacts the stability and security of business processes.
[0005] In the healthcare sector, intelligent medical devices and robots are increasingly being applied to scenarios such as surgical assistance, rehabilitation care, and hospital logistics. However, current technologies have limited capabilities for fusing multimodal information, making it difficult to achieve collaborative processing of multi-source information such as vision and audio. This results in medical robots being unable to accurately understand environmental changes and human instructions when performing critical tasks. Furthermore, existing motion planning methods lack flexibility and adaptability in the face of complex and ever-changing medical environments and emergencies, making it difficult for devices to adjust their motion strategies in a timely manner, thus affecting the safety and efficiency of operations. Summary of the Invention
[0006] The main objective of this invention is to provide a multimodal fusion feature-driven motion control method, apparatus, device, and storage medium, aiming to solve the technical problems in the prior art, such as decision-making easily getting trapped in local optima, insufficient multimodal information fusion, and lack of global optimality and environmental adaptability in motion planning.
[0007] To achieve the above objectives, the present invention provides a multimodal fusion feature-driven motion control method, comprising:
[0008] Acquire first modal input information and second modal input information, and generate a first modal feature vector based on the first modal input information and a second modal feature vector based on the second modal input information;
[0009] The first modality feature vector and the second modality feature vector are fused to generate a multimodal fusion feature;
[0010] Action commands are generated based on the multimodal fusion features;
[0011] Based on the action instructions, the current state of the device, the current environmental information, and the task objective, an initial action plan is generated;
[0012] Generate a globally optimal action sequence based on the initial action planning;
[0013] Control the device to execute the globally optimal action sequence.
[0014] Furthermore, to achieve the above objectives, the present invention provides a multimodal fusion feature-driven motion control device, comprising:
[0015] A multimodal feature extraction module is used to acquire first modal input information and second modal input information, and generate a first modal feature vector based on the first modal input information and a second modal feature vector based on the second modal input information;
[0016] A multimodal fusion module is used to fuse the first modality feature vector and the second modality feature vector to generate multimodal fusion features;
[0017] The action generation module is used to generate action instructions based on the multimodal fusion features;
[0018] The action planning module is used to generate an initial action plan based on the action command, the current state of the device, the current environmental information, and the task objective.
[0019] The sequence optimization module is used to generate a globally optimal action sequence based on the initial action planning;
[0020] The device control module is used to control the device to execute the globally optimal action sequence.
[0021] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and a multimodal fusion feature-driven motion control program stored in the memory and executable on the processor, wherein when the multimodal fusion feature-driven motion control program is executed by the processor, it implements the steps of the multimodal fusion feature-driven motion control method as described above.
[0022] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing a multimodal fusion feature-driven motion control program, wherein the multimodal fusion feature-driven motion control program, when executed by a processor, implements the steps of the multimodal fusion feature-driven motion control method as described above.
[0023] Beneficial Effects: This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as embodied intelligence, fintech, and healthcare. It discloses a multimodal fusion feature-driven action control method, device, equipment, and medium, comprising: acquiring first modal input information and second modal input information; generating a first modal feature vector based on the first modal input information; generating a second modal feature vector based on the second modal input information; fusing the first and second modal feature vectors to generate a multimodal fusion feature; generating action commands based on the multimodal fusion feature; generating an initial action plan based on the action command, the current state of the device, current environmental information, and the task objective; generating a globally optimal action sequence based on the initial action plan; and controlling the device to execute the globally optimal action sequence. This invention generates an initial action plan by combining the action command generated from the multimodal fusion feature with the device state, environmental information, and task objective, and optimizes and generates a globally optimal action sequence based on the initial action plan to control device execution. This improves the decision-making performance of intelligent entities in complex environments, achieves efficient fusion of multimodal information, enhances the globality and environmental adaptability of action planning, avoids local optima problems, and improves the accuracy and flexibility of overall task execution. Attached Figure Description
[0024] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings:
[0025] Figure 1 This is a schematic diagram of an application environment for a multimodal fusion feature-driven motion control method according to an embodiment of the present invention;
[0026] Figure 2 This is a flowchart illustrating an embodiment of the multimodal fusion feature-driven motion control method of the present invention;
[0027] Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the multimodal fusion feature-driven motion control device of the present invention;
[0028] Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention;
[0029] Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation
[0030] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0031] The multimodal fusion feature-driven motion control method provided in this invention can be applied to, for example... Figure 1 In this application environment, the user terminal communicates with the server via a network. The server can obtain first and second modal input information from the user terminal, generate a first modal feature vector based on the first modal input information, and generate a second modal feature vector based on the second modal input information; fuse the first and second modal feature vectors to generate multimodal fusion features; generate action commands based on the multimodal fusion features; generate initial action plans based on the action commands, current device state, current environmental information, and task objectives; generate a globally optimal action sequence based on the initial action plan; and control the device to execute the globally optimal action sequence. This invention generates initial action plans by combining action commands generated from multimodal fusion features with device state, environmental information, and task objectives, and optimizes and generates a globally optimal action sequence based on the initial action plans to control device execution. This improves the decision-making performance of intelligent entities in complex environments, achieves efficient fusion of multimodal information, enhances the globality and environmental adaptability of action planning, avoids local optima problems, and improves the accuracy and flexibility of overall task execution. The user terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The present invention will now be described in detail through specific embodiments.
[0032] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the multimodal fusion feature-driven motion control method provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0033] like Figure 2 As shown, the multimodal fusion feature-driven motion control method proposed in this invention includes the following steps:
[0034] S10, acquire first modal input information and second modal input information, generate a first modal feature vector based on the first modal input information, and generate a second modal feature vector based on the second modal input information;
[0035] In this embodiment, acquiring the first modal input information refers to acquiring data input with specific modal attributes through an information acquisition device. The first modal input information may include, but is not limited to, image data, video data, point cloud data, infrared information, depth information, or other data related to the spatial environment, object features, and visual perception. The type of information acquisition device can be selected according to application requirements, and may specifically include cameras, radar, laser scanning equipment, structured light sensors, depth cameras, 3D scanners, etc. The information acquisition process may include initializing the data source, setting parameters, adjusting image resolution or frame rate, and data caching and synchronization. During the acquisition of the first modal input information, a data fusion strategy from multiple sensor sources can be combined to improve information integrity and environmental adaptability.
[0036] Generating a first-modality feature vector based on first-modality input information refers to the process of transforming the acquired raw data into a structured form using feature extraction algorithms, extracting numerical representations that reflect environmental states, spatial structures, object attributes, or other visually relevant attributes. The first-modality feature vector can be a fixed-length or variable-length data array. The specific data structure can be designed according to subsequent processing needs. Feature extraction methods can include convolutional neural networks, visual transformers, graph neural networks, 3D point cloud coding structures, or other deep learning models with image information processing capabilities. The data transformation process can include multi-level feature aggregation, scale invariance enhancement, edge information extraction, or spatial relationship modeling. The generated first-modality feature vector should be able to express both static structures and dynamic changes in the environment, ensuring the accuracy and robustness of subsequent processing.
[0037] Acquiring second-modal input information refers to obtaining data input with independent information dimensions based on a data source different from the first modality. Second-modal input information may include audio data, speech information, vibration signals, temperature data, gas concentration, human physiological indicators, pressure sensing information, or other data related to hearing, sound, vibration, sound field environment, or environmental physical changes. Information sources may include microphone arrays, sound sensors, ambient sound monitoring equipment, voice acquisition devices, vibration sensing modules, or other multimodal information acquisition systems. The acquisition process may involve equipment calibration, signal amplification, filtering and noise reduction, data timing synchronization, or multi-channel data integration.
[0038] Generating second-modal feature vectors based on second-modal input information refers to the structural transformation of second-modal input information using signal processing algorithms and deep learning structures to extract parameter information representing changes in audio, speech, environment, or physical state. Second-modal feature vectors can be vectors, matrices, tensors, or other numerical forms suitable for multidimensional representation. Specifically, they can include spectral features, Mel-frequency cepstral coefficients, pitch parameters, temporal energy distribution, signal envelope information, or other feature expressions reflecting time-frequency structure and signal dynamics. The data transformation process can utilize one-dimensional convolutional networks, recurrent neural networks, acoustic transducer structures, end-to-end audio encoders, and time-frequency joint modeling modules to ensure that the second-modal feature vectors possess good generalization performance and environmental adaptability while maintaining the integrity of information representation.
[0039] High-resolution cameras can capture environmental images to generate first-modal input information. This first-modal feature vector is then generated using a multi-scale feature extraction structure based on convolutional neural networks. The visual encoding network parameters can be optimized and adjusted according to different environmental conditions; for example, an image enhancement module can be introduced in low-light environments to improve the quality of information representation. Alternatively, a microphone array can be used to collect ambient sound and voice commands in real time as second-modal input information. An end-to-end acoustic feature encoder converts the audio signals into second-modal feature vectors representing the command content and environmental state. The audio encoding structure can incorporate noise robustness design to adapt to the voice information acquisition needs in high-noise industrial environments or complex indoor spaces. Furthermore, the timestamps during the acquisition of visual and audio information can be strictly aligned to ensure the synchronization of multimodal data and improve subsequent fusion results.
[0040] Example Explanation: In the field of embodied intelligence, visual information can be used to acquire image data of obstacles, tools, equipment, and dynamic environmental changes within the workspace, serving as the first modality input. This first modality feature vector is then generated through a visual feature extraction network, representing the environmental structure, object position, and shape information. Simultaneously, audio information, including human operation commands, environmental acoustic feedback, or equipment operating status signals, is acquired as the second modality input. This second modality feature vector, extracted by an audio encoder, represents the content of voice commands, environmental noise interference, and potential risk warnings. This joint representation of multimodal features enables embodied intelligence systems to accurately identify the work environment and operational requirements, enhancing their understanding of dynamic changes in complex scenarios. It is widely applied in tasks such as path planning for warehousing and logistics robots, dynamic obstacle avoidance for industrial collaborative robots, and fault perception and voice interaction control for autonomous inspection robots, significantly improving the decision-making flexibility, task completion efficiency, and human-machine collaboration level of embodied intelligent devices in real-world environments.
[0041] In the field of healthcare, visual information can be used to obtain the patient's location, posture, and environmental layout in the ward, while audio information can be used to obtain the patient's voice commands or abnormal sound changes. By jointly generating the first modality feature vector and the second modality feature vector, the medical robot's understanding of the environmental state and the patient's needs can be enhanced, thereby improving its responsiveness and service quality in tasks such as medical assistance, rehabilitation companionship, and intelligent monitoring.
[0042] In the fintech business, visual information can be used to collect images of the office environment or customers' faces, and audio information can be used to obtain the content of customers' voice communication and the background sound of the environment. Combined with multimodal feature vector expression, the environmental perception and human-computer interaction capabilities in scenarios such as intelligent customer service, automatic guidance, and financial information security monitoring can be improved, so as to achieve highly reliable human-computer collaboration and information response in financial scenarios.
[0043] This embodiment, by simultaneously acquiring input information from different modalities and efficiently extracting features from each modal information, can effectively integrate visual and audio information in the environment, solve the problem of isolated information expression, improve the system's overall perception capability of complex and changing environments, provide an accurate and comprehensive input foundation for subsequent multimodal fusion, decision generation, and action planning, and significantly enhance the device's adaptability and decision accuracy in dynamic environments.
[0044] S20, fuse the first modality feature vector and the second modality feature vector to generate a multimodal fusion feature;
[0045] In this embodiment, the first modality feature vector and the second modality feature vector are fused to generate a multimodal fusion feature. Specifically, this refers to combining, associating, or transforming data representation information from different modalities at the feature level to obtain a fusion representation containing multidimensional information. The first modality feature vector originates from the feature extraction process for the first modality input information in the preceding operations. The first modality input information can be visual perception data such as images, videos, and point clouds. The first modality feature vector is typically stored in the form of a multidimensional array, vector sequence, dense matrix, or sparse representation, expressing information such as the spatial structure, texture information, and target attributes of the visual content. The second modality feature vector originates from the feature extraction process for the second modality input information in the preceding operations. The second modality input information can be speech, audio, vibration signals, radar echoes, temperature sequences, text data, etc. The second modality feature vector typically expresses non-visual information, such as semantic content, acoustic features, and environmental changes, in the form of acoustic feature parameters, text semantic encoding, or multidimensional vector arrays. The fusion process can be achieved through various techniques such as concatenation, weighted superposition, cross-modal attention mechanisms, feature transformation networks, joint encoding networks, semantic alignment networks, and cross-modal contrastive learning structures. Specific operations include, but are not limited to, concatenating the feature vectors of the first and second modalities along their dimensions to form combined features, inputting features from different modalities into a unified fusion neural network, obtaining a fusion representation through multi-layer mapping transformation, or dynamically adjusting the weights and combination methods of feature fusion based on cross-modal correlations to ultimately generate multimodal fusion features. Multimodal fusion features have a more comprehensive environmental understanding capability, and can simultaneously express structural, semantic, behavioral, and environmental information from different information sources, thereby improving the system's overall perception and decision-making capabilities in complex environments.
[0046] An end-to-end multimodal fusion network structure can be adopted, where the first and second modal feature vectors are input into a unified multi-layer neural network. After a series of operations such as convolution, fully connected layers, normalization, and activation function transformation, a fixed-dimensional multimodal fusion feature is output. Attention mechanisms can also be used to enhance the fusion effect. Through self-attention or cross-attention structures, the correlation weights between different modal features are dynamically calculated, highlighting important information, suppressing invalid interference, and improving the effectiveness and robustness of the fusion representation. Furthermore, a joint contrastive learning strategy can be employed. By constructing positive and negative sample pairs, the representation space distribution of the multimodal fusion features is optimized, strengthening the consistency and complementarity of different modal information, further improving the accuracy and stability of the multimodal fusion representation. For different types of first and second modal input information, the fusion network structure, feature transformation methods, and fusion strategies can be adjusted according to specific application requirements to achieve flexible adaptation.
[0047] Example: In the healthcare field, multimodal fusion features can be generated by fusing visual modal feature vectors and audio modal feature vectors. This helps medical robots accurately identify patient body parts, instrument positions, and doctors' verbal instructions during surgery, ensuring the accuracy of the operation path and the real-time nature of the interactive response, reducing the risk of misoperation, and improving the safety and efficiency of the medical process.
[0048] In the field of fintech, multimodal fusion features can be generated by integrating the feature vectors of image information and voice commands. These features can be applied to scenarios such as smart teller machines and virtual customer service to achieve a joint understanding of user operation behavior, identity verification information and environmental audio signals, thereby improving the human-computer interaction experience and security level of financial business systems.
[0049] In the field of embodied intelligence, multimodal fusion features can be generated by integrating environmental visual features and voice information features. These features can be applied to autonomous inspection robots, intelligent delivery equipment, etc., to improve the equipment's ability to jointly perceive dynamic environments, operation commands, and obstacle information, thereby enhancing the flexibility of task execution, the intelligence of path planning, and the reliability of environmental adaptation.
[0050] This embodiment generates multimodal fusion features by fusing the first modality feature vector and the second modality feature vector. This effectively integrates heterogeneous features from different information sources within a unified expression space, making up for the limitations of single modality information, improving the system's overall perception capability and task decision-making level in complex environments, avoiding misjudgments, omissions, or decision deviations caused by incomplete information, and enhancing the accuracy of environmental understanding and the reliability of operational strategies.
[0051] S30, Generate action instructions based on the multimodal fusion features;
[0052] In this embodiment, generating action commands based on multimodal fusion features specifically refers to inputting the fused multimodal fusion features into the command generation structure. After processes such as feature parsing, semantic mapping, and action parameter inference, action commands that can directly guide the device to perform operations are output. Multimodal fusion features are typically high-dimensional feature representations that incorporate multi-source data such as visual, audio, environmental, and semantic information, possessing comprehensive information capabilities for environmental understanding, intent recognition, and behavior reasoning. Action commands are sets of control parameters specific to a particular device, typically including information such as spatial position, posture changes, motion trajectory, velocity parameters, acceleration parameters, action type, and device state configuration, used to drive the device to perform specific operations.
[0053] Multimodal fusion features are input into the action generation structure, which may include a multi-layer neural network, a logical decision unit, a parameter parsing module, or a hybrid control unit. First, the multimodal fusion features are received through the input layer, and preliminary transformations or dimensional mappings are performed on the feature information to ensure that the feature representation matches the parameter space of the subsequent computational structure. Subsequently, deep information extraction and multi-level feature transformations are performed on the fusion features through hidden layers to further enhance the high-order semantic information and the correlation between the fusion representation and environmental behavior.
[0054] After completing multi-layer feature processing, preliminary motion parameter results can be generated through the output layer. Motion parameters may include, but are not limited to, target position, joint angles, operation paths, and motion category labels. Furthermore, the joint angle sequence within the motion parameters can be parsed, and based on the device's structural characteristics and spatial layout, the corresponding motion trajectory path can be calculated, achieving spatial mapping and trajectory planning of the operation path.
[0055] Simultaneously, speed information can be extracted from motion parameters to generate speed control curves. These curves describe the speed variation patterns during operation, ensuring the smoothness, coordination, and safety of motion execution. The motion trajectory path and speed control curves are then integrated to form basic motion commands, which express the complete control logic of the operation path and speed changes.
[0056] To adapt to the operational needs of different types of devices, device type identifiers can be added to the basic action instructions. The device type identifiers can indicate the specific device or device category to which the control instructions are adapted, ensuring the correct issuance and execution of action instructions in a multi-device system. Finally, a complete action instruction containing the device type identifier is generated. The complete action instruction has the characteristics of complete structure, clear information, and device executableness, and can be directly transmitted to the device execution module to drive the device to complete the corresponding operation behavior.
[0057] An end-to-end action generation network structure can be adopted, which takes multimodal fusion features as input and outputs action parameters directly through a multi-layer neural network structure. The input layer adjusts the feature representation through linear transformation and normalization structure, the hidden layer uses multi-layer perceptron, convolutional structure or attention mechanism network to extract deep feature information, and the output layer combines multi-task learning strategy to output position parameters, pose parameters and operation category information respectively.
[0058] When analyzing motion parameters, the three-dimensional spatial motion trajectory corresponding to the joint angle sequence can be calculated based on the kinematic model and equipment structural parameters. Combined with inverse kinematics methods, this ensures the trajectory path meets the accessibility and continuity requirements of the equipment's mechanical structure. During the generation of speed control curves, methods such as fifth-order polynomial interpolation, trapezoidal velocity programming, and S-curve programming can be used to generate smooth speed change curves that conform to dynamic constraints, avoiding mechanical shocks or control errors caused by sudden speed changes.
[0059] After the basic action instructions are generated, the corresponding device type identifier is added according to the device's system configuration table or communication protocol specifications to distinguish devices of different models, functions or uses, ensuring the effective issuance of instructions and the ability to coordinate control of multiple devices.
[0060] Example: In the healthcare field, action commands generated based on multimodal fusion features can be applied to surgical robot control scenarios. By integrating visual information, voice commands, and environmental perception information, action commands tailored to specific patient and surgical needs can be generated, ensuring the accuracy of surgical operation paths and the smoothness of speed control, reducing surgical risks, and improving the reliability and safety of medical operations.
[0061] In the fintech business, action commands generated based on multimodal fusion features can be applied to devices such as smart teller machines and financial service robots. By integrating user identity information, operation voice, and environmental monitoring information, comprehensive action commands that include identity verification, interactive operation, and security control can be generated, thereby improving the human-computer interaction capabilities and security protection level of financial devices and optimizing financial service processes.
[0062] In the field of embodied intelligence, action commands generated based on multimodal fusion features can be applied to autonomous inspection robots, intelligent logistics equipment, etc. By integrating environmental visual information, voice commands and sensor data, action commands that adapt to complex and dynamic environments can be generated, ensuring that the equipment can navigate autonomously, avoid obstacles, and operate efficiently in changing scenarios, thereby improving task completion rate and environmental adaptability.
[0063] This embodiment generates action commands based on multimodal fusion features, which fully utilizes the fusion of different information sources to improve the environmental adaptability, task understanding, and execution accuracy of the action commands. Combining trajectory paths and velocity curves ensures a smooth and continuous operation process, reducing motion errors and structural impacts, and improving equipment execution efficiency and operational safety. By introducing device type identifiers, precise control and efficient collaboration within multi-device systems are achieved, enhancing the overall system's flexibility and task adaptability to meet diverse operational needs in complex environments.
[0064] S40, Based on the action command, the current state of the device, the current environmental information, and the task objective, generate an initial action plan;
[0065] In this embodiment, generating an initial motion plan based on action commands, the current state of the device, current environmental information, and task objectives refers to establishing a set of spatial paths, dynamic parameters, and behavioral strategies oriented towards operation execution by fusing control commands, device operating status, environmental data, and task requirements. This results in an operation execution plan with environmental adaptability, dynamic constraints, and task orientation. Action commands are a set of control parameters fused and parsed from multimodal information, including joint motion information, path planning parameters, velocity change curves, and device type identifiers, possessing the ability to directly guide the device in performing operations. The current state of the device refers to its physical state, kinematic state, or functional configuration before the motion plan is generated. It typically includes information such as the device's position, attitude, joint angles, velocity, acceleration, operating mode, and hardware configuration, reflecting the device's real-time operating capabilities and physical boundary conditions.
[0066] Current environmental information refers to external environmental data acquired in the operational scenario. Sources include devices such as visual sensors, distance sensors, LiDAR, and depth cameras. It encompasses obstacle locations, spatial layout, dynamic changes, and environmental risk information, reflecting external constraints and potential interference factors in the space where the equipment is located. The task objective refers to the operational requirements and task indicators that the equipment must currently complete. It is typically expressed in the form of target location, target action type, task completion time, path preference, resource consumption limitations, and environmental safety requirements. It directly determines the task orientation and performance optimization standards of the action planning.
[0067] During the initial motion planning process, the environmental perception module collects current environmental information, and multi-source fusion technology is used to integrate the environmental data and model its spatial structure, generating environmental feature representations to describe the environmental spatial structure, obstacle distribution, and dynamic risk factors. Combined with the equipment's current state data, kinematic calculations and physical constraint analysis are used to obtain the equipment's reachability, motion limitations, and dynamic response characteristics, ensuring the feasibility of subsequent planning results and equipment safety.
[0068] The motion commands, current equipment state, environmental features, and task objectives are input into the environment modeling module to establish a comprehensive environment model based on the collaboration of equipment, environment, and task. This model comprehensively considers the impact of multi-source information on motion planning, resulting in an accurate and complete representation of the operating environment. By performing kinematic calculations on the current equipment state, real-time calculated state parameters are obtained, including joint angles, end effector positions, attitudes, reachability, and redundancy space, thereby improving the spatial adaptability and dynamic adjustment capabilities of motion planning.
[0069] Based on environmental feature representation, a three-dimensional spatial model of the environment is constructed. This model expresses the locations of obstacles, spatial passages, operable areas, and dynamic changes within the environment, providing the spatial basis for motion path planning. Within the three-dimensional spatial model, the feasible domain of the motion path is determined according to the equipment's reachability and dynamic constraints, eliminating obstacle interference and hazardous environmental areas to ensure the safety and feasibility of the planned path.
[0070] By combining the environmental model, solved state parameters, and feasible regions of action paths, the data is input into the constraint optimization module. This module performs path optimization, dynamic constraint adjustment, and performance trade-off analysis to generate an optimized action plan that possesses spatial reachability, dynamic continuity, and task orientation. The optimized action plan is then used as the initial action plan. This initial action plan exhibits comprehensive characteristics of environmental adaptability, equipment feasibility, and task satisfaction, providing fundamental support for the subsequent generation of the globally optimal action sequence.
[0071] Initial motion planning can be generated through a structure combining multi-layer neural networks and optimization algorithms. The environmental perception module can employ various devices such as depth cameras, LiDAR, and vision sensors to acquire spatial information in real time, and then use convolutional neural networks or graph neural networks to extract features and represent spatial structure from the environmental data. The current state data of the device is acquired through real-time sensors, encoders, and inertial measurement units, and coordinate transformation and dynamic parameter calculation are performed by the kinematics solution module.
[0072] The environment modeling module can employ multi-source data fusion technology to integrate target information from action commands, equipment status data, environmental feature representations, and task objectives, generating a unified spatial environment representation that supports dynamic changes and real-time model updates in complex environments. During kinematic calculations, forward and inverse kinematics methods combined with redundancy optimization techniques can enhance equipment spatial operability and path planning flexibility.
[0073] The construction of a 3D spatial model can employ voxel mesh modeling, point cloud reconstruction, or depth map fusion methods to represent the environmental spatial structure, obstacle locations, and operational space boundaries. During the determination of the feasible region for action paths, path search algorithms, spatial constraint detection, and dynamic risk assessment can be combined to eliminate inaccessible areas and potentially dangerous paths.
[0074] The constraint optimization module can generate an optimized action plan that balances multiple indicators by comprehensively considering path length, energy consumption, time efficiency, and operational safety, based on multi-objective optimization algorithms, constraint condition screening and dynamic adjustment mechanisms. After logical screening and parameter correction, the optimization results form an initial action plan that is executable and environmentally adaptable.
[0075] Example: In the healthcare field, initial motion planning can be generated based on action commands, current equipment status, current environmental information, and task objectives. This can be applied to intelligent rehabilitation robot systems. By integrating the patient's real-time posture, operating environment layout, and rehabilitation task requirements, highly adaptable and safe rehabilitation operation plans can be dynamically generated, ensuring the safety and personalization of the rehabilitation operation process and improving rehabilitation outcomes and medical efficiency.
[0076] In the field of fintech, it can be applied to the control process of self-service financial equipment and intelligent interactive terminals. By integrating user instructions, the current working status of the equipment, the ATM operating environment and service task requirements, it can dynamically plan the equipment's action path and operation strategy, thereby improving operational efficiency, optimizing user experience and ensuring the security of financial services.
[0077] In the field of embodied intelligence, initial motion planning is generated based on action commands, current equipment status, current environmental information, and task objectives. This can be applied to autonomous logistics robots, intelligent inspection equipment, etc. By integrating equipment control commands, spatial layout information, environmental dynamic data, and task requirements, efficient, continuous, and safe operation paths are dynamically planned. This ensures the equipment's autonomous navigation, task execution, and efficient operation capabilities in complex dynamic environments, and enhances the embodied intelligence system's adaptability and overall operational efficiency in changing environments.
[0078] This embodiment generates initial action plans based on action commands, current device state, current environmental information, and task objectives. This enables multi-source information fusion, dynamic environment modeling, and real-time task adaptation, improving the environmental adaptability of action planning, device execution reliability, and task completion efficiency. Through environmental perception and 3D modeling, it comprehensively expresses the operational space structure and dynamic risk information. Combined with kinematic calculations and path optimization, it generates safe, continuous, and task-oriented operation paths, avoiding path interference, execution conflicts, and task deviations. This enhances the stability, flexibility, and efficiency of the operating system in complex environments.
[0079] S50, Generate the globally optimal action sequence based on the initial action planning;
[0080] In this embodiment, generating a globally optimal action sequence based on initial action planning involves further optimizing the action sequence by using multi-path candidate selection, expected evaluation, and dynamic optimization, based on existing spatial paths, dynamic parameters, and task strategies. This process improves the global optimality of action execution and the system's adaptability. Initial action planning, as a dynamic path and strategy set integrating action instructions, equipment status, environmental information, and task objectives, provides the basic expression of operation paths, motion constraints, and task objectives. Therefore, the process of generating a globally optimal action sequence typically includes candidate action sequence generation, environmental state updating, candidate sequence performance evaluation, and optimal sequence selection.
[0081] Candidate action sequence generation refers to generating multiple action sequences with different path structures, dynamic parameters, or behavioral strategies based on the initial action planning, through strategy expansion, path transformation, or parameter adjustment. This expands the operational space and provides diverse action execution solutions. These action sequences are structurally based on the initial action planning framework, but differ in path, speed, energy consumption, or dynamic response, aiming to provide a basis for comparing multiple solutions.
[0082] During the environmental status update process, the environmental model, spatial structure representation, and task requirements parameters are dynamically adjusted by combining the real-time acquired current environmental information, equipment status, and task objectives. This ensures that the environmental conditions for candidate action sequence evaluation remain consistent with the actual operating environment, thereby improving the reliability and adaptability of the evaluation results.
[0083] Candidate sequence performance evaluation is conducted by setting expected reward values and combining environmental information, equipment status and task objectives to comprehensively evaluate each candidate action sequence in multiple dimensions such as task completion efficiency, path safety, energy consumption and time optimization, and generate corresponding expected reward values that reflect the overall performance and execution expectations of the action sequence.
[0084] Finally, by comparing the expected reward values of all candidate action sequences, the action sequence with the highest reward value is selected as the globally optimal action sequence. This ensures that the operating system selects the overall optimal execution plan among multiple paths and multiple strategy schemes, thereby improving operational efficiency, environmental adaptability, and task completion results.
[0085] Globally optimal action sequences can be dynamically generated by combining deep reinforcement learning, dynamic programming, or multi-objective optimization techniques with real-time environmental perception and device status monitoring. During the generation of candidate action sequences, path transformation, dynamic parameter adjustment, or policy network inference can be employed to generate diverse action sequences with differences in path structure and dynamic behavior, thus expanding the solution space.
[0086] Environmental status updates can rely on devices such as visual sensors, LiDAR, depth cameras, and inertial measurement units to acquire real-time information on dynamic environmental changes, obstacle status, and spatial layout adjustments. Combined with sensor data, the environmental model and task requirements can be dynamically updated to ensure the real-time performance and accuracy of environmental assessments.
[0087] During the performance evaluation process, a multi-dimensional reward function can be set to comprehensively score candidate action sequences by combining factors such as path length, operational safety, time efficiency, energy consumption, and task completion, thereby evaluating the execution effect of the sequence and the system performance.
[0088] When selecting the optimal action sequence, the optimal action sequence can be selected based on the maximum reward strategy, weighted multi-objective balance, or dynamic adjustment mechanism. This optimal action sequence can then be used as the final globally optimal action sequence to guide the equipment execution and improve the global optimality and environmental adaptability of the system operation.
[0089] Example: In the healthcare field, generating globally optimal motion sequences based on initial motion planning can be applied to the personalized training of rehabilitation robots. By generating and evaluating multiple motion sequences, the rehabilitation operation path and parameters can be dynamically optimized to adapt to changes in patient posture and environmental adjustments, thereby improving rehabilitation effectiveness and operational safety, and enhancing the intelligent adaptability of the rehabilitation system.
[0090] In the field of fintech, it can be applied to the operation control of smart teller equipment and self-service terminals. By generating and optimizing multi-path action sequences, it can dynamically adjust the equipment operation strategy and path layout to adapt to differences in user operations and changes in environmental layout, thereby improving operational efficiency and ensuring equipment stability and financial business continuity.
[0091] In the field of embodied intelligence, it can be applied to the path planning and motion control of autonomous mobile robots and intelligent operating equipment. By generating diverse candidate action sequences based on the initial motion planning, and combining real-time environmental data and equipment status, the optimal execution plan can be dynamically selected, thereby improving the global path optimization capability and dynamic adaptability of the embodied intelligence system under complex environments and multi-task requirements, and ensuring the system's efficient, safe and intelligent operation capability in changing environments.
[0092] This embodiment generates a globally optimal action sequence based on initial action planning. Through multi-scheme generation, dynamic environment updates, and multi-dimensional performance evaluation, it comprehensively improves the global optimality, environmental adaptability, and task completion efficiency of action execution schemes. It overcomes the limitations of a single planning scheme, avoids the trap of local optima, and ensures that the operating system selects the overall optimal action execution scheme under complex environments and multi-task requirements, thereby improving the flexibility, reliability, and intelligence level of the operating system.
[0093] S60, control the device to execute the globally optimal action sequence.
[0094] In this embodiment, the control device executes a globally optimal action sequence, which involves driving the actuators to complete corresponding physical actions based on the optimized action sequence, ensuring the accuracy, stability, and consistency of the action execution with the task objectives. The globally optimal action sequence refers to a set of action parameters optimized based on the initial action plan, combined with environmental information, device status, and task objectives. It typically includes spatial path information, dynamic speed parameters, action timing control data, and device operation commands. This sequence not only reflects the operation path and timing but also comprehensively considers device structure, task requirements, and environmental constraints, ensuring the optimality of the executed actions and the overall coordination of the system.
[0095] During control execution, the globally optimal action sequence is first parsed into a structured action instruction set, clearly defining each execution unit, execution step, action parameter, and time control node. The action instruction set can cover information such as equipment joint angles, movement paths, operating speeds, dynamic accelerations, and timing trigger conditions, ensuring clear instructions and complete data, facilitating subsequent identification and execution by the control unit.
[0096] After parsing, the action instruction set is converted into low-level drive signals by the device controller. Depending on the type of device and the control system, the drive signals usually include voltage control signals, current commands, pulse signals or digital control codes, which are used to directly control motors, hydraulic systems, servo devices or other execution modules to complete the physical realization of the action.
[0097] After the drive signal is sent to the actuator, the actuator performs specific mechanical actions, spatial position changes, or equipment function operations according to the received control signal, realizing the action execution process within the physical space. During the execution process, the execution status parameters are monitored synchronously to obtain the real-time position, velocity, acceleration, attitude, or other dynamic information of the actuator, ensuring that the action execution process is consistent with the globally optimal action sequence, promptly detecting deviations or anomalies, and ensuring operational safety and system stability.
[0098] By combining an industrial robot control system, automated equipment control unit, or intelligent execution platform with a globally optimal motion sequence generation module, structured parsing of motion sequences and generation of control instructions can be achieved. During the motion instruction set parsing process, a multi-level instruction parsing mechanism can be adopted to decompose the complex globally optimal motion sequence into specific instructions for each execution unit and distribute them to different execution mechanisms, thereby improving system control efficiency and module collaboration capabilities.
[0099] During the control signal conversion process, servo drive technology, bus communication protocols, or wireless control systems can be combined to generate standardized drive signals that are compatible with different execution devices, ensuring the real-time performance, accuracy, and device compatibility of signal transmission.
[0100] The actuator may include a multi-degree-of-freedom robotic arm, a mobile chassis, an end effector, a hydraulic system, or other intelligent devices. After the control signal is transmitted, the actuator completes the corresponding physical action according to the instruction, monitors the status parameters in real time, and dynamically corrects the execution process by combining sensor data, vision system, or position feedback device to ensure the accuracy of action execution and system safety.
[0101] Example: In the medical and health field, the control device executes the globally optimal action sequence, which can be applied to auxiliary operating systems such as rehabilitation training robots and exoskeleton devices. The driven device completes the precise guidance and dynamic assistance of the patient's limbs according to the optimized action sequence, improves the effectiveness and safety of rehabilitation training, and enhances the system's adaptability and operational stability in the personalized rehabilitation process.
[0102] In the field of fintech, it can be applied to the mechanical operation unit control of self-service terminals and smart teller machines. Based on the globally optimal action sequence, it controls robotic arms, transmission mechanisms or other equipment to complete operations such as financial card transmission, document printing, and currency deposit and withdrawal, thereby improving the accuracy and efficiency of equipment operation and ensuring the continuity of financial services and user service experience.
[0103] In the field of embodied intelligence, the control of devices to execute globally optimal action sequences can be applied to systems such as mobile robots and autonomous operating equipment. Based on the optimized action sequences, the devices are driven to complete complex operations such as autonomous navigation, environmental adaptation, and task execution, thereby improving the execution accuracy, operational safety, and overall intelligence level of embodied intelligence systems in dynamic environments, and enhancing the system's autonomous operation capabilities and efficient collaborative capabilities in changing environments.
[0104] This embodiment controls the device to execute the globally optimal action sequence, which can transform the optimized action sequence into an efficient and stable physical action execution process. This ensures that the path, speed, and timing of the device operation are highly matched with the task requirements, improves the accuracy, coordination, and intelligence of the system's action execution, reduces operational errors, and enhances the system's stability and task completion effect in complex environments.
[0105] This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as embodied intelligence, fintech, and healthcare. It discloses a multimodal fusion feature-driven action control method, device, equipment, and medium, comprising: acquiring first modal input information and second modal input information; generating a first modal feature vector based on the first modal input information; generating a second modal feature vector based on the second modal input information; fusing the first and second modal feature vectors to generate a multimodal fusion feature; generating action commands based on the multimodal fusion feature; generating an initial action plan based on the action command, the current state of the device, current environmental information, and the task objective; generating a globally optimal action sequence based on the initial action plan; and controlling the device to execute the globally optimal action sequence. This invention generates an initial action plan by combining the action command generated from the multimodal fusion feature with the device state, environmental information, and task objective, and optimizes the initial action plan to generate a globally optimal action sequence for device execution. This improves the decision-making performance of intelligent entities in complex environments, achieves efficient fusion of multimodal information, enhances the globality and environmental adaptability of action planning, avoids local optima problems, and improves the accuracy and flexibility of overall task execution.
[0106] In one embodiment, step S10 above includes:
[0107] S101, Obtain visual modal input information as the first modal input information;
[0108] S102, the first modal input information is processed by a visual encoder to generate coded visual features, and the coded visual features are used as the first modal feature vector;
[0109] S103, acquire audio modal input information as the second modal input information;
[0110] S104, the second modal input information is processed by an audio encoder to generate encoded audio features, and the encoded audio features are used as the second modal feature vector.
[0111] In this embodiment, acquiring visual modal input information as the first modal input information refers to obtaining visual information about the environment, target objects, or scenes through image sensors, cameras, optical systems, or other devices capable of acquiring image data. The visual information may include color images, grayscale images, depth images, or other data based on electromagnetic wave imaging. Color images may use RGB format, grayscale images may use single-channel matrix data, and depth images may use data collected by TOF sensors, structured light sensors, or lidar to represent spatial distance information. Visual information can be acquired through camera equipment installed at fixed locations, or dynamically through imaging devices mounted on mobile devices, robots, or wearable devices, ensuring the system can perceive environmental changes or target features in real time.
[0112] The visual encoder processes the first modality input information to generate encoded visual features. This involves using a specially designed neural network, convolutional neural network (CNN), visual Transformer model, or other deep learning structure to perform multi-level feature extraction, compression, and expression optimization on the acquired visual data. The visual encoder generates high-dimensional dense features or low-dimensional abstract expressions based on the spatial structure, color distribution, texture information, edge features, and other content of the image. The encoded visual features can be expressed in the form of vectors, matrices, or tensors, which facilitates subsequent information fusion and computational processing.
[0113] Using encoded visual features as the first modality feature vector means directly using the structured data output by the visual encoder as the first modality feature representation of the system. The first modality feature vector retains the core content of visual information, including spatial information, object category, positional relationship or environmental state, etc., providing visual data support for multimodal information fusion and improving the comprehensiveness and accuracy of the system's environmental perception.
[0114] Acquiring audio modal input information as the second modal input information refers to collecting sound signals from the environment through devices such as microphones, pickups, and audio sensors. Audio information can include voice commands, environmental noise, structural sound source signals, or other audio data that helps the system understand and make judgments. Audio input can be mono, stereo, or multi-channel array signals. Combined with microphone array technology, spatial localization, noise suppression, and sound source separation can be achieved, improving the quality of audio data and the system's perception capabilities.
[0115] The audio encoder processes the second modality input information to generate coded audio features. This involves using audio processing networks, temporal convolutional networks, recurrent neural networks, Transformer structures, or other deep learning architectures to perform time-series analysis, feature extraction, and signal representation optimization on the acquired audio data. The audio encoder can combine spectrogram analysis, short-time Fourier transform, Mel-spectral transform, or other acoustic feature processing methods to transform the original audio signal into a high-dimensional representation. The coded audio features can reflect semantic information, audio signal structure, rhythmic variations, or other attributes that help the system understand it.
[0116] Using encoded audio features as second modality feature vectors means directly using the data output by the audio encoder as the second modality feature expression of the system. The second modality feature vector reflects the temporal dynamics, semantic information, or environmental change status of the audio signal, providing an audio-level data foundation for multimodal information fusion and subsequent decision-making, and enhancing the system's information acquisition capability and responsiveness in complex environments.
[0117] This embodiment acquires visual and audio modal input information and generates first and second modal feature vectors based on the visual encoder and audio encoder, respectively. This enables efficient acquisition and representation of multi-source information. The independent processing of visual and audio information ensures the accuracy of their respective feature representations. The generation process of encoding visual and audio features improves the uniformity of the data structure and the efficiency of subsequent fusion. Overall, it enhances the system's ability to understand the environment, targets, and task instructions, improves the comprehensiveness of information perception and the level of multi-dimensional representation, optimizes the downstream multi-modal information fusion and intelligent decision-making process, and further enhances the system's adaptability and execution reliability in dynamic and complex environments.
[0118] In one embodiment, step S30 above includes:
[0119] S301, Input the multimodal fusion features into the action generator;
[0120] S302, The multimodal fusion features are processed through the input layer of the action generator to generate a primary feature representation;
[0121] S303, The primary feature representation is processed through the hidden layer of the action generator to generate a high-level feature representation;
[0122] S304, The high-level feature representation is processed through the output layer of the motion generator to generate joint motion parameters;
[0123] S305, parse the joint angle sequence in the joint motion parameters, and determine the motion trajectory path based on the joint angle sequence;
[0124] S306, Extract the joint velocity parameters from the joint motion parameters, and generate a velocity control curve based on the joint velocity parameters;
[0125] S307, The motion trajectory path and speed control curve are combined to form a basic motion command;
[0126] S308, Add a device type identifier to the basic action instruction, and generate a complete action instruction containing the device type identifier.
[0127] In this embodiment, the multimodal fusion features are input into the action generator. This means that the multimodal fusion features obtained through the multimodal information fusion module are used as a unified information representation and passed to the action generator, which is specifically designed to generate control commands. The multimodal fusion features contain comprehensive expressive information from visual, audio, or other perceptual sources, possessing high-dimensional expressive power and semantic richness. The action generator can be a deep neural network, a recurrent network, a transformer structure, or a combined multilayer network, possessing the ability to map complex features to specific control commands.
[0128] The input layer of the action generator processes multimodal fusion features to generate primary feature representations. This means that after receiving multimodal fusion features, the input layer of the action generator performs a first-stage structural transformation and information abstraction on high-dimensional, multi-source information based on fully connected networks, convolutional networks, or other feature mapping mechanisms to generate primary feature representations. The primary features retain the key expressive content of multimodal information and provide a unified feature foundation for subsequent deep information processing through dimensionality reduction, feature selection, or embedding operations.
[0129] The action generator processes primary feature representations through hidden layers to generate advanced feature representations. This means that in the hidden layer structure of the action generator, the primary features are further subjected to multi-layer nonlinear transformations, feature extraction, and information interaction to generate advanced feature representations with higher semantic expressive power and control relevance. The hidden layer can include multilayer perceptrons, attention mechanisms, autoencoders, or other network modules used for feature enhancement and complex information modeling to ensure that the system can extract deep expressions directly related to action control from the fused information.
[0130] The output layer of the motion generator processes high-level feature representations to generate joint motion parameters. This means that the output layer of the motion generator maps high-level features to joint motion parameters required for the control device to perform actions. Joint motion parameters include, but are not limited to, joint angles, joint velocities, joint accelerations, position coordinates, or attitude information. The output layer can adopt a fully connected structure, linear transformation, or distributed representation to ensure that the output parameters meet the input requirements of the control system or device actuator.
[0131] Analyzing the joint angle sequence in the joint motion parameters and determining the motion trajectory path based on the joint angle sequence means that after the system acquires the joint motion parameters, it calculates the motion trajectory path in the spatial coordinate system based on the joint angle sequence information, combined with the structural parameters and kinematic model of the equipment. The motion trajectory path reflects the sequence of position changes that the equipment should follow during execution, ensuring the spatial accuracy and continuity of the action execution. The calculation of the trajectory path can be based on forward kinematics, inverse kinematics, path planning algorithms, or simulation prediction technology.
[0132] Extracting joint velocity parameters from joint motion parameters and generating velocity control curves based on these parameters involves extracting parameter information describing the dynamic response speed of the joint from the joint motion parameters, and generating velocity control curves that reflect the changing trend of the equipment's motion speed through time series analysis, curve fitting, or dynamic control models. The velocity control curves are used to guide speed adjustment and dynamic smoothing during the execution of actions, ensuring the stability, efficiency, and operational safety of the equipment's motion process.
[0133] The basic motion command is formed by integrating the obtained spatial motion trajectory path and speed control curve. This means integrating the spatial motion trajectory path and the speed control curve in the time dimension, and generating basic motion commands for actual control execution through data structure combination, synchronization strategy or spatiotemporal mapping. The basic motion command includes a joint expression of position and speed, which can comprehensively reflect the path, speed and timing requirements of the equipment in the motion execution process, and ensure the coordination and controllability of the motion process.
[0134] Adding device type identifiers to basic action instructions generates complete action instructions that include device type identifiers. This means adding identifier information based on the specific device type, model, structural characteristics, or operating specifications involved in the control system, on the basis of basic action instructions. Device type identifiers can be coded labels, structural parameters, model information, or other control auxiliary information used to distinguish different devices. By carrying device type identifiers, complete action instructions ensure the adaptability and compatibility of instructions in multiple devices, heterogeneous systems, or complex environments, avoid misoperation and system conflicts, and improve the flexibility and reliability of the overall control system.
[0135] This embodiment inputs multimodal fusion features into the motion generator, which then processes the structured information through the input layer, hidden layer, and output layer. Combined with the analysis of joint angle sequences and joint velocity parameters, the system can accurately generate basic motion commands containing spatial position and dynamic velocity information. Furthermore, by adding device type identifiers, it ensures that the generated complete motion commands have adaptability and execution accuracy in multi-device environments. Overall, it improves the quality of motion command generation and device control effect driven by multimodal information, and enhances the system's multimodal understanding, dynamic control, and multi-device collaboration capabilities in complex environments.
[0136] In one embodiment, step S40 above includes:
[0137] S401 collects current environmental information through the environmental perception module;
[0138] S402, extract features from the current environmental information to generate an environmental feature representation;
[0139] S403, Obtain the current status of the device and the task objective;
[0140] S404, input the action command, the current state of the device, the environmental feature representation and the task objective into the environment modeling module, and construct an environment model through the environment modeling module;
[0141] S405, Perform kinematic calculation on the current state of the device to obtain the calculated state parameters;
[0142] S406, Based on the environmental feature representation, establish a three-dimensional spatial model of the environment, and determine the feasible domain of action paths in the three-dimensional spatial model of the environment;
[0143] S407, the environment model, the solution state parameters and the feasible domain of the action path are input into the constraint optimization module, the constraint optimization module generates an optimized action plan, and the optimized action plan is used as the initial action plan.
[0144] In this embodiment, the current environmental information is collected through the environmental perception module. This means that the environmental perception module, which is configured on the device or external system, is used to obtain environmental information related to the execution of actions within the spatial area where the device is located. The environmental information includes, but is not limited to, scene structure data, obstacle distribution information, spatial layout information, temperature and humidity parameters, or other environmental factors that can affect the execution of device actions. The environmental perception module may include visual sensors, LiDAR, depth cameras, ultrasonic sensors, temperature and humidity sensors, or multi-sensor fusion systems. The collection process is realized through data acquisition interfaces, real-time monitoring mechanisms, or sensor networks to ensure that the acquired environmental information is timely and accurate.
[0145] The process of extracting features from the current environmental information and generating an environmental feature representation refers to the system using feature extraction algorithms or models to extract a structured representation that reflects the key features of the environment from the aforementioned acquired environmental information. The environmental feature representation includes spatial structural features, obstacle location features, dynamic change information, or other environmental elements directly related to equipment action planning. The feature extraction process can be implemented through convolutional neural networks, point cloud processing algorithms, 3D reconstruction technology, or other feature encoding methods. The environmental feature representation is used for subsequent environmental model construction and action path planning, and has the characteristics of concise expression, complete information, and high correlation with the environmental state.
[0146] Acquiring the current status and task objectives of the equipment refers to the system collecting various parameter information related to the operating status of the equipment itself, and combining it with task objective information input by the task system or issued by the control system. The current status of the equipment includes the equipment position, attitude, speed, joint angle, power system status, or other data that can reflect the real-time operating status of the equipment. The task objectives include the specific task type, task content, task parameters, or constraints that the equipment needs to complete. The current status of the equipment is obtained through built-in sensors, control interfaces, or real-time feedback systems, while the task objectives are obtained through task scheduling systems, instruction input modules, or external system instructions, ensuring that the system has complete equipment operating information and task execution requirements.
[0147] The action commands, current device status, environmental feature representation, and task objectives are input into the environment modeling module. The environment modeling module then constructs an environment model. Based on the aforementioned multidimensional information, the environment modeling module performs information fusion, environmental structure reconstruction, and dynamic relationship reasoning to generate an environment model. The environment model is a mathematical or data expression describing the relationship between the device and the environment, the spatial structure layout, and the association of dynamic elements. The environment modeling module may include a spatial mapping system, a dynamic scene analysis module, a relationship reasoning network, or other structures capable of constructing a multidimensional environmental expression. The environment model provides a structured data foundation for subsequent path planning, action adjustment, and dynamic obstacle avoidance.
[0148] The system performs kinematic calculations on the current state of the equipment to obtain the calculated state parameters. This involves the system using the equipment's structural parameters, joint configuration, dynamic characteristics, and real-time state data to perform kinematic model inference and calculations. The results are used to obtain the specific position, attitude, velocity, or other dynamic parameters of each component of the equipment in the current state. The calculated state parameters include forward kinematics results, inverse kinematics solutions, redundancy calculations, or operational space mapping results. The kinematic calculations are implemented through analytical models, numerical algorithms, or data-driven models to ensure that the system can accurately grasp the current physical state and operational boundaries of the equipment.
[0149] The establishment of a three-dimensional spatial model of the environment based on environmental feature representation, and the determination of the feasible domain of action paths in the three-dimensional spatial model, refers to the construction of a three-dimensional spatial model reflecting the spatial structure and layout of the environment by combining the aforementioned environmental feature representation with spatial structure data and obstacle information. The three-dimensional spatial model of the environment includes spatial boundaries, obstacle locations, passable areas or other spatial constraint information. Based on this model, the system analyzes the feasible path areas of the equipment in space and determines the feasible domain of action paths. The feasible domain of action paths is used to describe the spatial range in which the equipment can perform actions without violating environmental constraints, ensuring the safety and spatial rationality of path planning.
[0150] The environmental model, solution state parameters, and feasible domain of action paths are input into the constraint optimization module. The constraint optimization module generates an optimized action plan, which is then used as the initial action plan. This means that the system integrates the environmental model, equipment state, and spatial feasible domain information, and uses the constraint optimization module to perform multi-objective constraint solving and action path optimization to generate an optimized action plan that meets the task requirements, environmental constraints, and equipment capability boundaries. The optimized action plan includes equipment paths, action sequences, speed parameters, and timing arrangements. The constraint optimization module can be implemented using mathematical programming algorithms, path optimization models, dynamic constraint adjustment mechanisms, or other structures used for action planning optimization. Finally, the optimized action plan is determined as the initial action plan and serves as the input basis for generating the subsequent globally optimal action sequence.
[0151] This embodiment collects environmental information and performs feature extraction through the environmental perception module, enabling the system to comprehensively grasp the external environmental state. Combining the current state of the equipment with the task target input, a structured environmental model is constructed using the environmental modeling module. Dynamic parameters of the equipment are obtained through kinematic calculations, and the feasible domain of the action path is determined by combining the three-dimensional spatial model of the environment. Furthermore, the constraint optimization module performs path optimization and action sequence planning, ultimately generating an initial action plan. Overall, it realizes multi-source information fusion, dynamic environmental modeling, action path spatial constraint analysis, and action planning optimization based on dynamic constraints. This improves the equipment's dynamic adaptability in complex environments, the flexibility of action execution, and the overall optimal level of path planning, ensuring that the equipment can safely, efficiently, and accurately complete the task target.
[0152] In one embodiment, step S50 above includes:
[0153] S501, Generate a set of candidate action sequences based on the initial action planning;
[0154] S502, The current environment information is processed by the feature extraction module to generate an environment feature representation;
[0155] S503, for each candidate action sequence in the candidate action sequence set, determine the expected reward value based on the environmental feature representation, the current state of the device, and the task objective;
[0156] S504. Compare the expected reward values of all candidate action sequences and select the candidate action sequence with the largest expected reward value as the globally optimal action sequence.
[0157] In this embodiment, the generation of a candidate action sequence set based on the initial action plan refers to the system reasoning or simulating various different action execution schemes based on the previously generated initial action plan, combined with the path information, dynamic parameters and spatial feasible domain in the action plan, forming a set of multiple candidate action sequences. The candidate action sequence set is used to provide diversified action execution paths and strategy selections. Each action sequence in the set corresponds to a set of action parameter combinations and time sequences that can be actually executed, and has a certain degree of difference and environmental adaptability. The generation process can adopt an action reasoning network, path mutation algorithm or sequence combination mechanism based on a sample library.
[0158] The feature extraction module processes the current environmental information to generate an environmental feature representation. This means that the system uses the algorithm structure in the feature extraction module to extract feature information reflecting key environmental states, obstacle distribution, spatial layout, or dynamic changes from the real-time collected environmental information. The environmental feature representation expresses the external environmental state in a structured and highly condensed data form. The feature extraction module can use image recognition networks, 3D point cloud processing algorithms, semantic segmentation models, or multimodal perception systems to ensure the accuracy and real-time performance of the environmental feature representation. The environmental feature representation is used for subsequent action sequence effect evaluation and reward value calculation.
[0159] For each candidate action sequence in the candidate action sequence set, the system jointly analyzes the execution effect of each action sequence under the current environment and equipment conditions based on environmental feature representation, current equipment state, and task objective. It comprehensively considers path feasibility, environmental adaptability, action efficiency, and task achievement degree, and uses an effect evaluation model or reward calculation function to quantify the expected reward value of each action sequence. The expected reward value reflects the overall quality of the action sequence. The calculation process fully combines environmental state, real-time equipment capabilities, and task objective constraints to ensure that the reward value evaluation results are scientific and have decision-making reference value.
[0160] The system compares the expected reward values of all candidate action sequences, sorts them according to their reward values, and prioritizes selecting the candidate action sequence with the highest reward value. The sequence with the highest reward value represents the best action execution plan under the current environment, equipment status, and task objective. This sequence is determined as the globally optimal action sequence, and the system executes subsequent control and action plans accordingly. The selection process can use sorting algorithms, maximum value retrieval, or optimal selection mechanisms based on policy networks to ensure that the finally determined globally optimal action sequence has the maximum effect of execution safety, environmental adaptability, and task completion.
[0161] This embodiment generates a set of candidate action sequences based on initial action planning, which expands the diversity of action execution strategies. By combining real-time environmental feature information and device status, it can comprehensively evaluate the environmental adaptability and task execution effect of each action sequence. Using the expected reward value as a quantitative standard, it effectively compares the merits of different sequences and finally selects the optimal sequence as the global action plan. This improves the system's action decision-making level under complex environments and dynamic task conditions, enhances the device's dynamic adaptive capability and global optimal task completion efficiency, avoids the technical defects of traditional single-path strategies that are prone to getting trapped in local optima, and ensures the efficient unity of global optimization of device action strategies and environmental adaptation.
[0162] In one embodiment, step S60 above includes:
[0163] S601, the globally optimal action sequence is parsed into a time-series action instruction set;
[0164] S602, the timing action instruction set is converted into drive signals through the device controller;
[0165] S603, the drive signal is sent to the execution device to drive the execution device to perform a physical action;
[0166] S604, monitor the real-time motion parameters of the execution device as execution status parameters.
[0167] In this embodiment, the globally optimal action sequence is parsed into a time-series action instruction set. This means that after the system generates the globally optimal action sequence, each action node is further divided into temporal segments and parameter decompositions based on the action nodes, action parameters, and execution order contained in the action sequence, forming a discrete action instruction set arranged in chronological order. The time-series action instruction set includes one or more action execution units ordered by time. Each action execution unit contains basic control parameters such as the displacement parameters, angle parameters, execution speed, and motion duration of the target device. This parsing process can be completed based on action sequence decomposition algorithms, path time mapping rules, or discretization functions based on motion trajectories, ensuring that the continuity of the action sequence is reasonably transformed into time-series control instructions that can be parsed and executed in real time by the device controller.
[0168] The device controller converts the timing action instruction set into drive signals. This means that the parsed timing action instruction set is input into the device controller. The device controller calculates the corresponding electrical signals, current waveforms, or control pulses based on the action parameters in the instruction set. The control algorithm completes the conversion from high-level action instructions to low-level drive signals. Drive signals may include PWM pulses, motor voltage, current intensity, or servo control parameters. The device controller can adapt different signal output formats and electrical standards according to the device type to ensure that the instruction information is accurately converted within the controller and transmitted to the execution device. The drive signals match the electrical response parameters of the device to achieve precise device response.
[0169] Sending drive signals to the actuators drives them to perform physical actions. This means that the device controller outputs the generated drive signals to the connected actuators in real time, such as motors, robotic arms, servo motors, and mobile chassis. The drive signals directly act on the drive units inside the actuators, controlling them to complete specific physical actions such as joint rotation, end effector displacement, gripper opening and closing, and path movement. The drive process is synchronized with the time step of the timing action instruction set to ensure that the actions are executed accurately according to the predetermined rhythm, path, and parameters. The physical action results of the actuators respond in real time to changes in the drive signals, forming a continuous process of device movement.
[0170] Monitoring the real-time motion parameters of the executing device as execution status parameters refers to the continuous collection of motion status data of the device during the execution of physical actions by the system through monitoring components such as motion sensors, position detection modules, angle encoders, and speed measurement units installed on the device body or moving structure. These data include parameters such as real-time position, joint angle, execution speed, and acceleration. The collected motion parameters are fed back to the system data path in real time as the actual execution status of the device's actions. These real-time motion parameters are defined as execution status parameters. Execution status parameters are used for subsequent action deviation analysis, action effect confirmation, or action adjustment calculations. During the motion parameter monitoring process, synchronization with the drive signal is maintained to ensure that the execution status of the device is known in real time.
[0171] This embodiment effectively ensures the temporal continuity of motion control and the electrical compatibility of device execution by parsing the globally optimal motion sequence into a time-series motion instruction set and accurately converting the instruction set into drive signals through the device controller. By sending the drive signals to the execution device in real time and driving the device to complete the physical motion, the efficient conversion and precise execution of device motion instructions are achieved. By monitoring the real-time motion parameters of the execution device and collecting them as execution status parameters, the actual motion state of the device during execution can be continuously tracked, establishing a dynamic closed loop between motion instructions and device state. This provides a real-time data foundation for subsequent motion adjustment or execution effect verification, thereby improving the real-time performance, accuracy, and execution reliability of the device motion control link and enhancing the device's adaptability to complex motion sequences.
[0172] In one embodiment, after step S60 above, the method further includes:
[0173] S701, during the execution of the globally optimal action sequence, the execution state parameters are obtained as action execution data;
[0174] S702, determine the execution difference value based on the action execution data and the initial action plan;
[0175] S703, when the execution difference value exceeds a preset threshold, an adjusted action instruction is generated based on the execution difference value and the updated current environment information;
[0176] S704, control the device to execute the adjusted action command.
[0177] In this embodiment, during the execution of the globally optimal action sequence, execution state parameters are acquired as action execution data. This means that during the actual physical actions performed by the device according to the globally optimal action sequence, the device's motion state information is collected in real time through a built-in or external state monitoring module. The execution state parameters may include the device's position, speed, acceleration, attitude, joint angle, path deviation, or other physical quantities that reflect the actual operating status of the device. These parameters are acquired through sensor arrays, position tracking systems, visual feedback units, or other detection methods. The system summarizes the above parameters to form structured action execution data. The action execution data reflects the actual execution effect of the device and is used for subsequent comparison with the theoretical planning results.
[0178] Determining the execution difference value based on action execution data and initial action planning involves the system comparing the acquired action execution data with parameters such as theoretical action path, target position, action sequence, and equipment status in the previously generated initial action plan. A difference calculation model or deviation analysis algorithm is used to quantitatively analyze the degree of deviation between the two. The execution difference value is a quantified deviation index that reflects the degree of difference between the current actual execution state of the equipment and the theoretical planned state. The difference value can be expressed using Euclidean distance, trajectory offset, time error, or a comprehensive multidimensional index. During the difference calculation process, environmental changes, equipment status fluctuations, and external interference factors are considered to ensure that the difference value results truly reflect the equipment execution deviation.
[0179] When the execution difference value exceeds a preset threshold, an adjusted action command is generated based on the execution difference value and the updated current environmental information. The preset threshold set by the system is the maximum acceptable deviation range for the equipment. If the execution difference value exceeds this threshold, it indicates that there is a significant deviation between the actual execution state of the equipment and the theoretical plan. Based on this, the system combines the updated current environmental information acquired in real time, including data such as environmental structure, obstacle distribution, spatial constraints, and external dynamic interference. The system comprehensively analyzes the execution deviation and environmental change factors, and dynamically regenerates the adjusted action command. The adjusted action command is a new targeted control command that corrects the equipment's action path, adjusts the execution parameters, or optimizes the action strategy to ensure that the equipment maintains normal operation under environmental changes or execution deviations.
[0180] The control device executes the adjusted action command, which means that the system transmits the generated adjusted action command to the device control module. The control module converts the command into a specific drive signal and applies it to the control interface or drive unit of the execution device. The execution device adjusts its current physical motion state according to the adjusted action command, including path fine-tuning, position correction, attitude control or other motion compensation operations. The adjustment process is completed dynamically without stopping the overall motion flow, realizing the device's adaptive correction of motion under the condition of execution deviation or environmental changes, and improving the robustness of device operation and the stability of task completion.
[0181] Example Description: In an embodied intelligent robot control system, to perform multimodal fusion perception and adaptive motion control tasks, the system first acquires visual modal input information and audio modal input information. The visual modal input information is collected by a high-definition imaging unit configured at the robot's front end, mainly including images of the operating environment, images of the target object, and spatial layout information. The audio modal input information is collected by a multi-microphone array, including user voice commands, ambient background sound, or on-site sound cues. The system inputs the visual modal input information into a visual encoder. The visual encoder uses a deep convolutional neural network structure to extract image feature information layer by layer, generate coded visual features, and use the coded visual features as the first modal feature vector. The audio modal input information is input into an audio encoder. The audio encoder uses time-frequency domain joint modeling to extract the voiceprint features and semantic information of the audio, generate coded audio features, and use them as the second modal feature vector.
[0182] The system integrates the first-modal feature vector and the second-modal feature vector, employing a multi-layer cross-modal information interaction mechanism. At the fusion layer, deep semantic association of different modal information is achieved, generating multi-modal fusion features. These multi-modal fusion features are input to a motion generator, which includes an input layer, a hidden layer, and an output layer. The input layer performs primary feature extraction on the multi-modal fusion features. The hidden layer enhances feature representation through deep nonlinear transformation. The output layer generates joint motion parameters, including joint angle sequences and joint velocity parameters. The system parses the joint angle sequences and calculates the complete motion trajectory path based on the sequence information and the mechanical structure model. Joint velocity parameters are extracted, and a velocity control curve is generated. The motion trajectory path and velocity control curve are fused to form basic motion commands. Subsequently, a device type identifier is added to ensure that the motion commands match the specific embodied intelligent robot model and control interface, ultimately outputting complete motion commands.
[0183] The system acquires the current state of the equipment, including parameters such as the joint state of the robotic arm, the position of the end effector, and the posture of the equipment. It collects current environmental information in real time through the environmental perception module, utilizing a combination of multiple sensors such as vision, radar, and ultrasound to acquire information on the environmental layout, obstacles, and dynamic interference factors. It then extracts features from the current environmental information to generate an environmental feature representation and obtains the task objective. The task objective originates from the upper-level task planning module, clarifying the operational purpose, such as moving a target object, obstacle avoidance paths, or interaction positions. The system inputs the action commands, the current state of the equipment, the environmental feature representation, and the task objective into the environmental modeling module. Based on the environmental information and equipment parameters, the environmental modeling module constructs an environmental model, completing the digital representation of the environmental structure, spatial topology, and dynamic information. It further performs kinematic calculations on the current state of the equipment to obtain the calculated state parameters. Combined with the environmental feature representation, it establishes a three-dimensional spatial model of the environment and determines the feasible region of the action path within the spatial model. Based on the environmental model, the calculated state parameters, and the feasible region of the action path, it inputs the data into the constraint optimization module. The constraint optimization module comprehensively considers equipment constraints, environmental limitations, and task requirements to optimize the action strategy and generate an optimized action plan, which serves as the initial action plan.
[0184] Based on the initial action planning, the system generates a set of candidate action sequences. The current environmental information is processed by the feature extraction module to generate an environmental feature representation. For each candidate action sequence, the expected reward value is calculated by combining the environmental feature representation, the current state of the device, and the task objective. The reward value is comprehensively evaluated based on the task completion degree, path efficiency, and energy consumption indicators. The system compares the expected reward values of all candidate action sequences and selects the candidate action sequence with the largest reward value to determine the globally optimal action sequence.
[0185] The system parses the globally optimal action sequence into a time-series action instruction set. The instruction set arranges specific control commands according to time order and inputs them into the device controller. The controller converts the time-series action instruction set into drive signals. The drive signals are sent to the execution device of the embodied intelligent robot through the communication bus, driving each execution unit of the execution device to complete the actual physical action. During the process, the system monitors the motion state of the execution device in real time, obtains real-time motion parameters such as position, speed, and attitude, and forms execution state parameters.
[0186] During the execution of the globally optimal action sequence, the system acquires execution status parameters as action execution data. This data is then compared with the initial action plan. The system calculates the execution difference value. If the difference value exceeds a preset threshold, it indicates that the current device is deviating from the plan. The system combines the execution difference value with the updated environmental information to dynamically generate adjusted action instructions, control the device to perform the adjustment operation, correct the action deviation, and ensure that the embodied intelligent robot can stably, adaptively, and efficiently complete multimodal fusion task instructions in complex dynamic environments.
[0187] In the control process of the medical and health service robot, visual and audio modal input information from the medical and nursing scene is first collected. Visual modal input information is acquired through a high-definition camera mounted on the robot, obtaining images of the ward environment, patient posture, and the layout of medical equipment. Audio modal input information is acquired through a sound pickup unit configured on the robot, including on-site sound signals such as patient voice requests, nursing instructions, or monitoring alarm sounds. The system inputs the visual modal input information into a visual encoder, which uses an image processing network based on a deep convolutional architecture to extract environmental structure, patient limb posture, and the location of target medical equipment, generating coded visual features as the first modal feature vector. Audio modal input information is input into an audio encoder, which uses a time-frequency fusion extraction algorithm to identify command language, sound source location, and alarm sound characteristics, generating coded audio features as the second modal feature vector.
[0188] The system fuses the feature vectors of the first and second modalities using a multimodal attention mechanism, integrating the spatial relationships and semantic features of different modalities to generate multimodal fused features. These multimodal fused features are input to the action generator, which performs preliminary structuring of the feature information through the input layer, deep feature abstraction through the hidden layer, and generates joint motion parameters for the medical robot through the output layer. These motion parameters include joint angle sequences and joint velocity parameters. The system parses the joint angle sequences, generates operation paths based on the robot's motion structure, extracts joint velocity parameters to generate continuous velocity control curves, and fuses the motion paths and velocity control curves to form basic action commands. Combined with device type identifiers in the medical scenario, such as nursing robots, drug delivery robots, or rehabilitation assistive devices, complete action commands are generated, ensuring that the action execution matches the device type.
[0189] The system collects the current status of the equipment, including the positions of each joint of the medical robot, the status of the actuators, and operational load parameters. Simultaneously, the environmental perception module collects real-time environmental information of the medical facility, including bed locations, obstacle distribution, dynamic personnel trajectories, and emergency status markers in the ward. The environmental information is processed by the feature extraction module to generate environmental feature representations. The system integrates nursing task objectives, derived from task requests issued by the nursing system, such as patient item delivery, body positioning adjustment, and rehabilitation assistance. The system inputs action commands, the current equipment status, environmental feature representations, and task objectives into the environmental modeling module. This module integrates spatial layout, equipment dynamic attributes, and task objective parameters to generate a structured environmental model. Kinematic calculations are then performed to obtain state calculation parameters under equipment motion constraints. Combined with the environmental feature representation, a three-dimensional spatial model of the environment is established, and feasible action path ranges are identified in this three-dimensional space. The system inputs the environmental model, calculated state parameters, and feasible action path domains into the constraint optimization module. Taking into account ward space limitations, medical safety zones, and patient comfort constraints, an optimized action plan is generated, which serves as the initial action plan.
[0190] Based on the initial action planning, the system generates a set of candidate action sequences. The feature extraction module processes the current environmental information in real time and updates the environmental feature representation. For each candidate action sequence in the set, the system combines the environmental feature representation, the current state of the equipment, and the task objective to determine the expected reward value. The expected reward value is calculated based on the efficiency of medical task completion, the stability of equipment operation, and the smoothness of the path. The system compares the expected reward values of all candidate action sequences and selects the candidate action sequence with the largest reward value as the globally optimal action sequence.
[0191] The system analyzes the globally optimal action sequence, generates a time-series action instruction set, and inputs the time-series action instruction set into the device controller. The device controller converts the action instruction set into drive signals, which are transmitted to the execution device of the medical robot through the control bus to drive the execution device to complete the action. The action includes object handling, delivery of medical supplies, or adjustment of patient position. During the execution of the action, the system monitors the motion parameters of the execution device in real time, including the actual motion path, speed changes, and load status of the execution device, and forms execution status parameters.
[0192] During the execution of the globally optimal action sequence, the system continuously acquires execution status parameters as action execution data and calculates the difference between the data and the initial action plan. The system calculates the execution difference value, and when the execution difference value exceeds a preset threshold, it indicates that there is a significant deviation between the actual action of the device and the pre-planned action. The system combines the current execution difference value with the re-acquired current environmental information to generate adjusted action instructions. The device controller then re-drives the device to perform the adjustment operation, ensuring that the medical robot continuously follows real-time changes in the dynamic medical environment, thus guaranteeing patient care safety, task execution accuracy, and the continuity of medical operations.
[0193] In the fintech business, the system first acquires first-modal and second-modal input information. Specifically, visual modal input information is obtained through image acquisition devices deployed in branches, lobbies, or automated teller terminals. This information includes customer behavior images, teller layout images, and risk warning signs. Simultaneously, audio modal input information is acquired through audio pickup devices, covering customer voice communication, environmental noise, and financial instruction announcements. The system inputs the visual modal input information into a visual encoder. The visual encoder uses an image feature extraction network to extract customer identity features, teller layout structure, and potential anomaly markers, generating coded visual features as the first-modal feature vector. Audio modal input information is input into an audio encoder. The audio encoder uses semantic recognition and sound source localization technology to extract voice instructions, communication semantics, and environmental background noise, generating coded audio features as the second-modal feature vector.
[0194] The system integrates the first-modality feature vector and the second-modality feature vector, and combines them with a multimodal feature fusion network to associate spatial information with speech and semantic information, generating multimodal fused features. These multimodal fused features are input to an action generator. The action generator structures the multimodal data through the input layer, deeply mines potential feature associations through the hidden layer, and generates joint motion parameters with operational guidance functions through the output layer. These parameters include the displacement sequence and response speed indicators of the system's actuators. The system analyzes the joint angle sequence, combines it with the financial terminal's structural design to determine the operation path, extracts joint speed parameters to generate a smooth speed control curve, and fuses the operation path and speed control curve to form basic action commands. The system distinguishes different financial business execution devices based on device type identifiers, such as smart teller machines, service robots, or risk warning terminals, and generates complete action commands to ensure that various devices accurately respond to business needs.
[0195] The system acquires the current status of the equipment, including the mechanical structure status of the terminal equipment, the location of the execution module, and the trigger status of the safety mechanism. The system also acquires current environmental information through an environmental perception module, covering customer aggregation and distribution, counter risk area identification, and real-time business flow data. The system processes this environmental information using a feature extraction module to form an environmental feature representation. Simultaneously, it acquires the task objectives within the financial business scenario, including specific customer guidance, risk warning push notifications, or financial document delivery requirements. The system inputs action instructions, the current equipment status, the environmental feature representation, and the task objectives into the environment modeling module. Combining branch layout, equipment attributes, and task requirements, it constructs an environmental model for the financial business scenario. The system further performs kinematic calculations on the current equipment status to obtain the calculated state parameters under structural constraints. Combining these with the environmental feature representation, it constructs a three-dimensional spatial model and determines the feasible operational area. Based on the environmental model, the calculated state parameters, and the feasible region of the action path, the system inputs them into a constraint optimization module. Combining customer activity distribution, business operation safety boundaries, and environmental change trends, it generates an optimized action plan, which serves as the initial action plan.
[0196] The system generates a set of candidate action sequences based on the initial action planning, processes the current environmental information in real time, updates the environmental feature representation, and calculates the expected reward value of the candidate action sequence by combining the environmental feature representation, the current state of the device and the task objective. The reward value takes into account operational efficiency, customer experience and business security. The system compares the reward values of all candidate action sequences and selects the candidate action sequence with the largest reward value as the global optimal action sequence.
[0197] The system analyzes the globally optimal action sequence, generates a time-series action instruction set, and converts the time-series action instruction set into drive signals through the device controller. The drive signals are transmitted to the execution device, and the device completes business-related physical operations according to the drive signals, including adjusting the customer guidance path, pushing risk information prompts, or adjusting the actions of the self-service terminal module. The system monitors the motion parameters of the execution device in real time, obtains data on device position changes, execution speed, and execution error during the operation, and forms execution status parameters.
[0198] During the execution of the globally optimal action sequence, the system continuously collects execution status parameters to form action execution data. The system compares the action execution data with the initial action plan and calculates the execution difference value. If the execution difference value exceeds a preset threshold, the system combines the execution difference value with the updated current environment information to dynamically adjust the action instructions and re-control the equipment to execute the adjusted actions. This ensures that the system has dynamic response capabilities in financial business scenarios, effectively improving the service efficiency of financial outlets, the accuracy of customer guidance, and the real-time and reliability of risk warnings.
[0199] This embodiment acquires execution status parameters and generates action execution data in real time during the execution of the globally optimal action sequence, enabling a comprehensive understanding of the actual operating status of the device. Based on the difference between the action execution data and the initial action plan, the device's action deviation is accurately quantified. When the deviation exceeds a set threshold, the device dynamically generates adjusted action instructions in conjunction with real-time environmental information and controls the device to perform the adjustment operation. This achieves dynamic closed-loop correction and intelligent adaptive adjustment of the device's actions, significantly improving the stability of the device's task execution and the robustness of its action planning in complex environments. It ensures that the device can still accurately and reliably complete task instructions under uncertain environments and dynamic changes.
[0200] In one embodiment, a multimodal fusion feature-driven motion control device is provided, which corresponds one-to-one with the multimodal fusion feature-driven motion control method described in the above embodiments. (Refer to...) Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the multimodal fusion feature-driven motion control device of the present invention. The modules include a multimodal feature extraction module 10, a multimodal fusion module 20, a motion generation module 30, a motion planning module 40, a sequence optimization module 50, and a device control module 60. Detailed descriptions of each functional module are as follows:
[0201] The multimodal feature extraction module 10 is used to acquire first modal input information and second modal input information, and generate a first modal feature vector based on the first modal input information and a second modal feature vector based on the second modal input information;
[0202] The multimodal fusion module 20 is used to fuse the first modality feature vector and the second modality feature vector to generate multimodal fusion features;
[0203] Action generation module 30 is used to generate action instructions based on the multimodal fusion features;
[0204] The motion planning module 40 is used to generate an initial motion plan based on the motion command, the current state of the device, the current environmental information, and the task objective.
[0205] Sequence optimization module 50 is used to generate a globally optimal action sequence based on the initial action planning;
[0206] The device control module 60 is used to control the device to execute the globally optimal action sequence.
[0207] In one embodiment, the multimodal feature extraction module 10 is specifically used for:
[0208] Obtain visual modal input information as the first modal input information;
[0209] The first modal input information is processed by a visual encoder to generate encoded visual features, and the encoded visual features are used as the first modal feature vector.
[0210] The audio modal input information is obtained as the second modal input information;
[0211] The second modal input information is processed by an audio encoder to generate encoded audio features, and the encoded audio features are used as the second modal feature vector.
[0212] In one embodiment, the action generation module 30 is specifically used for:
[0213] The multimodal fusion features are input into the action generator;
[0214] The multimodal fusion features are processed through the input layer of the action generator to generate a primary feature representation;
[0215] The primary feature representation is processed through the hidden layer of the action generator to generate a high-level feature representation;
[0216] The high-level feature representation is processed by the output layer of the motion generator to generate joint motion parameters;
[0217] The joint angle sequence in the joint motion parameters is analyzed, and the motion trajectory path is determined based on the joint angle sequence;
[0218] Extract the joint velocity parameters from the joint motion parameters, and generate a velocity control curve based on the joint velocity parameters;
[0219] The motion trajectory path and speed control curve are combined to form basic motion commands;
[0220] Add a device type identifier to the basic action instruction to generate a complete action instruction that includes the device type identifier.
[0221] In one embodiment, the action planning module 40 is specifically used for:
[0222] The current environmental information is collected through the environmental sensing module;
[0223] The current environmental information is used to extract features and generate an environmental feature representation;
[0224] Obtain the current status of the device and the task objective;
[0225] The action command, the current state of the device, the environmental feature representation, and the task objective are input into the environment modeling module, and an environment model is constructed through the environment modeling module.
[0226] Perform kinematic calculations on the current state of the device to obtain the calculated state parameters;
[0227] A three-dimensional spatial model of the environment is established based on the environmental feature representation, and a feasible domain for action paths is determined in the three-dimensional spatial model of the environment.
[0228] The environment model, the solution state parameters, and the feasible region of the action path are input into the constraint optimization module. The constraint optimization module generates an optimized action plan, which is then used as the initial action plan.
[0229] In one embodiment, the sequence optimization module 50 is specifically used for:
[0230] A candidate action sequence set is generated based on the initial action planning;
[0231] The current environment information is processed by the feature extraction module to generate an environment feature representation;
[0232] For each candidate action sequence in the candidate action sequence set, the expected reward value is determined based on the environmental feature representation, the current state of the device, and the task objective.
[0233] Compare the expected reward values of all candidate action sequences, and select the candidate action sequence with the largest expected reward value as the globally optimal action sequence.
[0234] In one embodiment, the device control module 60 is specifically used for:
[0235] The globally optimal action sequence is parsed into a time-series action instruction set;
[0236] The device controller converts the timing action instruction set into drive signals.
[0237] The drive signal is sent to the execution device to drive the execution device to perform physical actions;
[0238] The real-time motion parameters of the execution device are monitored as execution status parameters.
[0239] In one embodiment, the device control module 60 is specifically used for:
[0240] During the execution of the globally optimal action sequence, execution state parameters are obtained as action execution data;
[0241] The execution difference value is determined based on the action execution data and the initial action plan;
[0242] When the execution difference value exceeds a preset threshold, an adjusted action instruction is generated based on the execution difference value and the updated current environment information.
[0243] Control the device to execute the adjusted action command.
[0244] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used for communication with external user terminals via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a multimodal fusion feature-driven motion control method on the server side.
[0245] In one embodiment, a computer device is provided, which may be a user terminal, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides determination and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a multimodal fusion feature-driven motion control method on the user side.
[0246] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:
[0247] Acquire first modal input information and second modal input information, and generate a first modal feature vector based on the first modal input information and a second modal feature vector based on the second modal input information;
[0248] The first modality feature vector and the second modality feature vector are fused to generate a multimodal fusion feature;
[0249] Action commands are generated based on the multimodal fusion features;
[0250] Based on the action instructions, the current state of the device, the current environmental information, and the task objective, an initial action plan is generated;
[0251] Generate a globally optimal action sequence based on the initial action planning;
[0252] Control the device to execute the globally optimal action sequence.
[0253] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:
[0254] Acquire first modal input information and second modal input information, and generate a first modal feature vector based on the first modal input information and a second modal feature vector based on the second modal input information;
[0255] The first modality feature vector and the second modality feature vector are fused to generate a multimodal fusion feature;
[0256] Action commands are generated based on the multimodal fusion features;
[0257] Based on the action instructions, the current state of the device, the current environmental information, and the task objective, an initial action plan is generated;
[0258] Generate a globally optimal action sequence based on the initial action planning;
[0259] Control the device to execute the globally optimal action sequence.
[0260] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0261] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0262] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0263] It should be noted that if any software tools or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The embodiments described above are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A multimodal fusion feature-driven motion control method, characterized in that, Includes the following steps: Acquire first modal input information and second modal input information, and generate a first modal feature vector based on the first modal input information and a second modal feature vector based on the second modal input information; The first modality feature vector and the second modality feature vector are fused to generate a multimodal fusion feature; Action commands are generated based on the multimodal fusion features; Based on the action command, the current state of the device, the current environmental information, and the task objective, an initial action plan is generated, including: collecting current environmental information through an environmental perception module; extracting features from the current environmental information to generate an environmental feature representation; obtaining the current state of the device and the task objective; inputting the action command, the current state of the device, the environmental feature representation, and the task objective into an environmental modeling module to construct an environmental model; performing kinematic calculations on the current state of the device to obtain calculated state parameters; establishing a three-dimensional spatial model of the environment based on the environmental feature representation, and determining the feasible region of the action path in the three-dimensional spatial model; inputting the environmental model, the calculated state parameters, and the feasible region of the action path into a constraint optimization module to generate an optimized action plan, and using the optimized action plan as the initial action plan. Generate a globally optimal action sequence based on the initial action planning; Control the device to execute the globally optimal action sequence.
2. The multimodal fusion feature-driven motion control method as described in claim 1, characterized in that, Acquiring first modal input information and second modal input information, generating a first modal feature vector based on the first modal input information, and generating a second modal feature vector based on the second modal input information, includes: Obtain visual modal input information as the first modal input information; The first modal input information is processed by a visual encoder to generate encoded visual features, and the encoded visual features are used as the first modal feature vector. The audio modal input information is obtained as the second modal input information; The second modal input information is processed by an audio encoder to generate encoded audio features, and the encoded audio features are used as the second modal feature vector.
3. The multimodal fusion feature-driven motion control method as described in claim 1, characterized in that, Based on the multimodal fusion features, action commands are generated, including: The multimodal fusion features are input into the action generator; The multimodal fusion features are processed through the input layer of the action generator to generate a primary feature representation; The primary feature representation is processed through the hidden layer of the action generator to generate a high-level feature representation; The high-level feature representation is processed by the output layer of the motion generator to generate joint motion parameters; The joint angle sequence in the joint motion parameters is analyzed, and the motion trajectory path is determined based on the joint angle sequence; Extract the joint velocity parameters from the joint motion parameters, and generate a velocity control curve based on the joint velocity parameters; The motion trajectory path and speed control curve are combined to form basic motion commands; Add a device type identifier to the basic action instruction to generate a complete action instruction that includes the device type identifier.
4. The multimodal fusion feature-driven motion control method as described in claim 1, characterized in that, Generate a globally optimal action sequence based on the initial action planning, including: A candidate action sequence set is generated based on the initial action planning; The current environment information is processed by the feature extraction module to generate an environment feature representation; For each candidate action sequence in the candidate action sequence set, the expected reward value is determined based on the environmental feature representation, the current state of the device, and the task objective. Compare the expected reward values of all candidate action sequences, and select the candidate action sequence with the largest expected reward value as the globally optimal action sequence.
5. The multimodal fusion feature-driven motion control method as described in claim 1, characterized in that, Controlling the device to execute the globally optimal action sequence includes: The globally optimal action sequence is parsed into a time-series action instruction set; The device controller converts the timing action instruction set into drive signals. The drive signal is sent to the execution device to drive the execution device to perform physical actions; The real-time motion parameters of the execution device are monitored as execution status parameters.
6. The multimodal fusion feature-driven motion control method as described in claim 1, characterized in that, After controlling the device to execute the globally optimal action sequence, the method further includes: During the execution of the globally optimal action sequence, execution state parameters are obtained as action execution data; The execution difference value is determined based on the action execution data and the initial action plan; When the execution difference value exceeds a preset threshold, an adjusted action instruction is generated based on the execution difference value and the updated current environment information. Control the device to execute the adjusted action command.
7. A motion control device driven by multimodal fusion features, characterized in that, The multimodal fusion feature-driven motion control device includes: A multimodal feature extraction module is used to acquire first modal input information and second modal input information, and generate a first modal feature vector based on the first modal input information and a second modal feature vector based on the second modal input information; A multimodal fusion module is used to fuse the first modality feature vector and the second modality feature vector to generate multimodal fusion features; The action generation module is used to generate action instructions based on the multimodal fusion features; The action planning module is used to generate an initial action plan based on the action command, the current state of the device, the current environmental information, and the task objective. This includes: collecting current environmental information through an environmental perception module; extracting features from the current environmental information to generate an environmental feature representation; obtaining the current state of the device and the task objective; inputting the action command, the current state of the device, the environmental feature representation, and the task objective into an environmental modeling module to construct an environmental model; performing kinematic calculations on the current state of the device to obtain calculated state parameters; establishing a three-dimensional environmental spatial model based on the environmental feature representation and determining the feasible region of the action path within the three-dimensional environmental spatial model; inputting the environmental model, the calculated state parameters, and the feasible region of the action path into a constraint optimization module to generate an optimized action plan, and using the optimized action plan as the initial action plan. The sequence optimization module is used to generate a globally optimal action sequence based on the initial action planning; The device control module is used to control the device to execute the globally optimal action sequence.
8. A computer device, characterized in that, The computer device includes a memory, a processor, and a multimodal fusion feature-driven motion control program stored in the memory and executable on the processor. When the multimodal fusion feature-driven motion control program is executed by the processor, it implements the steps of the multimodal fusion feature-driven motion control method as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The storage medium stores a multimodal fusion feature-driven motion control program, which, when executed by a processor, implements the steps of the multimodal fusion feature-driven motion control method as described in any one of claims 1-6.
Citation Information
Patent Citations
Data decision-making method and system based on multi-modal large model analysis
CN119494079A
Text-guided multi-modal relationship extraction method and apparatus
WO2025130069A1