Action instruction generation and optimization method and device, equipment and medium
By collecting multimodal information and performing dynamic weighted fusion and decision optimization, the problem of insufficient perception and decision-making in embodied intelligence systems in complex environments has been solved, achieving efficient and flexible action execution capabilities and improving the intelligence level of the system in fields such as fintech and healthcare.
Patent Information
- Application Number
- CN202511051739.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-11-14
AI Technical Summary
Existing embodied intelligence systems lack the ability to dynamically fuse multimodal information and make unified decisions, making it unable to efficiently adapt to complex and changing environments. This results in poor coordination between perception, decision-making, and action, affecting the practicality and reliability of the system in fields such as fintech and healthcare.
The system collects environmental visual information, audio information, and action state information, performs dynamic weighted fusion through a multimodal attention mechanism, generates comprehensive decision features, inputs them into the decision network to generate action instructions, executes the instructions, and collects environmental feedback information to optimize the decision network.
It improves the perception, decision-making, and action coordination of embodied intelligent systems in complex environments, enables robots to make efficient decisions and execute flexibly in changing scenarios, and enhances the system's environmental adaptability and operational flexibility.
Smart Images

Figure CN120952043A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device, and storage medium for generating and optimizing action instructions. Background Technology
[0002] Despite the rapid development of embodied intelligence technology, robots still exhibit significant shortcomings in autonomous decision-making and multimodal information fusion capabilities in complex environments. Existing embodied intelligence systems typically rely on single-modal or low-level information input, lacking mechanisms for effectively integrating multi-source data such as visual, audio, and their own motion states. This makes it difficult to achieve efficient and accurate autonomous decision-making and motion control in dynamic environments. These deficiencies in information fusion and decision-making capabilities severely restrict the practicality and reliability of robots in complex application scenarios.
[0003] In the fintech sector, embodied intelligent devices are increasingly being integrated into systems such as smart branches, smart tellers, and financial service robots to assist customers with transactions, data collection, and risk alerts. However, limited by current technology, embodied intelligent devices in the financial field experience a disconnect between perception and decision-making when faced with multi-source information. These devices often fail to autonomously adjust their service strategies based on real-time environmental changes, leading to low transaction efficiency, poor customer experience, and operational conflicts and logical inconsistencies in multi-tasking scenarios, impacting the continuity and security of financial services.
[0004] In the healthcare sector, embodied intelligent systems are widely used in rehabilitation assistance, surgical collaboration, and intelligent nursing, requiring extremely high operational flexibility and environmental awareness. Existing systems generally suffer from single-mode information processing and fragmented decision-making processes, making it difficult to achieve effective collaboration of multimodal information. Especially in complex medical environments, they cannot comprehensively judge patient needs and operational feedback based on real-time visual, audio, and motion status information, resulting in slow system response, inaccurate operation, and difficulty in ensuring the safety and intelligence level of the medical process.
[0005] In summary, existing embodied intelligence systems generally lack efficient multimodal information dynamic fusion and unified decision-making capabilities. When faced with complex and changing environments, these systems struggle to autonomously generate precise actions and optimize feedback, failing to effectively improve their environmental adaptability, operational flexibility, and multitasking capabilities. These technical deficiencies directly impact the intelligence level and practical application effectiveness of embodied intelligence systems in fintech, healthcare, and other complex application scenarios. Summary of the Invention
[0006] The main objective of this invention is to provide a method, apparatus, device, and storage medium for generating and optimizing action instructions, aiming to solve the technical problems of the shortcomings of existing technologies in multimodal information fusion and adaptive decision optimization, which result in poor coordination between perception, decision-making, and action, and inability to efficiently adapt to dynamic environmental changes.
[0007] To achieve the above objectives, the present invention provides a method for generating and optimizing action instructions, comprising:
[0008] Collect environmental visual information, environmental audio information, and motion status information;
[0009] The environmental visual information and the environmental audio information are processed to obtain visual features and audio features, respectively.
[0010] A multimodal attention mechanism is used to dynamically weight and fuse the visual features, audio features, and action state information to generate comprehensive decision features;
[0011] The comprehensive decision features are input into the decision network to generate action instructions;
[0012] Execute the action instructions and collect environmental feedback information;
[0013] The decision network is optimized based on the environmental feedback information to generate an updated decision network.
[0014] Furthermore, to achieve the above objectives, the present invention provides an action command generation and optimization apparatus, comprising:
[0015] The perception and acquisition module is used to collect environmental visual information, environmental audio information, and motion status information;
[0016] The feature extraction module is used to process the environmental visual information and the environmental audio information to obtain visual features and audio features respectively;
[0017] The fusion decision module is used to dynamically weight and fuse the visual features, audio features, and action state information using a multimodal attention mechanism to generate comprehensive decision features;
[0018] The instruction generation module is used to input the comprehensive decision features into the decision network and generate action instructions;
[0019] The behavior execution module is used to execute the action instructions and collect environmental feedback information;
[0020] The learning and updating module is used to optimize the decision network based on the environmental feedback information and generate an updated decision network.
[0021] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and an action instruction generation and optimization program stored in the memory and executable on the processor, wherein when the action instruction generation and optimization program is executed by the processor, it implements the steps of the action instruction generation and optimization method as described above.
[0022] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing an action instruction generation and optimization program, wherein when the action instruction generation and optimization program is executed by a processor, it implements the steps of the action instruction generation and optimization method described above.
[0023] Beneficial Effects: This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as embodied intelligence, fintech, and healthcare. It discloses a method, apparatus, device, and medium for generating and optimizing action commands, including: collecting environmental visual information, environmental audio information, and action state information; processing the environmental visual and audio information to obtain visual features and audio features respectively; using a multimodal attention mechanism to dynamically weight and fuse the visual features, audio features, and action state information to generate comprehensive decision features; inputting the comprehensive decision features into a decision network to generate action commands; executing the action commands and collecting environmental feedback information; optimizing the decision network based on the environmental feedback information; and generating an updated decision network. This invention improves the perception, decision-making, and action coordination level of embodied intelligence systems in complex environments through the dynamic fusion and unified decision-making of multimodal information, realizing efficient decision-making and flexible execution capabilities of robots in changing scenarios. Attached Figure Description
[0024] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings:
[0025] Figure 1 This is a schematic diagram of an application environment for the action instruction generation and optimization method in one embodiment of the present invention;
[0026] Figure 2 This is a flowchart illustrating an embodiment of the action instruction generation and optimization method of the present invention;
[0027] Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the motion instruction generation and optimization device of the present invention;
[0028] Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention;
[0029] Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation
[0030] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0031] The action instruction generation and optimization method provided in this invention can be applied to, for example... Figure 1 In this application environment, the user terminal communicates with the server via a network. The server can collect environmental visual information, environmental audio information, and action state information from the user terminal. It processes the environmental visual and audio information to obtain visual and audio features respectively. A multimodal attention mechanism is used to dynamically weight and fuse the visual features, audio features, and action state information to generate comprehensive decision features. These comprehensive decision features are then input into a decision network to generate action commands. The action commands are executed, and environmental feedback information is collected. Based on the environmental feedback information, the decision network is optimized to generate an updated decision network. This invention improves the perception, decision-making, and action coordination level of embodied intelligent systems in complex environments through dynamic fusion and unified decision-making of multimodal information, achieving efficient decision-making and flexible execution capabilities for robots in changing scenarios. The user terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.
[0032] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the action instruction generation and optimization method provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0033] like Figure 2 As shown, the motion instruction generation and optimization method proposed in this invention includes the following steps:
[0034] S10 collects environmental visual information, environmental audio information, and motion status information;
[0035] In this embodiment, the process of acquiring environmental visual information, environmental audio information, and action state information is geared towards the comprehensive acquisition of multimodal information. It relies on the collaborative work of different types of sensing devices to form the multidimensional information input foundation for the embodied intelligent system. Environmental visual information refers to image information acquired through visual sensors that express the spatial state of the environment. Visual sensors may include RGB cameras, depth cameras, or composite perception modules combining color and depth information. When acquiring environmental visual information, the visual sensors acquire a continuous sequence of image frames covering a specified spatial range. Combined with 3D spatial modeling methods, structural information, object boundaries, size parameters, and spatial relative positions in the environment are extracted, providing data sources including dimensions such as color, texture, shape, and depth. The acquisition of environmental visual information is not limited to two-dimensional image data; it may also include spatial point cloud information formed based on structured light, time-of-flight principles, or binocular stereo vision technology to enhance the system's 3D spatial understanding capabilities.
[0036] Environmental audio information refers to various sound source information in the environment collected by audio sensors. These sensors can include single-channel pickup devices, array microphone devices, or spatial audio sensing components with directional sound source identification capabilities. Audio information acquisition encompasses continuous monitoring of environmental background noise, speech, mechanical equipment operation sounds, or other identifiable sound source signals. Combined with real-time filtering, spectrum analysis, and signal enhancement techniques, high-fidelity audio data is extracted to ensure the accuracy of subsequent information processing. The time synchronization and spatial positioning capabilities of audio information can be achieved through multi-channel signal fusion and sound source localization algorithms, improving the system's understanding of the spatial correlation of audio information.
[0037] Motion status information indicates the dynamic parameter status of various components during the robot's execution of actions. This information includes, but is not limited to, the spatial position, attitude angles, joint angles, velocity, acceleration, and mechanical feedback parameters of the robotic arm or end effector. Acquiring motion status information relies on the collaborative work of motion sensors, position encoders, inertial measurement units, or torque sensors to monitor the dynamic response process of the robot's motion system in real time. The acquisition of motion status information is not limited to monitoring the parameters of a single component; it can also cover the overall dynamic process of multiple components working collaboratively, achieving high-precision status feedback for complex tasks.
[0038] The acquisition process of environmental visual information, environmental audio information, and motion status information must ensure the temporal synchronization and spatial consistency of multimodal data. Specifically, the acquisition timing and data format of various types of information are coordinated through a timestamp synchronization module, a data caching mechanism, and a synchronization calibration strategy between sensors to avoid delays, offsets, or error accumulation during the information acquisition process, and to ensure the effectiveness and consistency of multi-source information in subsequent fusion and decision-making processes.
[0039] In practical implementation, environmental visual information can be acquired by deploying RGB-D cameras with a resolution greater than 1920×1080 pixels and depth perception capabilities. Combined with structured light illumination and depth calculation units, high-precision spatial point cloud information and color image information within the coverage area are obtained. For audio information acquisition, a microphone array consisting of at least four directional pickup units is used, combined with beamforming algorithms to improve the recognition rate of environmental voice commands. For low signal-to-noise ratio environments, an adaptive noise reduction module can be equipped to optimize audio signal quality in real time. For motion state information acquisition, multi-turn absolute encoders and high-sensitivity inertial measurement units are deployed at each joint of the robot arm to acquire angle changes, posture adjustments, and joint torque information in real time. Force sensing modules are placed at the end effector to provide feedback on the force and contact state at the end effector, enabling comprehensive monitoring of dynamic parameters in complex operating environments. Time synchronization of multimodal information is controlled by a global clock and a unified data bus, ensuring that visual, audio, and motion state information are transmitted and stored synchronously on the same time reference.
[0040] Example Explanation: In the field of embodied intelligence, for robotic systems operating in complex environments, the acquisition of environmental visual information typically involves integrating high-resolution RGB-D cameras, wide-angle vision units, and 3D LiDAR devices to collect real-time spatial images and depth information covering the operating area. This assists the robot in constructing environmental maps, identifying the geometric features and spatial positions of target objects, and improving its operational path planning and obstacle avoidance capabilities. Audio information acquisition relies on arrayed high-sensitivity microphones, combined with sound source localization and speech content recognition algorithms, to monitor voice commands, mechanical operating sounds, and external abnormal sound sources in the environment in real time. This helps the robot understand complex voice interactions, judge external risk signals, and enhance environmental adaptability and human-robot collaboration capabilities. Motion state information is acquired by configuring high-precision encoders, inertial measurement units, and force control sensing modules at various mobile joints, actuators, and end effectors on the robot body. This provides real-time feedback on the robot's spatial position, posture changes, motion trajectory, and contact mechanics parameters, supporting the robot in performing high-precision, continuous, and dexterous operational tasks in unknown, dynamic, or complex environments, improving the overall system's stability, flexible control, and operational safety in embodied intelligence scenarios. By combining the multimodal synchronous acquisition of the above information, the embodied intelligence system can quickly build a comprehensive and accurate multidimensional perception model in a dynamic environment, which serves as the basis for subsequent fusion decision-making and autonomous action planning, thereby achieving human-level environmental understanding and autonomous interaction capabilities.
[0041] In the healthcare field, during the acquisition of environmental visual information, medical robots utilize high-precision depth cameras to acquire three-dimensional spatial structural information around the operating table, assisting in the positioning of surgical instruments, recognition of human tissue contours, and planning of operational paths. In the audio information acquisition phase, a microphone array monitors medical personnel's voice commands in real time, and combined with speech recognition algorithms, accurately interprets operational requests or warning messages, improving the intelligent response level of the medical process. For the acquisition of motion status information, the robot's end effector's high-sensitivity force sensing module provides real-time feedback on mechanical changes during operation, ensuring the safety and accuracy of the operation.
[0042] In the fintech sector, intelligent robot systems serving financial service outlets combine facial recognition and spatial positioning technologies to acquire environmental visual information, obtaining real-time customer location and behavioral information to assist in intelligent guidance and identity verification. Audio information is acquired through directional microphone arrays to recognize customer voice inquiries, enabling multilingual and multi-dialect natural language interaction. The acquisition of motion status information monitors the robot's posture and the position of the interactive terminal, ensuring the accuracy of interface adjustments and the stability of service actions.
[0043] This embodiment, through the coordinated acquisition of environmental visual information, environmental audio information, and action status information, enables the system to obtain multimodal data sources that comprehensively reflect environmental status, acoustic information, and its own dynamic parameters. This overcomes the limitations of single perception methods in spatial structure understanding, voice information parsing, and action feedback monitoring, and improves the accuracy and robustness of multimodal information fusion and intelligent decision-making.
[0044] S20, process the environmental visual information and the environmental audio information to obtain visual features and audio features respectively;
[0045] In this embodiment, environmental visual information refers to environmental image data acquired during the preliminary data collection phase, which includes spatial structure information, texture detail information, and depth information. This includes, but is not limited to, RGB images, RGB-D images, stereo vision images, or environmental image sequences obtained through multi-sensor fusion. This information originates from the robot's external visual sensing system, such as a high-resolution camera, a wide-angle vision module, a depth camera, or a 3D vision imaging device. The essence of environmental visual information is to reflect the appearance characteristics of objects in the robot's workspace, spatial layout information, and environmental changes. Specific applications may include the shape of packaging materials in logistics scenarios, the layout of operating tables in medical scenarios, or the structure of counter areas in financial service environments.
[0046] Environmental audio information refers to the sound data related to the operating environment obtained through an external audio sensor array, covering environmental background noise, command voice signals, equipment operation sounds, and external abnormal sound source signals. The sources of this information include, but are not limited to, omnidirectional microphone arrays, distributed acoustic sensors, or sound pickup devices with speech recognition capabilities. The specific role of environmental audio information is to assist the system in recognizing language inputs, monitoring abnormal sounds, and analyzing environmental dynamic changes, ensuring that the robot can complete multimodal fusion analysis tasks on the premise of auditory perception.
[0047] Visual features are the high-dimensional information expression results for describing image content formed after data preprocessing, feature extraction, structure encoding, and feature mapping operations based on environmental visual information. Visual features can include spatial structure features, texture pattern features, boundary contour features, color distribution features, geometric shape features, etc. The generation of visual features usually relies on convolutional structures, feature pyramid structures, or graph neural network structures in deep neural networks to ensure that the output has good spatial expression ability and semantic parsing ability to meet the requirements of subsequent multimodal information fusion and decision-making generation.
[0048] Audio features are a set of parameters for expressing audio content formed after time-domain and frequency-domain analysis, acoustic feature extraction, and high-dimensional mapping based on environmental audio information. Audio features can include Mel spectrum coefficients, speech content vectors, sound source direction features, sound energy distribution features, or temporal structure features. The generation of this feature usually depends on recurrent neural networks, attention mechanisms, or transform structures based on spectral analysis to ensure that audio features can effectively express language information, abnormal sound patterns, and spatial acoustic structures in the environment, supporting the multimodal linkage analysis of the system.
[0049] In the specific actual implementation process, first, it is necessary to perform data normalization processing, noise suppression, and image enhancement operations on the environmental visual information obtained by a high-resolution image sensor to improve the quality of image data. Subsequently, spatial structure features, boundary contour features, and local texture features are extracted through a convolutional neural network to form an initial visual feature expression. For the initial visual features, a feature enhancement module can be further applied, through generative adversarial networks, variational autoencoders, or feature transformation networks, to improve the robustness and discriminability of the visual features, and finally output the visual features for subsequent steps.
[0050] In the processing of environmental audio information, signal sampling, noise reduction filtering, and short-time Fourier transform are first performed to obtain the time-frequency distribution characteristics of the audio. Then, a Mel filter bank and spectrogram generation module are used to extract audio feature representations from the speech signal. Further, a long short-term memory network or Transformer structure is used to optimize the sequential representation and environmental adaptability of the audio features, outputting audio features that maintain a unified high-dimensional representation structure with visual features, thus meeting the requirements of subsequent multimodal fusion.
[0051] In embodied intelligence systems, the generation of visual features can employ different image processing network structures depending on the specific task requirements. For example, in applications requiring high-precision object recognition, deep convolutional neural network structures can be used for multi-scale feature extraction, combined with attention mechanisms to enhance feature representation. In large-scale environmental perception tasks, lightweight image coding networks and depth estimation modules can be applied to quickly extract spatial structural features, balancing computational efficiency and feature quality.
[0052] During the generation of audio features, parameters can be optimized and structures adjusted based on different audio processing techniques. For example, in complex noisy environments, adaptive noise reduction modules and multi-channel beamforming technology can be used to enhance effective speech signals; in scenarios with intensive voice interaction, end-to-end speech coding structures and semantic analysis networks can be introduced to improve the expressive power of audio features and the accuracy of semantic parsing; in applications requiring spatial sound source localization, spatial information from microphone arrays can be combined to generate high-dimensional audio features containing the location and spatial distribution of sound sources, enhancing the system's environmental understanding capabilities.
[0053] The specific dimensions, representation forms, and data structures of visual and audio features can be flexibly adjusted according to different system architectures and task requirements. For example, a unified multimodal embedding space, serialized high-dimensional feature vectors, or graph structure feature sets can be adopted to ensure that the information of each modality has good alignment and fusion, thereby improving the overall system performance in the process of multimodal information analysis and fusion decision-making.
[0054] Example Description: In the healthcare field, for scenarios involving robot-assisted surgery, laboratory sample processing, and intelligent rehabilitation equipment operation, environmental visual information is acquired through high-precision endoscopes, external auxiliary imaging devices, or 3D structured light cameras to form visual data expressing the structure of the surgical area, the shape of sample containers, or the status of rehabilitation equipment. Through spatial structure feature extraction and feature enhancement, visual features with high spatial resolution and geometric structure representation capabilities are generated. Combined with audio information of doctor's commands acquired through a voice interaction system, multi-dimensional acoustic feature extraction and semantic optimization generate audio features expressing the speech content and environmental sound status. This assists the system in accurately understanding operational needs and dynamically adapting to the operating environment in high-risk, delicate operating environments, ensuring the system's operational accuracy and safety.
[0055] In the fintech business sector, targeting scenarios such as smart counter services, automated customer reception, and remote operation and maintenance of financial equipment, environmental visual information is acquired through a multi-camera system to obtain counter structure, customer behavior, and the status of the surrounding environment. Combined with acoustic sensing units to obtain voice commands and environmental sounds, the system quickly establishes multimodal information expression through an efficient visual and audio feature generation process, improving the perception and understanding of business scenarios. This assists intelligent systems in providing higher processing efficiency and service experience in tasks such as financial business processing, customer service response, and abnormal status early warning.
[0056] This embodiment processes environmental visual and audio information independently to obtain high-dimensional information representations of spatial structure and environmental acoustic features. This effectively enhances the embodied intelligent system's multidimensional perception capability of complex environments, ensuring that the system acquires visual and audio features with high expressive and discriminative capabilities before performing multimodal information fusion. This provides a stable and accurate perceptual information foundation for subsequent fusion analysis and action decision-making, thereby enhancing the overall system's adaptability and operational reliability in dynamic, complex, and highly uncertain environments.
[0057] S30, a multimodal attention mechanism is used to dynamically weight and fuse the visual features, audio features, and action state information to generate comprehensive decision features;
[0058] In this embodiment, the multimodal attention mechanism refers to a processing method that dynamically models the associations, adjusts weights, and fuses information from different perceptual sources in the feature space by introducing a cross-modal information interaction structure. Based on an attention structure, this mechanism generates a dynamic weight matrix by calculating the correlation and importance distribution between different information modalities, thereby achieving a weighted combination of information and improving the overall synergy and discriminative ability of information expression. The technical sources of the multimodal attention mechanism include, but are not limited to, Transformer structures, two-stream interaction networks, cross-modal alignment structures, or adaptive information fusion frameworks, and it is widely used in comprehensive perception, semantic understanding, and fusion decision-making tasks in multi-information-source environments.
[0059] Visual features refer to a set of high-dimensional parameters extracted from environmental visual information that express information such as the spatial structure of the environment, the geometric features of objects, texture distribution, and spatial layout. They are typically derived from the outputs of convolutional neural networks, feature pyramid structures, or spatial graph neural networks. In embodied intelligence tasks, visual features are used to express the geometric structure of the operational space, the position and shape of target objects, and the dynamic changes in the environment, ensuring that the system possesses stable and accurate spatial perception capabilities.
[0060] Audio features refer to a set of parameters extracted from environmental audio information that express speech content, environmental sound source features, sound temporal structure, and spatial acoustic state. These features typically originate from spectrum analysis modules, acoustic feature extraction structures, or temporal information coding networks. In multimodal information analysis, audio features are used to assist systems in understanding speech commands, perceiving environmental dynamics, and identifying abnormal sounds, thereby improving the system's auditory perception and speech interaction capabilities.
[0061] Action state information refers to a high-dimensional set of information expressing the structural state, motion parameters, and spatial position of each execution unit in the current robot system. This includes, but is not limited to, parameters such as joint angles, end effector positions, motion trajectories, and force control states. Action state information typically originates from internal motion sensors, position encoders, and force sensing modules, reflecting the robot's own motion state and operational execution, ensuring the system possesses real-time and accurate self-state perception capabilities.
[0062] Dynamic weighted fusion refers to the joint calculation and information integration of visual features, audio features, and action state information based on the intermodal weight relationships calculated using a multimodal attention mechanism, and generating a fusion result with the collaborative expression capability of multi-source information. Dynamic weighted fusion adaptively adjusts the contribution of each modal information, strengthens the expression of key information, suppresses redundant and noisy information, and improves the overall consistency and discriminative ability of multimodal information expression.
[0063] Comprehensive decision features refer to high-dimensional representations generated by dynamically weighted fusion of visual features, audio features, and action state information to support the decision network in making action decisions and generating strategies. Comprehensive decision features possess the ability to align multiple information sources, collaboratively express semantic information, and adapt to dynamic environments. Serving as the foundational input for subsequent action decisions, strategy optimization, and operation execution, they ensure the system has efficient and accurate autonomous decision-making capabilities in complex and ever-changing environments.
[0064] In the specific implementation, visual features, audio features, and action state information are first input into the multimodal attention structure, and linear transformation operations are performed on each to map them to a unified high-dimensional feature space, ensuring that the information from each modality has good alignment and interactivity. Subsequently, based on the attention weight calculation module, the correlations between visual features and audio features, visual features and action state information, and audio features and action state information are comprehensively considered. Attention weight parameters among the three modalities are calculated through dot product similarity or multi-head attention mechanisms, dynamically reflecting the importance distribution of each modality in the current task environment.
[0065] By combining the calculated multimodal attention weight matrix, the system performs a weighted summation operation on the transformed visual features, audio features, and action state information, integrating them to form a comprehensive decision feature that includes the fusion results of multimodal information. This comprehensive decision feature possesses the capabilities of collaborative expression from multiple information sources, dynamic weight adaptive adjustment, and efficient integration of semantic information. As input information for the decision network in subsequent steps, it fully supports the system's action decision-making and strategy generation processes.
[0066] The input of visual features can come from the output of image coding networks with different structures. For example, a multi-layer convolutional network based on ResNet can be used to extract spatial texture and geometric structure features, or a pyramid structure combined with a spatial graph neural network can be used to enhance the ability to represent environmental information at different scales. The input of audio features can be combined with Mel-spectrum analysis and sequence modeling structures, such as applying LSTM networks or Transformer coding structures, to extract speech content features and spatial sound source information, ensuring the efficiency and robustness of the system in language interaction and environmental acoustic perception.
[0067] Action state information can be derived from a combination of multi-dimensional information from the internal sensing system, including but not limited to the real-time angle values of each joint, the spatial position of the end effector, the motion path trajectory and the current force state. Specific parameters come from the position encoder, inertial measurement unit and force control sensor. Through data normalization and high-dimensional mapping structure, action state information that expresses the system's self-state is formed.
[0068] In the design of multimodal attention mechanisms, different attention structures can be selected based on task requirements. For example, in complex scenarios with rapidly changing multiple information sources, a multi-head attention structure combined with residual connections and layer normalization techniques can be used to improve the stability of attention distribution and the robustness of information representation. In scenarios where the weights of information modalities change relatively little, a single-head attention structure and adaptive gating mechanisms can be used to reduce computational complexity and system load, ensuring that the system has good real-time performance and responsiveness.
[0069] In the dynamic weighted fusion operation, the weight matrix can be calculated based on the standard dot product attention structure, combined with positional encoding information and a modality importance adaptive adjustment module, to dynamically optimize the consistency and discriminative ability of the fusion result. The output structure of the comprehensive decision features can be designed as a fixed-dimensional high-dimensional vector, a serialized feature matrix, or a structured multi-information-source joint expression result, depending on the system architecture and task requirements, ensuring that the comprehensive decision features have good scalability and system compatibility.
[0070] Example Description: In the healthcare business field, for scenarios such as robot-assisted intelligent surgery, automated laboratory sample processing, and intelligent rehabilitation equipment operation, the system uses a multimodal attention mechanism to dynamically and weightedly fuse visual features expressing the surgical environment structure, audio features of the doctor's voice commands, and motion state information of the equipment. This generates comprehensive decision features that express the collaborative information of surgical needs, environmental state, and equipment motion parameters. This assists the system in accurately generating operation strategies and action decisions under high-risk, complex structures, and dynamic environmental changes, thereby improving operational accuracy and patient safety.
[0071] In the field of fintech business, targeting intelligent counter service processing, remote control of automated equipment and intelligent customer service systems, the system uses a multimodal attention mechanism to dynamically integrate visual features reflecting the counter layout and environmental structure, audio features of customer voice interaction information and action status information of real-time equipment operation status, to generate comprehensive decision features. This helps the system to intelligently generate service strategies and equipment operation decisions under conditions of multiple users, complex service needs and dynamic environmental changes, thereby improving business processing efficiency, service response capabilities and the stability and security of the overall system.
[0072] This embodiment employs a multimodal attention mechanism to dynamically weight and fuse visual features, audio features, and action state information. This effectively enhances the embodied intelligence system's comprehensive perception capability and information collaborative expression level in multi-information source environments. It ensures that the system has the ability to dynamically adjust information weights, adaptively integrate multimodal information, and efficiently generate fused expression results in complex environments. This provides a complete and accurate information input foundation for subsequent action decisions and strategy optimization, and enhances the system's autonomous decision-making capability and operational reliability in highly dynamic, highly uncertain, and highly complex operating environments.
[0073] S40, input the comprehensive decision features into the decision network to generate action instructions;
[0074] In this embodiment, the comprehensive decision feature is a multi-dimensional representation generated by fusing visual features, audio features, and action state information and dynamically weighting them through a multimodal attention mechanism. It is primarily used to support the system's action decisions and strategy outputs. The comprehensive decision feature internally includes multi-information source structure alignment, cross-modal information collaborative expression, and adaptability to dynamic environmental changes, serving as the fundamental information input to ensure the system's efficient and accurate action generation capabilities.
[0075] Decision networks are structured decision systems used to generate action instructions based on comprehensive decision features. They typically include multi-level information encoding structures, attention mechanism modules, feedforward information transformation networks, and output instruction generation units, possessing the capabilities of multi-dimensional information representation, semantic structure capture, and dynamic policy output. Decision networks can originate from deep neural networks, sequence modeling networks based on Transformer structures, or multi-information fusion networks combining graph structure information, and are widely used in the action generation, policy optimization, and operation instruction output processes of embodied intelligent systems.
[0076] Inputting comprehensive decision features into a decision network refers to using the comprehensive representation output of a multimodal attention mechanism as input information to drive the network's internal information encoding, feature transformation, and policy generation processes. This enables dynamic action decisions and command generation based on environmental states and task requirements. This operation ensures that the system can generate operational commands in real time and accurately based on the comprehensive representation results under dynamic changes in multi-source information, thereby improving the system's autonomous decision-making and efficient operation capabilities.
[0077] Generating action instructions refers to the process by which a decision network, based on comprehensive decision-making characteristics, performs multi-layered information processing, strategy analysis, and result reasoning to output standardized operation instructions that express specific operational requirements, action parameters, and execution paths, driving each execution unit of the system to complete the corresponding action. Action instructions typically include, but are not limited to, high-dimensional information sets such as joint angle values, end-effector position parameters, trajectory control information, and dynamic adjustment parameters, ensuring that the system possesses good operational accuracy, execution flexibility, and adaptability to dynamic environments.
[0078] In its implementation, the system inputs comprehensive decision features into the multi-level encoding structure of the decision network. First, through multi-level information encoding and expression compression, semantic information and structural features within the features are extracted, improving the system's information expression efficiency and structural adaptability. Subsequently, the multi-head attention mechanism within the decision network dynamically captures the correlation and importance distribution among information elements for different information dimensions and feature sequences, generating attention-optimized features to ensure consistency and efficient collaboration among multiple information sources.
[0079] In the feedforward neural network structure, the system performs nonlinear mapping and representation space transformation on attention-optimized features, enhancing the complexity of feature representation and policy reasoning ability, thereby improving the system's policy adaptability and information discrimination ability in complex environments. Finally, through the output layer of the decision network, the system maps the transformed features into standardized action instructions that express specific operational requirements and control parameters, driving the system's actuators to perform specific operations according to the action instructions.
[0080] The input of comprehensive decision features can be expressed using multidimensional high-dimensional vectors, sequential structures, or graph structures. The appropriate feature organization form can be selected based on task requirements to improve the system's compatibility and operational flexibility in information representation. The structure of the decision network can be flexibly adjusted according to changes in the scenario and system performance requirements. For example, in highly dynamic environments with complex information structures and varied operational needs, a deep sequence modeling network based on the Transformer structure can be used, combined with multi-layer information encoding, layer normalization, and residual connections, to improve the system's representational stability and policy generation robustness.
[0081] The design of multi-head attention mechanisms can dynamically adjust the number and structural layout of attention heads based on the complexity of information modalities and system computing resources, ensuring that the system controls computational load and real-time requirements while maintaining information representation capabilities. The structure of feedforward neural networks can combine different activation functions, information transformation layers, and network widths to dynamically optimize the system's information representation capabilities and policy inference efficiency.
[0082] The output format of motion commands can be flexibly set according to the type of actuator and operational requirements. For example, in a joint-type actuator system, motion commands can include the real-time angle values and position adjustment parameters of each joint. In a space operating system, motion commands can include the position coordinates, trajectory path, and attitude parameters of the end effector, ensuring the accuracy and efficiency of system operation.
[0083] Example Description: In the field of healthcare, for scenarios such as robotic surgical assistance, automated sample processing, and intelligent rehabilitation equipment operation, the system dynamically generates action instructions that express surgical needs and equipment operation parameters based on comprehensive decision-making features. This ensures that the equipment has high flexibility and high precision in operation under complex surgical environments and biological sample handling conditions, thereby improving surgical safety and operational efficiency, and optimizing the stability and flexibility of experimental operations.
[0084] In the field of fintech business, for intelligent counter business processing, remote control of automated equipment and customer service systems, the system generates action instructions in real time based on comprehensive decision-making characteristics, expressing the equipment's operating needs and service response parameters. This drives the equipment to intelligently generate operation strategies and service response actions under conditions of multi-task concurrency, complex service needs and dynamic environmental changes, thereby improving business processing efficiency, system stability and service intelligence, and ensuring that the system has a good customer interaction experience and service quality.
[0085] This embodiment inputs comprehensive decision features into a decision network to generate action instructions that express specific operational needs. The system can generate dynamic operation strategies and execution instructions in real time and accurately based on the collaborative expression results of multiple information sources. This enhances the autonomous decision-making ability and operational flexibility of the embodied intelligent system in complex environments, ensuring that the system has efficient and stable operation execution capabilities and task adaptability under changing environments and complex task conditions, thereby enhancing the overall intelligence level and environmental adaptability of the system.
[0086] S50, execute the action command and collect environmental feedback information;
[0087] In this embodiment, executing an action command refers to transmitting the action commands generated by the decision network to the actuators of the equipment through the control system, driving the equipment or mechanical system to complete the corresponding physical action. The action commands specifically include information such as operating parameters, execution path, movement speed, and posture adjustment. They can be single-dimensional control signals or multi-dimensional joint control parameters, dynamically configured according to different execution requirements. The sources of the action commands include, but are not limited to, decision results generated by fusing visual features, audio features, and action state information, ensuring that the commands possess targeted, adaptive, and high-precision control capabilities.
[0088] An actuator is a hardware system or physical module capable of performing actions. It can be a robotic arm, end effector, mobile platform, or multi-degree-of-freedom manipulation module. It adjusts its spatial position, motion posture, or functional state according to action commands to achieve the desired operational effect. The structure and type of the actuator are designed based on the task environment and operational requirements, possessing high dynamic response, high-precision control, and environmental adaptability, enabling it to complete flexible and varied operational tasks in dynamic and complex environments.
[0089] Collecting environmental feedback information refers to the process by which the system acquires data on changes in the environmental state through various sensing methods during the execution of action commands, comprehensively evaluates the effectiveness of action execution and environmental response, and forms multi-dimensional feedback information for system optimization and strategy adjustment. Environmental feedback information includes, but is not limited to, a multi-modal and multi-level set of information such as visual environmental change information, audio environmental change information, environmental state change data, and reward signals.
[0090] Environmental visual change information refers to information about changes in the state, spatial structure, or visual features of objects in the environment after an action is performed, acquired through visual sensors. This information originates from high-precision cameras, depth cameras, or image sensing units, and is combined with image processing algorithms to extract change data in real time, reflecting the dynamic response of the environment. Environmental audio change information refers to information about changes in ambient sound, voice commands, or background noise acquired through audio sensors. This information is combined with signal processing and audio analysis algorithms to assess the impact of changes in the environmental sound field on action execution.
[0091] Environmental state change data refers to information on changes in environmental physical attributes, spatial parameters, or functional indicators acquired through the state monitoring module. This includes multi-dimensional data such as temperature, humidity, pressure, location, velocity, and acceleration, reflecting the overall trend of environmental state changes and their operational impact. Reward signals are performance evaluation data dynamically generated by the reward function module, combining the action execution effect with environmental feedback results. These signals reflect the contribution of the operational results to the overall system objective, supporting subsequent system optimization and strategy updates.
[0092] In practice, the system first receives and parses action commands through the execution device, driving each functional unit to complete spatial movement, posture adjustment, or functional operation, ensuring the accuracy and stability of the operation. During the execution of the action, the vision sensor collects environmental images and spatial structure data in real time, and combines them with vision processing algorithms to generate environmental visual change information, dynamically monitoring changes in the environmental visual state.
[0093] The audio sensor synchronously acquires environmental sound field information, extracts speech signals, environmental noise, and background sound features to form environmental audio change information, reflecting the environmental sound field response in real time. The status monitoring module acquires environmental physical parameters and functional status through multiple sensing units, and combines data fusion algorithms to generate environmental status change data, comprehensively perceiving the trend of environmental status changes and their impact on operations.
[0094] The system, in conjunction with the reward function module, dynamically calculates reward signals based on environmental feedback, operational effects, and expected goals, quantitatively evaluating the effectiveness of action execution and the degree of system contribution. Ultimately, environmental visual changes, environmental audio changes, environmental state changes, and reward signals are integrated through the information integration module to form environmental feedback information, supporting the system's self-optimization and dynamic strategy adjustment.
[0095] When executing motion commands, the actuator can be a robotic arm, end effector, or mobile platform with multi-degree-of-freedom control capabilities. Different degrees of freedom, control precision, and response speeds can be configured according to task requirements to ensure the flexibility and stability of the operation. The parameter structure of the motion commands can be expressed based on joint space, Cartesian space, or task space, and the information organization can be dynamically adjusted according to different task environments to improve the efficiency and adaptability of the operation.
[0096] Environmental visual change information can be acquired using high-resolution RGB cameras, depth cameras, or multimodal vision systems, combined with edge computing platforms or high-performance image processing units to complete visual data acquisition and change analysis in real time, ensuring the accuracy and timeliness of visual feedback information. Environmental audio change information can be acquired using directional microphone arrays, array-type audio sensing systems, or multi-source sound field sensing devices, combined with speech recognition and acoustic analysis algorithms to dynamically extract environmental sound change features.
[0097] The acquisition of environmental state change data can be combined with inertial measurement units, force sensing systems, temperature and humidity monitoring modules, and environmental monitoring networks. Monitoring strategies and data acquisition frequencies can be dynamically adjusted based on environmental physical parameters and state change trends, improving the comprehensiveness and timeliness of environmental state change data. Reward signal generation can be achieved by combining task objectives, operating procedures, and system performance indicators to design multi-level, multi-indicator reward functions. This allows for dynamic evaluation of action execution effectiveness and supports the system's self-learning and strategy optimization.
[0098] Example Description: In the healthcare business field, the system drives rehabilitation robots, surgical aids, or automated biological sample processing equipment to perform fine motor operations based on motion commands. It acquires visual change information of the surgical field, operation area, or experimental environment through visual sensors, and extracts environmental sound field response and voice interaction signals by combining audio sensors. It dynamically acquires environmental state change data, and evaluates the operation effect and medical task achievement by combining reward signals. This improves the dexterity, flexibility, and safety of equipment operation, and ensures stable operation execution capability and efficient task adaptability in highly dynamic and complex environments.
[0099] In the field of fintech business, the system drives smart terminals, automated service equipment or remote operation platforms through execution devices to complete business operations and environmental interactions. Based on visual change information, audio change information and environmental status data, it monitors changes in the operating environment and customer needs in real time. Combined with reward signals, it dynamically optimizes business processes and operation strategies, improves the system's autonomous service capabilities, operational efficiency and customer satisfaction, and ensures that fintech equipment has good service continuity, operational stability and intelligent interaction level in complex business environments.
[0100] This embodiment, by executing action commands and collecting environmental feedback information, enables the system to acquire multi-dimensional environmental response information in real time and comprehensively based on the action execution results and dynamic changes in the environment. It can dynamically evaluate the operation effect and system contribution, form multi-information fusion environmental feedback information, support the system's self-optimization, strategy updates and action adjustments, improve the system's operational accuracy, environmental adaptability and autonomous decision-making ability in complex environments, and enhance the overall intelligence level and task execution efficiency of the system.
[0101] S60, optimize the decision network based on the environmental feedback information to generate an updated decision network.
[0102] In this embodiment, the decision-making network is optimized based on environmental feedback information to generate an updated decision-making network, which includes multiple stages. The aim is to dynamically adjust the structural parameters and decision-making strategies of the decision-making network through environmental feedback information, thereby improving the system's adaptability and intelligence level in complex environments. Environmental feedback information refers to a multimodal, multi-dimensional comprehensive feedback result formed through data fusion and information integration, combining visual change information, audio change information, environmental state change data, and reward signals, reflecting the action execution effect and environmental response.
[0103] Optimizing a decision network refers to dynamically adjusting the parameter structure, weight distribution, or learning strategy within the decision network based on environmental feedback information, thereby enhancing the network's ability to perceive and make decisions in response to environmental changes. A decision network is a deep learning network with a multi-level coding structure, multi-head attention mechanisms, and feedforward neural network modules. It combines multimodal information input to generate high-precision action commands, supporting autonomous decision-making and task execution in embodied intelligent systems.
[0104] Generating an updated decision network refers to creating a version of the decision network with stronger environmental adaptability, better decision-making efficiency, and higher operational accuracy through parameter adjustment, structural reconstruction, or strategy updates during the optimization process, ensuring that the system can continuously improve its operational performance and environmental coordination capabilities.
[0105] In its implementation, the system first extracts real-time observation data and reward signals from environmental feedback information. The real-time observation data comes from environmental state change data, reflecting dynamic environmental changes and operational impacts. The reward signals are dynamically generated based on the action execution effect and task achievement degree, quantifying the operational contribution and system performance.
[0106] Based on real-time observation data, the system employs an environmental perception model to perform state probabilistic inference and generate updated environmental state probability distributions. The environmental perception model can be a dynamic Bayesian network, a hidden Markov model, or a deep probabilistic graphical structure, combining historical data and real-time observation information to dynamically infer the uncertainty and changing trends of the current environmental state.
[0107] Based on the updated environmental state probability distribution, the system determines the current environmental state, forming a quantitatively expressed environmental state variable to support targeted decision optimization. Combining the reward signal with the current environmental state, the system updates the state value function, employing temporal difference learning, reinforcement learning, or policy gradient-based optimization methods to dynamically adjust the state-action mapping relationship, thereby improving the system's decision-making accuracy and operational performance in the current environment.
[0108] The system optimizes the decision network using the updated state value function. Based on the backpropagation algorithm, gradient update mechanism, or structure adaptation strategy, it dynamically adjusts the parameter weights, network structure, or activation mode of the decision network to generate an updated version of the decision network, thereby improving the system's environmental adaptability, decision accuracy, and operational flexibility.
[0109] In optimizing the decision-making network, a parameter update strategy based on deep reinforcement learning can be adopted. This involves dynamically adjusting network parameters based on environmental feedback information to improve the system's adaptability to environmental changes and the robustness of operational strategies. The environmental perception model can be a multi-level dynamic Bayesian network structure that combines visual, audio, and state information to dynamically infer trends in environmental state changes, supporting efficient perception and accurate decision-making in complex dynamic environments.
[0110] The state value function can be updated using Q-learning, SARSA, or deep deterministic policy methods based on policy gradients. By observing data and reward signals in real time, the state-action value mapping relationship can be dynamically adjusted to improve the system's ability to optimize operational policies in uncertain environments.
[0111] The updated decision network can be based on a structurally variable neural network, an attention-enhanced structure, or a multimodal fusion mechanism, combined with dynamically adjusted network parameters and structural strategies, to form an intelligent network structure with stronger environmental perception, better information fusion capabilities, and higher decision-making efficiency, supporting the system to achieve flexible and efficient autonomous decision-making and task execution in complex environments.
[0112] Example Description: In the healthcare business field, the system dynamically optimizes the decision network structure of rehabilitation robots, surgical aids, or intelligent nursing systems based on environmental feedback information. By combining patient movement status, changes in environmental status, and operational feedback, the system adjusts operational strategies and control parameters in real time, improving the autonomous operation capability, intelligent collaboration level, and human-computer interaction performance of medical equipment, and ensuring safety, flexibility, and efficiency in dynamic medical environments.
[0113] In the fintech business, the system combines environmental feedback information to dynamically update the decision network of smart terminals, remote operation devices, or business automation systems. Based on changes in customer needs, fluctuations in the business environment, and feedback on operational results, the system optimizes system strategies and service processes in real time, improving the environmental adaptability, business response speed, and customer service level of the fintech system. This ensures that the system has continuous and stable intelligent decision-making capabilities and efficient business execution capabilities in complex and ever-changing financial business scenarios.
[0114] This embodiment optimizes the decision network based on environmental feedback information to generate an updated decision network. In dynamic and complex environments, the system can combine multimodal environmental feedback information and task execution results to adjust the decision network parameter structure and optimization strategy in real time, continuously improving environmental adaptability, information fusion efficiency and decision accuracy. This enhances the overall intelligence level, task completion ability and environmental coordination performance of the system, enabling the embodied intelligent system to achieve efficient self-adaptation and autonomous optimization in changing environments.
[0115] This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as embodied intelligence, fintech, and healthcare. It discloses a method, apparatus, device, and medium for generating and optimizing action commands, including: collecting environmental visual information, environmental audio information, and action state information; processing the environmental visual and audio information to obtain visual features and audio features respectively; using a multimodal attention mechanism to dynamically weight and fuse the visual features, audio features, and action state information to generate comprehensive decision features; inputting the comprehensive decision features into a decision network to generate action commands; executing the action commands and collecting environmental feedback information; optimizing the decision network based on the environmental feedback information; and generating an updated decision network. This invention improves the perception, decision-making, and action coordination level of embodied intelligence systems in complex environments through the dynamic fusion and unified decision-making of multimodal information, realizing efficient decision-making and flexible execution capabilities of robots in changing scenarios.
[0116] In one embodiment, step S10 above includes:
[0117] S101, acquires environmental visual information including spatial depth information through a visual sensor;
[0118] S102, acquires environmental audio information containing spectral characteristics through an audio sensor;
[0119] S103 acquires joint angle information through a motion sensor and motion trajectory information through a position sensor;
[0120] S104, integrate the joint angle information and motion trajectory information to form motion state information;
[0121] S105, perform a timestamp alignment operation on the environmental visual information, environmental audio information and action state information to generate aligned environmental visual information, aligned environmental audio information and aligned action state information respectively.
[0122] In this embodiment, environmental visual information, environmental audio information, and motion state information are collected to obtain multi-source data reflecting the external environment and the motion state of the operating subject, serving as the foundation for multimodal information fusion and subsequent intelligent decision-making. Environmental visual information refers to image data acquired by visual sensors, which can be imaging devices with depth perception capabilities, such as structured light cameras, binocular vision systems, or LiDAR integrated modules. Spatial depth information refers to data reflecting the relative spatial distance, three-dimensional structural contours, or stereoscopic shapes between objects and sensors in the environment. This data originates from multi-angle or multi-frequency acquisition results of the environmental scene by the visual sensors and provides data support for subsequent spatial structure modeling and dynamic environmental analysis.
[0123] Environmental audio information is sound data acquired through audio sensors, including single microphones, array microphones, or integrated acoustic modules. Spectral characteristics refer to the data extracted from the frequency distribution, energy distribution, or time-domain characteristics of sound signals through Fourier transform, Mel-frequency cepstral coefficient extraction, or other audio signal processing algorithms. These characteristics reflect background noise, voice commands, the operating status of mechanical equipment, or other acoustic information in the environment, providing supplementary audio-level information for subsequent multimodal fusion and environmental state analysis.
[0124] Motion state information is a data set reflecting the state of the operating entity, composed of joint angle information and motion trajectory information. Joint angle information originates from real-time measurements by motion sensors, which can be rotary encoders, angle sensors, or inertial measurement units, recording the spatial attitude changes of each joint in the mechanical structure in real time. Motion trajectory information is acquired through position sensors, including laser positioning devices, visual odometry, or GPS modules, used to capture spatial position changes at the end effector or key nodes, forming complete motion trajectory data.
[0125] In the process of integrating joint angle information and motion trajectory information to form motion state information, the system, based on time synchronization mechanism and data fusion algorithm, combines the output data of various sensors to construct a spatial motion description under a unified time reference, thereby achieving dynamic, continuous and comprehensive perception of the state of the operating subject.
[0126] To ensure temporal consistency and spatial correlation of different data types during multimodal fusion, the system performs timestamp alignment on environmental visual information, environmental audio information, and action state information. Timestamp alignment adjusts the timestamps of various data types using a unified time base or synchronization signal, ensuring that visual, audio, and action data correspond to the same environmental state and action at a unified time, eliminating sensor latency differences and inconsistent data acquisition frequencies. After timestamp alignment, the system generates aligned environmental visual information, aligned environmental audio information, and aligned action state information, ensuring temporal synchronization and spatial correspondence of data inputs for subsequent information fusion, feature extraction, and intelligent decision-making processes, thereby improving the accuracy of information processing and the overall decision-making effectiveness of the system.
[0127] In this embodiment, the environmental visual information, environmental audio information, and motion state information acquired through the above methods possess spatial structural integrity, acoustic information accuracy, and motion state continuity. The introduction of spatial depth information enhances the system's ability to perceive the three-dimensional structure of the environment and the spatial position of objects. The extraction of spectral features strengthens the system's ability to distinguish between sound environments, speech information, and noise interference. The integration of motion state information enables comprehensive perception of the operator's posture and trajectory. Through timestamp alignment, the system ensures the temporal synchronization and spatial correspondence of various types of information, eliminating time deviations between different data sources, improving the reliability and accuracy of multimodal data fusion, and ultimately providing a high-quality, synchronous, and consistent data foundation for multimodal information fusion and intelligent decision-making.
[0128] In one embodiment, step S20 above includes:
[0129] S201, Perform spatial feature extraction operation on the environmental visual information to generate initial visual features;
[0130] S202, Perform feature enhancement operation on the initial visual features to generate enhanced visual features, and use the enhanced visual features as the final visual features;
[0131] S203, Perform a spectral feature extraction operation on the environmental audio information to generate initial audio features;
[0132] S204, Perform feature optimization operation on the initial audio features to generate optimized audio features, and use the optimized audio features as the final audio features.
[0133] In this embodiment, environmental visual information refers to image data related to the external environment, spatial structure, and object contours collected by a visual sensor. Environmental visual information can be a single two-dimensional image sequence, or three-dimensional point cloud data or RGB-D image information with depth information. The system performs spatial feature extraction on the environmental visual information, specifically using feature extraction models based on convolutional neural networks, graph neural networks, or visual Transformer structures to extract spatial features related to object edges, spatial structure, geometric contours, and lighting textures. The spatial feature extraction operation can extract local texture features and global spatial structure information layer by layer for multi-scale image information, ultimately generating initial visual features. The initial visual features are a first-stage abstract expression of the environmental visual information, preserving the morphological features and spatial distribution information in the environmental scene, providing a data foundation for subsequent enhancement processing.
[0134] During the feature enhancement process on the initial visual features, the system introduces generative adversarial networks, variational autoencoders, or feature reconstruction networks based on regularization constraints. This addresses issues such as missing information, noise interference, or insufficient spatial resolution in the initial visual features by performing multi-level feature supplementation, information enhancement, and noise suppression. The feature enhancement operation generates enhanced visual features by reconstructing complete spatial structure information and optimizing the robustness and discriminativeness of feature representation. These enhanced visual features retain key geometric structures and spatial relationships from the original environmental visual information while improving the stability of feature representation under different lighting conditions and complex background environments. The system then uses these enhanced visual features as the final visual feature input to subsequent multimodal fusion or decision networks.
[0135] Ambient audio information refers to environmental sound data collected by audio sensors, including various acoustic signals such as background noise, speech, and equipment operating noise. Spectral feature extraction is performed on this environmental audio information. Specifically, this involves using signal processing algorithms such as Mel-frequency transform, short-time Fourier transform, and continuous wavelet transform to convert the time-domain audio signal into a frequency-domain representation, extracting initial audio features that include the frequency distribution, energy spectrum distribution, and temporal variation trends of the audio signal. These initial audio features reflect the sound source structure, speech characteristics, or background noise patterns in the environmental audio information, possessing the ability to describe the acoustic environment state and audio information characteristics.
[0136] During the feature optimization process on the initial audio features, the system applies feature modeling methods based on flow models, probabilistic generative models, or deep representation learning networks to filter out, compress, and reconstruct noise interference, information redundancy, and unstable representations in the initial audio features, thereby enhancing the discriminative ability and robustness of the audio features. The feature optimization operation generates optimized audio features, which possess highly stable and expressive acoustic information description capabilities, effectively preserving key spectral features and dynamic trends in the environmental audio information. The system uses these optimized audio features as the final audio feature input to subsequent information fusion or decision-making modules.
[0137] In this embodiment, the system performs multi-level processing and enhancement on the structural and spectral features of both environmental visual and audio information using the methods described above. The final output visual features possess complete spatial structure representation and stable geometric information, while the output audio features possess accurate spectral information and highly robust acoustic representation capabilities. Spatial feature extraction and feature enhancement operations improve the stability and discriminative ability of environmental visual information under complex backgrounds and dynamic lighting conditions. Spectral feature extraction and feature optimization operations enhance the information retention and accuracy of environmental audio information in multi-noise interference environments. Overall, this improves the data foundation quality for multimodal information fusion and intelligent decision-making processes, enhancing the system's adaptability and decision accuracy in dynamic and complex environments.
[0138] In one embodiment, step S30 above includes:
[0139] S301, Perform a first linear transformation operation on the visual features to generate a first transformed feature;
[0140] S302, Perform a second linear transformation operation on the audio features to generate second transformed features;
[0141] S303, Perform a third linear transformation operation on the action state information to generate a third transformation feature;
[0142] S304, determine the first similarity weight between the first transformation feature and the second transformation feature, the second similarity weight between the first transformation feature and the third transformation feature, and the third similarity weight between the second transformation feature and the third transformation feature;
[0143] S305, Construct a multimodal attention weight matrix based on the first similarity weight, the second similarity weight, and the third similarity weight;
[0144] S306, the first transformation feature, the second transformation feature, and the third transformation feature are weighted and fused according to the multimodal attention weight matrix to generate a comprehensive decision feature.
[0145] In this embodiment, visual features refer to high-dimensional feature vectors that characterize the spatial structure, geometric contours, and morphological information in environmental visual information after spatial feature extraction and enhancement. Audio features refer to feature vectors that describe the spectral distribution, energy changes, and acoustic patterns in environmental audio information after spectral feature extraction and optimization. Action state information is a data set formed based on the action execution process, combined with joint angle information and motion trajectory information, used to reflect the current motion state, position changes, and action process parameters of the actuator.
[0146] Performing the first linear transformation operation on visual features refers to mapping the visual features to a unified feature space through matrix multiplication, linear mapping networks, or fully connected structures, generating the first transformed features. This operation achieves dimensional adjustment and spatial structure standardization of visual features, ensuring the comparability of information from different modalities within the same mathematical space. The first transformed features preserve the key geometric information and spatial representation structure of the visual features, facilitating subsequent multimodal fusion computation.
[0147] During the second linear transformation operation on the audio features, the system employs matrix mapping, feature reconstruction networks, or parametric linear transformations similar to the first linear transformation structure to map the audio features into a unified feature space, generating the second transformed features. While preserving the spectral information and dynamic trends of the audio features, the second transformed features adjust the feature representation structure and dimensions, providing input data with a unified format for multimodal information fusion.
[0148] During the third linear transformation operation on the motion state information, the system constructs a linear mapping function or feature compression network for the joint angle and motion trajectory data in the motion state information, and converts the motion state information into a third transformation feature with a standardized structure and consistent dimensions, so as to ensure that the motion state information can participate in the unified feature fusion calculation with visual features and audio features.
[0149] In determining the first similarity weight between the first transformed feature and the second transformed feature, the system quantifies the correlation between visual and audio features through dot product operations, cosine similarity calculations, or a scoring function based on an attention mechanism, thereby generating the first similarity weight. The first similarity weight reflects the degree of interaction and correlation between visual and audio information in the current environment.
[0150] In determining the second similarity weight between the first and third transformation features, the system uses a scoring function for similar structures and a correlation calculation method to quantify the degree of correlation between visual features and action state information, and generates the second similarity weight to reflect the matching relationship between visual information and the current action state.
[0151] In determining the third similarity weight between the second and third transformation features, the system also applies the correlation quantification method to calculate the correlation between audio features and motion state information, and generates the third similarity weight to describe the interaction between acoustic information and motion state.
[0152] Based on the first similarity weight, the second similarity weight, and the third similarity weight, the system constructs a multimodal attention weight matrix. This matrix has a symmetric or asymmetric structure, and its internal elements are composed of the similarity weights between the features of each modality, reflecting the dynamic weight ratio of multimodal information in the overall fusion process.
[0153] Based on the multimodal attention weight matrix, the system performs a weighted fusion operation on the first, second, and third transformation features. This includes matrix multiplication, feature concatenation, weighted summation, or multi-level information fusion based on a self-attention mechanism. During the fusion process, different modal features contribute information according to their respective weight ratios, ultimately generating a comprehensive decision feature. This comprehensive decision feature integrates the spatial structure information of visual features, the spectral information of audio features, and the motion state description of action state information, achieving dynamic weighted fusion of multimodal information and improving the system's perception and decision-making capabilities in complex environments.
[0154] In this embodiment, the system performs a dynamic weighted fusion operation on visual features, audio features, and action state information using a multimodal attention mechanism. The multimodal attention weight matrix, constructed based on similarity weights, effectively quantifies the correlation and relative importance between different modalities. During the fusion process, the contribution of each modal feature is dynamically adjusted according to its weight ratio. The resulting comprehensive decision feature integrates multi-dimensional environmental information and real-time action state information. Compared to decision-making models based on a single information source, this approach enhances the comprehensive perception of environmental states and the system's own action processes by fusing multi-source heterogeneous information, thereby improving the system's adaptive decision-making level and the accuracy of action execution in dynamic and complex environments.
[0155] In one embodiment, step S40 above includes:
[0156] S401, input the comprehensive decision features into the decision network;
[0157] S402, The comprehensive decision features are processed through the multi-head attention mechanism of the decision network to generate attention-optimized features;
[0158] S403, the feedforward neural network of the decision network is applied to transform the attention optimization features to generate transformed features;
[0159] S404, the transformation features are processed through the output layer of the decision network to generate action instruction codes, and the action instruction codes are parsed into action instructions.
[0160] In this embodiment, the comprehensive decision feature is a high-dimensional feature vector formed by fusing visual features, audio features, and action state information through a multimodal attention mechanism. It integrates spatial structure information, acoustic information, and action process parameters of the environment, possessing the ability to express multi-source heterogeneous information and comprehensively reflecting the current environment and its own state. The decision network is a deep neural network structure with feature processing and action decision-making capabilities. It includes a multi-level encoding structure, a multi-head attention mechanism, a feedforward neural network, and an output layer. The overall structure possesses information abstraction, feature extraction, and action generation functions.
[0161] During the process of inputting comprehensive decision features into the decision network, the system transmits the comprehensive decision features to the inside of the decision network through data interfaces, feature input modules, or embedding layers, ensuring that the information can be effectively received and processed in the network structure.
[0162] In processing comprehensive decision features using a multi-head attention mechanism within a decision network, the system, based on a multi-head parallel structure, performs multiple independent attention calculations for different information dimensions of the comprehensive decision features, generating attention outputs within multiple subspaces. The multi-head attention mechanism employs query vectors, key-value vectors, and scoring functions to calculate the dynamic correlation weights between elements within the comprehensive decision features, capturing high-order dependencies between information and enhancing the richness and relevance of feature representation. The results from multiple attention heads are merged through concatenation or weighted operations to generate attention-optimized features. These attention-optimized features, while maintaining the original information structure, enhance the interactive expressive capabilities between multi-dimensional information.
[0163] In the process of transforming attention-optimized features using a feedforward neural network applied to a decision network, the system employs nonlinear activation functions, hierarchical structures, and parameterized mapping operations to perform nonlinear transformations and representation reconstructions on the attention-optimized features in a high-dimensional feature space, generating transformed features. The feedforward neural network includes one or more layers of parameter matrices, bias vectors, and nonlinear activation units, possessing the functions of feature compression, information recombination, and representation enhancement, ensuring that the transformed features meet the requirements for action command generation in terms of dimensional adjustment and information representation structure.
[0164] During the processing of transformed features through the output layer of the decision network, the system uses parameter matrices, mapping functions, or specific output structures to convert the transformed features into action command codes. Action command codes are intermediate representations of action commands, typically in vector form, containing information such as specific action categories, action parameters, and target states. In the parsing of action command codes into action commands, the system, based on preset encoding parsing rules or mapping functions, restores the information in the action command codes into standardized action commands with execution meaning, specifically including robotic arm position adjustment, gripper control, and end effector motion planning. These action commands have the ability to directly drive the actuators to complete specific operations, ultimately acting on the actual environment to complete the predetermined action task.
[0165] This embodiment achieves high-order information representation, dynamic correlation modeling, and action parameter generation through a multi-layered structure within the decision network. A multi-head attention mechanism enhances the interaction between different information dimensions, while a feedforward neural network optimizes the structure and content of information representation. The output layer and encoding / parsing process efficiently transform high-dimensional feature information into executable standard action instructions. This process realizes a complete information flow path from multi-source heterogeneous information to specific action instructions, improving the system's action planning capability and execution accuracy in complex environments, and enhancing its adaptive response to changing environments.
[0166] In one embodiment, step S50 above includes:
[0167] S501, the action corresponding to the action command is executed by the execution device;
[0168] S502, collect environmental visual change information after the action is performed using a visual sensor;
[0169] S503, collect environmental audio change information after the action is performed using an audio sensor;
[0170] S504, Obtain environmental state change data after the action is executed through the state monitoring module;
[0171] S505, obtain the reward signal after the action is executed through the reward function module;
[0172] S506, integrate the environmental visual change information, environmental audio change information, environmental state change data and reward signal to generate environmental feedback information.
[0173] In this embodiment, the action command is control information with clear execution parameters generated based on multimodal information fusion and decision networks, specifically manifested as position adjustment, posture control, motion trajectory planning, or mechanical action commands. The execution device is a hardware structure with physical action capabilities, including a robotic arm, end effector, mobile platform, or other operating components of an embodied intelligent system. During the execution of the action command corresponding to the action command by the execution device, the system inputs the action command to the control unit of the execution device, driving the corresponding mechanical structure to complete position adjustment, grasping, transporting, placing, or interactive operations, ensuring the actual output of the action command at the physical level.
[0174] Environmental visual change information refers to the visible changes that occur in the environment during or after an action is performed, and it originates from visual sensors. Visual sensors may include RGB cameras, depth cameras, structured light devices, or LiDAR. During the acquisition of environmental visual change information through visual sensors, the system sets a reference image frame before the action is performed and acquires updated image data after the action is performed. Based on comparative analysis, spatial reconstruction, or change detection algorithms, information such as positional changes, object state changes, and spatial layout adjustments in the environment is extracted to form environmental visual change information.
[0175] Environmental audio change information refers to acoustic changes in the environment during the execution of an action, originating from audio sensors. Audio sensors include single-microphone arrays, distributed pickup modules, or high-sensitivity audio detection devices. During the acquisition of environmental audio change information through audio sensors, the system monitors audio signals in real time during the execution of the action, identifying mechanical operating sounds, ambient background noise, operational feedback sounds, or audio markers of specific events, extracting changes in audio features, and forming environmental audio change information.
[0176] Environmental state change data refers to the changes in the environment's physical state, dynamic parameters, or spatial relationships after an action is performed, and it originates from the state monitoring module. The state monitoring module includes an inertial measurement unit, force control sensors, a positioning system, or an environmental sensor network. During the acquisition of environmental state change data, the system compares environmental state parameters before and after the action, analyzes position information, velocity data, acceleration information, mechanical feedback, or changes in system stability, and forms quantitatively expressed environmental state change data.
[0177] The reward signal is a numerical feedback parameter that measures the effectiveness of action execution and the quality of environmental response, originating from the reward function module. Based on preset task objectives, environmental response standards, and a system evaluation model, the reward function module dynamically calculates the reward signal according to real-time environmental changes and action results. During the acquisition of the reward signal, the system quantifies and generates positive or negative reward signals based on visual environmental changes, audio environmental changes, and environmental state changes, combined with task objective deviation, action completion rate, and environmental adaptability, to guide subsequent optimization processes.
[0178] Environmental feedback information is a structured integration of the aforementioned environmental change information and reward signals. In integrating visual environmental change information, audio environmental change information, environmental state change data, and reward signals, the system, based on a unified data interface and timestamp mechanism, fuses multi-source information to form environmental feedback information with temporal consistency, spatial correlation, and state integrity. This environmental feedback information comprehensively reflects the effectiveness of action execution and environmental changes, serving as the input basis for optimizing the decision-making network and improving system performance.
[0179] This embodiment establishes a complete feedback loop between action and environment through the physical execution of action commands, multi-dimensional real-time perception of environmental changes, and structured integration of multi-source information. Visual and audio information about environmental changes enhances the system's ability to perceive changes in the external environment, environmental state change data provides a quantitative expression of the system's internal interaction with the environment, and reward signals enable dynamic evaluation of action effects. By integrating various feedback information, the system can efficiently and accurately grasp changes in the environmental state after action execution, enhancing its adaptability and self-optimization capabilities in complex dynamic environments, and improving the environmental interaction efficiency and task execution reliability of the embodied intelligence system.
[0180] In one embodiment, step S60 above includes:
[0181] S601, extract real-time observation data and reward signals from the environmental feedback information;
[0182] S602, Based on the real-time observation data, the environmental perception model performs state probability inference to generate an updated environmental state probability distribution;
[0183] S603, determine the current environmental state based on the updated environmental state probability distribution;
[0184] S604, Update the state value function based on the reward signal and the current environmental state;
[0185] S605, optimize the decision network using the updated state value function to generate an updated decision network.
[0186] In this embodiment, the environmental feedback information is a structured collection of environmental change data, state monitoring results, and reward feedback information acquired by the multi-source sensing system after the action is executed. It possesses completeness, temporal correlation, and multimodal fusion characteristics. During the extraction of real-time observation data and reward signals from the environmental feedback information, the system uses a data parsing module to extract parameters directly related to the environmental state from the environmental visual change information, environmental audio change information, and environmental state change data as real-time observation data. This real-time observation data includes, but is not limited to, spatial location parameters, object state parameters, system dynamics parameters, and environmental constraints. The reward signal originates from quantitative feedback information generated by the reward function module, characterizing the quality level of the current action execution result.
[0187] The environmental perception model is a probabilistic inference structure for environmental states built based on real-time observation data. It is implemented using Bayesian networks, dynamic Bayesian networks, or Markov decision process-based state inference structures. The environmental perception model performs probabilistic state inference based on real-time observation data. The system combines historical observation information with current real-time observation data to establish a state transition model and an observation probability model, calculating the current state probability distribution of the environment. The environmental state probability distribution refers to the set of probability values for different environmental states, reflecting the uncertainty and multiple possibilities of environmental states. In generating an updated environmental state probability distribution, the system uses methods such as filtering inference, Bayesian updating, or maximum a posteriori estimation based on the previous environmental state probability distribution, the state transition model, and current observation information to output the updated environmental state probability distribution.
[0188] In determining the current environmental state, the system, based on the updated environmental state probability distribution, selects or combines the decision rules of the task scenario according to the maximum probability principle and confidence interval, and identifies the environmental state that best matches the current observation data and the evolution pattern of historical states. The current environmental state is the basic state variable for subsequent value assessment and decision optimization.
[0189] The state value function is a mathematical mapping relationship that quantifies the potential value of an environment state and the expected return level. In the process of updating the state value function based on the reward signal and the current environment state, the system adjusts the parameters or mapping relationship of the state value function according to the temporal difference update rule, Bellman equation or policy gradient estimation in reinforcement learning methods, combined with the current environment state and the corresponding reward signal, so that the state value function more accurately reflects the actual value level of the current environment and the long-term return expectation.
[0190] The updated state value function serves as the basis for optimizing the decision network. During the optimization process, the system combines the state-action value function, policy gradient information, or advantage function with gradient descent, parameter updates, or structural adjustments to enhance the decision network's decision-making performance and adaptability in the current environment. The optimized decision network possesses superior feature representation capabilities, policy output capabilities, and environmental adaptability. In generating the updated decision network, the system dynamically adjusts the network's structural parameters, connection weights, or output mapping relationships to ensure that the updated network exhibits higher intelligence and better environmental interaction in subsequent task executions.
[0191] Example: In the field of embodied intelligence, robots rely on the fusion perception of multimodal information, intelligent decision-making, and real-time self-optimization capabilities when performing dynamic tasks.
[0192] In complex logistics and warehousing environments, dual-arm collaborative robots undertake intelligent handling and packing tasks for irregular items. First, the system acquires environmental visual information within the current operating area through visual sensors, specifically including RGB images, depth images, and spatial structure data, comprehensively reflecting the geometric appearance and three-dimensional contours of the manipulated objects. Simultaneously, audio sensors collect environmental audio information from the operating environment, covering ambient noise, voice commands, or mechanical feedback sounds, assisting the system in understanding external changes and human-robot collaboration information. The robot's motion sensors continuously acquire joint angle information, recording the posture changes of the multi-degree-of-freedom robotic arms in real time, while position sensors synchronously monitor the motion trajectory of the end effector, ensuring the continuity of spatial position and dynamic path. The system integrates joint angle information and motion trajectory information to form complete motion state information, accurately expressing the robot's current motion state and spatial posture. To ensure the consistency of multi-source data, the system performs timestamp alignment on environmental visual information, environmental audio information, and motion state information. Based on a unified time reference, it generates aligned environmental visual information, aligned environmental audio information, and aligned motion state information, eliminating time delay differences between multimodal data and ensuring the consistency of the information fusion foundation.
[0193] Based on the aligned data, the system first performs spatial feature extraction on the environmental visual information, using a convolutional neural network structure to extract target contours, surface textures, and spatial layout information to form initial visual features. Subsequently, the system performs feature enhancement on these initial visual features, employing generative adversarial networks or variational autoencoders to improve their robustness and discriminative power, ultimately outputting enhanced visual features as the visual features for subsequent system fusion. During audio information processing, the system performs spectral feature extraction on the environmental audio information, extracting the frequency domain distribution and temporal dynamic characteristics of the audio signal to form initial audio features. Combining flow models and high-dimensional mapping techniques, the system performs feature optimization on the initial audio features, improving the ability to distinguish between speech, noise, and environmental sounds, ultimately generating optimized audio features as the audio features for system fusion processing.
[0194] To achieve dynamic fusion of multimodal information, the system employs a multimodal attention mechanism. First, linear transformations are performed on visual features, audio features, and action state information, generating first, second, and third transformed features. The system then calculates the similarity weights of three modal pairs—visual and audio, visual and action state, and audio and action state—using dot product operations and similarity functions to quantify the correlation strength between each modality. Based on the calculation results, a multimodal attention weight matrix is constructed to comprehensively reflect the dynamic correlation of multi-source information. The system then performs weighted fusion of each transformed feature according to this weight matrix, outputting a comprehensive decision feature that integrates information about the current environment, action state, and external changes.
[0195] The system inputs comprehensive decision features into a decision network, which includes a multi-level encoding structure and a multi-head attention mechanism. Through multi-level information abstraction and multi-dimensional attention calculation, the system generates attention-optimized features. A feedforward neural network further performs non-linear mapping and feature transformation on the optimized features, generating transformed features. Through the output layer, the system encodes the transformed features into standardized action command codes. The parsing process restores the encoded content to specific action command parameters and spatial paths, thus generating standard action commands.
[0196] The robot executes specified actions via actuators based on motion commands, such as adjusting the robotic arm's posture and manipulating grippers to perform grasping, handling, and packing tasks. During action execution, a vision sensor collects real-time visual information about environmental changes, an audio sensor acquires information about environmental audio changes, and a state monitoring module records environmental state change data, reflecting changes in the position, posture, and spatial layout of the manipulated object. The system uses a reward function module to generate reward signals based on task completion, operational accuracy, and environmental change results, quantifying the effectiveness of the action and the system's performance. The system integrates multi-source feedback data to generate structured environmental feedback information, comprehensively expressing environmental changes, system state, and operational effectiveness.
[0197] The system extracts real-time observation data and reward signals based on environmental feedback information. Through an environmental perception model combining historical data and current observation information, it performs state probabilistic inference to calculate the updated environmental state probability distribution, accurately characterizing the uncertainty and dynamic changes of the environmental state. Combining the state distribution and reward signals, the system updates the state value function, dynamically adjusts the environmental state value assessment model, and quantifies the expected return level under different states. Based on the updated state value function, the system optimizes the structural parameters and output mapping of the decision network, dynamically improving the network's decision-making performance and environmental adaptability, generating an updated decision network. The updated decision network possesses stronger multimodal information processing capabilities, dynamic decision-making capabilities, and adaptability to complex environments. The system performs subsequent operations based on this network, enabling the robot to autonomously learn, dynamically optimize, and make intelligent decisions in changing environments, significantly improving task completion efficiency, operational accuracy, and the overall intelligence level of the system.
[0198] In the field of healthcare, embodied intelligent systems are widely used in complex scenarios such as assisted surgery, minimally invasive interventions, rehabilitation training, and intelligent nursing.
[0199] During the operation of the minimally invasive surgical robot, the system first acquires environmental visual information of the surgical area through high-precision vision sensors, including three-dimensional spatial structure, tissue boundaries, and surface morphology information of the target organ, ensuring that the system has a comprehensive grasp of the geometric details of the surgical environment. Simultaneously, high-sensitivity audio sensors collect environmental audio information, monitoring in real time the sounds generated during surgery, such as instrument contact sounds, patient physiological sounds, and other environmental sound sources, assisting in assessing tissue contact status and potential risks. To monitor the motion state of the operating end in real time, the system acquires joint angle information of the robotic arm through motion sensors, accurately reflecting the dynamic posture of the multi-degree-of-freedom mechanical structure. At the same time, position sensors acquire motion trajectory information of the surgical instrument's end effector, recording the operation path and spatial position in real time. The system integrates joint angle information and motion trajectory information to form motion state information expressing the complete motion state. To avoid temporal deviations in multi-source data, the system performs timestamp alignment on environmental visual information, environmental audio information, and motion state information, ensuring that all types of information are updated synchronously based on a unified timeline, eliminating data delay and asynchrony issues, and providing a reliable foundation for downstream multimodal fusion.
[0200] Based on time-aligned data, the system performs spatial feature extraction on environmental visual information, utilizing deep neural networks to extract the spatial relationships between tissue structure, anatomical features, and the surgical environment to form initial visual features. To improve the system's robustness under complex lighting and tissue deformation conditions, feature enhancement operations are further performed on the initial visual features. Based on autoencoder networks or generative adversarial structures, the feature representation capability is optimized, outputting enhanced visual features. For audio information, the system performs spectral feature extraction, analyzing the frequency domain changes of audio signals in the surgical environment to obtain key acoustic features, forming initial audio features. Through feature optimization operations, the stability of the audio features under background noise and complex sound source interference is enhanced, outputting optimized audio features.
[0201] The system employs a multimodal attention mechanism to fuse visual features, audio features, and action state information. First, it performs linear transformation operations to generate corresponding first, second, and third transformation features, quantifying the representation of each modality in the feature space. Based on intermodal similarity calculations, the system determines three sets of similarity weights: visual and audio, visual and action state, and audio and action state. These weights are then combined to construct a multimodal attention weight matrix, dynamically reflecting the importance distribution of each modality in specific operational scenarios. The system then uses this matrix to weight and fuse the multimodal information, generating comprehensive decision features that express the overall perception result.
[0202] The system inputs comprehensive decision features into the decision network. Through a multi-level coding structure and a multi-head attention mechanism, it mines the deep correlations of multimodal information and outputs attention-optimized features. The feedforward neural network further performs a nonlinear transformation on these features to generate transition features for motion control. The system encodes these transition features into standardized motion command codes through the output layer and parses them into specific motion command parameters to guide the surgical robot in performing delicate operations.
[0203] During surgical procedures, the system controls the robotic arm and end-effectors via an actuator to perform tissue cutting, suturing, positioning, and other auxiliary operations. The system uses visual sensors to collect real-time information on environmental visual changes, monitoring the surgical field and target tissue status. Audio sensors collect acoustic feedback during surgery to help determine tissue contact or equipment malfunctions. A status monitoring module acquires environmental status change data, and combined with mechanical structure feedback and environmental change information, the system dynamically adjusts its operational strategies. Based on the degree of operation completion, tissue protection, and process safety, a reward function module generates a reward signal to evaluate the current operational effect and system behavior. The system integrates visual, audio, status change, and reward signals to form multi-dimensional environmental feedback information, comprehensively reflecting the operational effect and environmental changes.
[0204] Based on environmental feedback, the system extracts real-time observation data and reward signals, and performs state probabilistic inference using an environmental perception model to generate a dynamically updated environmental state probability distribution. The system accurately judges the organization's state, spatial location, and environmental change trends based on this distribution. Combining the updated environmental state and reward signals, the system updates the state value function, optimizes the expected return assessment under different environmental states, and further adjusts the decision network structure and parameters through the state value function, outputting an updated decision network. This updated decision network possesses stronger multimodal information fusion capabilities, dynamic decision-making capabilities, and environmental adaptability. The system performs subsequent surgical operations based on this network, improving the robot's autonomous decision-making level, operational stability, and safety in complex minimally invasive surgical environments, significantly enhancing surgical efficiency, operational accuracy, and patient safety.
[0205] This embodiment extracts real-time observation data and reward signals from environmental feedback information, enabling the system to effectively acquire multi-source information reflecting environmental changes and action execution effects. The environmental perception model combines observation data to complete state probability reasoning, enhancing the system's ability to express the uncertainty of complex dynamic environments. The updated environmental state probability distribution accurately characterizes the multiple possibilities of the current environmental state. Based on the update of the state value function of the current environmental state and reward signals, the system's ability to quantify the value of environmental states is dynamically improved. The updated state value function is used to optimize the decision network, enhancing the system's autonomous learning and decision optimization capabilities in changing environments. The generated updated decision network significantly improves the system's environmental adaptability, decision accuracy, and execution effect in complex task scenarios. Overall, the system possesses a stronger level of intelligence and task completion capability.
[0206] In one embodiment, a motion command generation and optimization apparatus is provided, which corresponds one-to-one with the motion command generation and optimization method described in the above embodiments. (Refer to...) Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the action command generation and optimization device of the present invention. The modules include a perception and acquisition module 10, a feature extraction module 20, a fusion decision module 30, a command generation module 40, a behavior execution module 50, and a learning and update module 60. Detailed descriptions of each functional module are as follows:
[0207] The perception and acquisition module 10 is used to collect environmental visual information, environmental audio information, and motion status information.
[0208] Feature extraction module 20 is used to process the environmental visual information and the environmental audio information to obtain visual features and audio features respectively;
[0209] The fusion decision module 30 is used to dynamically weight and fuse the visual features, audio features, and action state information using a multimodal attention mechanism to generate comprehensive decision features;
[0210] Instruction generation module 40 is used to input the comprehensive decision features into the decision network and generate action instructions;
[0211] Behavior execution module 50 is used to execute the action instructions and collect environmental feedback information;
[0212] The learning and updating module 60 is used to optimize the decision network based on the environmental feedback information and generate an updated decision network.
[0213] In one embodiment, the sensing and acquisition module 10 is specifically used for:
[0214] Acquire environmental visual information, including spatial depth information, using visual sensors;
[0215] Acquire environmental audio information containing spectral characteristics using an audio sensor;
[0216] Joint angle information is obtained through motion sensors, and motion trajectory information is obtained through position sensors;
[0217] The joint angle information and motion trajectory information are integrated to form motion state information;
[0218] Perform a timestamp alignment operation on the environmental visual information, environmental audio information, and action state information to generate aligned environmental visual information, aligned environmental audio information, and aligned action state information, respectively.
[0219] In one embodiment, the feature extraction module 20 is specifically used for:
[0220] Perform spatial feature extraction on the environmental visual information to generate initial visual features;
[0221] Perform feature enhancement operations on the initial visual features to generate enhanced visual features, and use the enhanced visual features as the final visual features;
[0222] Perform a spectral feature extraction operation on the environmental audio information to generate initial audio features;
[0223] Perform feature optimization on the initial audio features to generate optimized audio features, and use the optimized audio features as the final audio features.
[0224] In one embodiment, the fusion decision module 30 is specifically used for:
[0225] Perform a first linear transformation operation on the visual features to generate a first transformed feature;
[0226] Perform a second linear transformation operation on the audio features to generate second transformed features;
[0227] Perform a third linear transformation operation on the action state information to generate a third transformation feature;
[0228] Determine the first similarity weight between the first transformation feature and the second transformation feature, the second similarity weight between the first transformation feature and the third transformation feature, and the third similarity weight between the second transformation feature and the third transformation feature;
[0229] A multimodal attention weight matrix is constructed based on the first similarity weight, the second similarity weight, and the third similarity weight;
[0230] The first transformation feature, the second transformation feature, and the third transformation feature are weighted and fused according to the multimodal attention weight matrix to generate a comprehensive decision feature.
[0231] In one embodiment, the instruction generation module 40 is specifically used for:
[0232] The comprehensive decision-making features are input into the decision network;
[0233] The comprehensive decision features are processed through the multi-head attention mechanism of the decision network to generate attention-optimized features;
[0234] The attention optimization features are transformed using the feedforward neural network of the decision network to generate transformed features;
[0235] The transformation features are processed by the output layer of the decision network to generate action instruction codes, and the action instruction codes are parsed into action instructions.
[0236] In one embodiment, the behavior execution module 50 is specifically used for:
[0237] The action corresponding to the action command is executed by the execution device;
[0238] The visual sensor collects information on environmental visual changes after the action is performed.
[0239] The environmental audio changes after the action is performed are collected using an audio sensor.
[0240] The environmental state change data after the action is executed is obtained through the state monitoring module;
[0241] The reward signal after the action is executed is obtained through the reward function module;
[0242] By integrating the environmental visual change information, environmental audio change information, environmental state change data, and reward signals, environmental feedback information is generated.
[0243] In one embodiment, the learning update module 60 is specifically used for:
[0244] Real-time observation data and reward signals are extracted from the environmental feedback information;
[0245] Based on the real-time observation data, the environmental perception model performs state probability inference to generate an updated environmental state probability distribution.
[0246] The current environmental state is determined based on the updated environmental state probability distribution;
[0247] Update the state value function based on the reward signal and the current environmental state;
[0248] The updated decision network is then optimized using the updated state value function to generate an updated decision network.
[0249] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown. The computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used for communication with external user terminals via a network connection. When the computer program is executed by the processor, it implements the functions or steps of an action instruction generation and optimization method on the server side.
[0250] In one embodiment, a computer device is provided, which may be a user terminal, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements the functions or steps of an action instruction generation and optimization method on the user side.
[0251] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:
[0252] Collect environmental visual information, environmental audio information, and motion status information;
[0253] The environmental visual information and the environmental audio information are processed to obtain visual features and audio features, respectively.
[0254] A multimodal attention mechanism is used to dynamically weight and fuse the visual features, audio features, and action state information to generate comprehensive decision features;
[0255] The comprehensive decision features are input into the decision network to generate action instructions;
[0256] Execute the action instructions and collect environmental feedback information;
[0257] The decision network is optimized based on the environmental feedback information to generate an updated decision network.
[0258] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:
[0259] Collect environmental visual information, environmental audio information, and motion status information;
[0260] The environmental visual information and the environmental audio information are processed to obtain visual features and audio features, respectively.
[0261] A multimodal attention mechanism is used to dynamically weight and fuse the visual features, audio features, and action state information to generate comprehensive decision features;
[0262] The comprehensive decision features are input into the decision network to generate action instructions;
[0263] Execute the action instructions and collect environmental feedback information;
[0264] The decision network is optimized based on the environmental feedback information to generate an updated decision network.
[0265] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0266] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0267] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0268] It should be noted that if any software tools or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The embodiments described above are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for generating and optimizing action instructions, characterized in that, Includes the following steps: Collect environmental visual information, environmental audio information, and motion status information; The environmental visual information and the environmental audio information are processed to obtain visual features and audio features, respectively. A multimodal attention mechanism is used to dynamically weight and fuse the visual features, audio features, and action state information to generate comprehensive decision features; The comprehensive decision features are input into the decision network to generate action instructions; Execute the action instructions and collect environmental feedback information; The decision network is optimized based on the environmental feedback information to generate an updated decision network.
2. The motion instruction generation and optimization method as described in claim 1, characterized in that, Collect environmental visual information, environmental audio information, and action status information, including: Acquire environmental visual information, including spatial depth information, using visual sensors; Acquire environmental audio information containing spectral characteristics using an audio sensor; Joint angle information is obtained through motion sensors, and motion trajectory information is obtained through position sensors; The joint angle information and motion trajectory information are integrated to form motion state information; Perform a timestamp alignment operation on the environmental visual information, environmental audio information, and action state information to generate aligned environmental visual information, aligned environmental audio information, and aligned action state information, respectively.
3. The motion instruction generation and optimization method as described in claim 1, characterized in that, Processing the environmental visual information and the environmental audio information to obtain visual features and audio features respectively includes: Perform spatial feature extraction on the environmental visual information to generate initial visual features; Perform feature enhancement operations on the initial visual features to generate enhanced visual features, and use the enhanced visual features as the final visual features; Perform a spectral feature extraction operation on the environmental audio information to generate initial audio features; Perform feature optimization on the initial audio features to generate optimized audio features, and use the optimized audio features as the final audio features.
4. The motion instruction generation and optimization method as described in claim 1, characterized in that, A multimodal attention mechanism is employed to dynamically weight and fuse the visual features, audio features, and action state information to generate comprehensive decision features, including: Perform a first linear transformation operation on the visual features to generate a first transformed feature; Perform a second linear transformation operation on the audio features to generate second transformed features; Perform a third linear transformation operation on the action state information to generate a third transformation feature; Determine the first similarity weight between the first transformation feature and the second transformation feature, the second similarity weight between the first transformation feature and the third transformation feature, and the third similarity weight between the second transformation feature and the third transformation feature; A multimodal attention weight matrix is constructed based on the first similarity weight, the second similarity weight, and the third similarity weight; The first transformation feature, the second transformation feature, and the third transformation feature are weighted and fused according to the multimodal attention weight matrix to generate a comprehensive decision feature.
5. The motion instruction generation and optimization method as described in claim 1, characterized in that, The comprehensive decision features are input into the decision network to generate action instructions, including: The comprehensive decision-making features are input into the decision-making network; The comprehensive decision features are processed through the multi-head attention mechanism of the decision network to generate attention-optimized features; The attention optimization features are transformed using the feedforward neural network of the decision network to generate transformed features; The transformation features are processed by the output layer of the decision network to generate action instruction codes, and the action instruction codes are parsed into action instructions.
6. The motion instruction generation and optimization method as described in claim 1, characterized in that, Executing the action instructions and collecting environmental feedback information includes: The action corresponding to the action command is executed by the execution device; The visual sensor collects information on environmental visual changes after the action is performed. The environmental audio changes after the action is performed are collected using an audio sensor. The environmental state change data after the action is executed is obtained through the state monitoring module; The reward signal after the action is executed is obtained through the reward function module; By integrating the environmental visual change information, environmental audio change information, environmental state change data, and reward signals, environmental feedback information is generated.
7. The motion instruction generation and optimization method as described in claim 1, characterized in that, The decision network is optimized based on the environmental feedback information to generate an updated decision network, including: Real-time observation data and reward signals are extracted from the environmental feedback information; Based on the real-time observation data, the environmental perception model performs state probability inference to generate an updated environmental state probability distribution. The current environmental state is determined based on the updated environmental state probability distribution; Update the state value function based on the reward signal and the current environmental state; The updated decision network is then optimized using the updated state value function to generate an updated decision network.
8. A motion command generation and optimization device, characterized in that, The action command generation and optimization device includes: The perception and acquisition module is used to collect environmental visual information, environmental audio information, and motion status information; The feature extraction module is used to process the environmental visual information and the environmental audio information to obtain visual features and audio features respectively; The fusion decision module is used to dynamically weight and fuse the visual features, audio features, and action state information using a multimodal attention mechanism to generate comprehensive decision features; The instruction generation module is used to input the comprehensive decision features into the decision network and generate action instructions; The behavior execution module is used to execute the action instructions and collect environmental feedback information; The learning and updating module is used to optimize the decision network based on the environmental feedback information and generate an updated decision network.
9. A computer device, characterized in that, The computer device includes a memory, a processor, and an action instruction generation and optimization program stored in the memory and executable on the processor. When the action instruction generation and optimization program is executed by the processor, it implements the steps of the action instruction generation and optimization method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The storage medium stores an action instruction generation and optimization program, which, when executed by the processor, implements the steps of the action instruction generation and optimization method as described in any one of claims 1-7.
Citation Information
Cited By
Self-training architecture and end-side chip-based intelligent robot training method and system, terminal and storage medium
CN122033996A