Robot trajectory control method, device and equipment based on multiple modes and medium

By employing multimodal fusion and contact force signal feedback mechanisms, the problem of incoordination between trajectory generation and force control feedback in robotic ultrasonic scanning was solved, enabling stable adhesion scanning and high-quality imaging of the robot on complex soft tissue surfaces.

CN121552341APending Publication Date: 2026-02-24SHENZHEN BEAUTIFUL RUBIKS CUBE ROBOT CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511731949.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-24
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Existing technologies lack multimodal fusion and adaptive closed-loop mechanisms in robotic ultrasonic scanning, resulting in a lack of coordination between trajectory generation and force control feedback, making it difficult to achieve stable adhesion scanning on complex soft tissue surfaces.

Method used

By acquiring natural language commands, environmental visual information, and robot state information, feature fusion is performed to generate visual-language fusion features. The impedance parameters are adaptively adjusted using contact force signals, and the trajectory generation model is corrected based on the contact force signals, thereby achieving coordinated control of multimodal perception and dynamic force control.

Benefits of technology

This enhances the robot's autonomous perception and execution coordination in complex contact environments, ensuring the stability of the scanning process and the quality of imaging.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121552341A_ABST
    Figure CN121552341A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of robots and intelligent control, and discloses a robot trajectory control method, device and equipment based on multiple modalities and a medium, and the method comprises the steps: obtaining a natural language instruction, environment visual information and current state information of a robot, carrying out the fusion of visual and language features, and generating a fusion feature, generating a track sequence through a track generation model according to the fusion features and the robot state information, collecting a contact force signal of the tail end of the robot in contact with the target surface, adaptively adjusting impedance parameters by using the contact force signal, and generating a control instruction; and based on the contact force signal, feedback correction is carried out on the track generation model to drive the tail end of the robot to move, and multi-mode sensing and force control feedback cooperative control is achieved. According to the invention, track and force control cooperation is realized through multi-modal fusion and contact force feedback, so that the robot keeps stable fitting and adaptive motion in a complex contact scene, and the track generation precision and the imaging quality are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robotics and intelligent control technology, and in particular to a method, apparatus, device and medium for robot trajectory control based on multimodality. Background Technology

[0002] Currently, ultrasound imaging has become an important non-invasive examination method in clinical diagnosis. However, traditional ultrasound examinations still mainly rely on doctors holding the probe for scanning. Because doctors need to maintain a proper probe posture, control contact force, and assess image quality in real time during the procedure, this process is highly dependent on individual experience and hand stability, making it difficult to achieve consistency and repeatability. Especially in soft tissue examinations such as those of the breast, abdomen, and heart, differences in operation between doctors can significantly affect image quality, leading to subjective fluctuations in diagnostic results.

[0003] With the development of robotic arms and medical robotics technologies, researchers are attempting to use robots to automate ultrasound scanning, reducing the burden of manual operation and improving imaging stability. Early solutions often employed offline trajectory planning or teaching-based replication, typically by pre-acquiring point clouds of the patient's body surface and generating a scanning path, which the robotic arm then followed. However, these methods rely on fixed postures and geometric models; when the patient's position or surface morphology changes, the pre-set trajectory easily becomes invalid, lacking real-time adaptability and making long-term application in clinical settings difficult.

[0004] To overcome path errors caused by changes in body position, some studies have introduced visual perception and point cloud reconstruction techniques, enabling robotic arms to adjust their posture based on real-time visual information during scanning. However, existing methods are mostly based on low-level visual geometric features, which cannot recognize complex semantic commands or the doctor's operational intentions, nor can they flexibly adjust scanning strategies according to different examination sites. In addition, these methods often rely solely on visual information for correction, failing to fully integrate other modal information (such as voice commands and the robot's own state), resulting in insufficient system interactivity and adaptability.

[0005] On the other hand, when a robot performs a scanning task, the control of the contact force between the end effector and the patient's skin directly affects image quality. Traditional force control strategies often employ fixed impedance or constant force modes, which cannot adaptively adjust according to real-time changes in contact force. When the probe is subjected to excessive force, it can easily cause patient discomfort or image distortion; when the force is insufficient, it is difficult to form a clear image, affecting diagnostic accuracy. In addition, existing path generation models are often disconnected from contact force feedback, failing to form an effective closed-loop control mechanism. This prevents the robot from dynamically correcting its trajectory based on mechanical feedback, making it difficult to achieve stable contact scanning on complex soft tissue surfaces. Summary of the Invention

[0006] The main objective of this invention is to provide a multimodal robot trajectory control method, device, equipment, and storage medium, aiming to solve the technical problem that the existing technology lacks a unified multimodal fusion and adaptive closed-loop mechanism between trajectory generation and force control feedback, and cannot dynamically correct the scanning trajectory according to real-time contact force signals, resulting in a lack of stability and intelligent adaptability in the scanning process.

[0007] To achieve the above objectives, the present invention provides a multimodal robot trajectory control method, comprising: Acquire natural language commands, environmental visual information, and the robot's current state information; The natural language instructions and the environmental visual information are subjected to feature fusion processing to generate visual-language fusion features; Based on the visual language fusion features and the robot's current state information, a trajectory generation model is used to predict and generate a trajectory sequence. Acquire the contact force signal generated when the robot end effector comes into contact with the target surface; The impedance parameters are adaptively adjusted based on the contact force signal, and control commands are generated based on the impedance parameters and the trajectory sequence. The contact force signal is used to correct the trajectory generation process of the trajectory generation model, and the robot end effector is driven to move according to the control command.

[0008] Furthermore, to achieve the above objectives, the present invention provides a multimodal robot trajectory control device, comprising: The multimodal input acquisition module is used to acquire natural language commands, environmental visual information, and the robot's current state information; The visual language feature fusion module is used to perform feature fusion processing on the natural language instructions and the environmental visual information to generate visual language fusion features; The trajectory prediction and generation module is used to predict and generate a trajectory sequence based on the visual language fusion features and the robot's current state information using a trajectory generation model. The contact force signal acquisition module is used to acquire the contact force signal generated when the robot end effector comes into contact with the target surface; An impedance adaptive control module is used to adaptively adjust the impedance parameters according to the contact force signal, and generate control commands based on the impedance parameters and the trajectory sequence; The feedback correction and end effector module is used to use the contact force signal to perform feedback correction on the trajectory generation process of the trajectory generation model, and drive the robot end effector to move according to the control command.

[0009] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and a multimodal robot trajectory control program stored in the memory and executable on the processor, wherein the multimodal robot trajectory control program, when executed by the processor, implements the steps of the multimodal robot trajectory control method as described above.

[0010] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing a multimodal robot trajectory control program, wherein the multimodal robot trajectory control program, when executed by a processor, implements the steps of the multimodal robot trajectory control method described above.

[0011] Beneficial Effects: This invention relates to the field of robotics and intelligent control technology, and discloses a multimodal robot trajectory control method, device, equipment, and medium. The method includes: acquiring natural language commands, environmental visual information, and robot current state information; performing visual and language feature fusion to generate fused features; generating a trajectory sequence based on the fused features and robot state information through a trajectory generation model; acquiring contact force signals generated by the robot's end effector contacting a target surface; adaptively adjusting impedance parameters and generating control commands using the contact force signals; and then performing feedback correction on the trajectory generation model based on the contact force signals to drive the robot's end effector motion, thereby achieving coordinated control of multimodal perception, dynamic force control, and trajectory updating. This invention introduces a multimodal fusion mechanism into the trajectory generation and force control processes, enabling the robot to generate trajectories with greater semantic understanding by integrating natural language, visual, and state information; simultaneously, it utilizes contact force signals to form an adaptive feedback closed loop, achieving dynamic correction of the trajectory generation model, thus maintaining stable contact scanning in complex contact environments and improving the coordination and imaging stability of the robot's autonomous perception and execution. Attached Figure Description

[0012] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a schematic diagram of an application environment for a multimodal robot trajectory control method according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating an embodiment of the multimodal robot trajectory control method of the present invention; Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the multimodal robot trajectory control device of the present invention; Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention; Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0013] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0014] The multimodal robot trajectory control method provided in this invention can be applied to, for example... Figure 1 In this application environment, the client communicates with the server via a network. The server can obtain natural language commands, environmental visual information, and the robot's current state information from the client. It performs visual and language feature fusion to generate fused features, and generates a trajectory sequence based on the fused features and the robot's state information through a trajectory generation model. It obtains the contact force signal generated when the robot's end effector contacts the target surface, adaptively adjusts the impedance parameters using the contact force signal, generates control commands, and then performs feedback correction on the trajectory generation model based on the contact force signal to drive the robot's end effector movement, achieving coordinated control of multimodal perception, dynamic force control, and trajectory update. This invention introduces a multimodal fusion mechanism in the trajectory generation and force control process, enabling the robot to generate trajectories with greater semantic understanding by integrating natural language, visual, and state information. At the same time, it uses the contact force signal to form an adaptive feedback closed loop to achieve dynamic correction of the trajectory generation model, thereby maintaining stable contact scanning in complex contact environments and improving the coordination and imaging stability of the robot's autonomous perception and execution. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster composed of multiple servers. The invention will be described in detail below through specific embodiments.

[0015] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the multimodal robot trajectory control method provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0016] like Figure 2 As shown, the robot trajectory control method based on multimodal operation proposed in this invention includes the following steps: S10: Acquire natural language commands, environmental visual information, and the robot's current state information; In this embodiment, the acquisition process of natural language commands, environmental visual information, and robot current state information forms the foundation of the multimodal input layer. Natural language commands originate from the operator's voice input signal. This voice signal is acquired by a microphone array and then processed by front-end speech denoising and echo cancellation to generate a temporally clean signal. The signal is then converted into text format after joint inference by an acoustic model and a language model, for subsequent use by the semantic parsing module. The text-based commands are further processed through word segmentation, dependency parsing, and intent extraction to form standardized language expression vectors, providing semantic input for multimodal fusion.

[0017] Environmental visual information is acquired by a color depth camera, which simultaneously acquires color and depth images to achieve joint perception of spatial texture and geometric information. After depth information is measured using infrared structured light or Time-of-Flight (ToF) mechanisms, it undergoes spatial calibration and distortion correction to form aligned point cloud data. The point cloud data and color images are then pixel-registered in a unified coordinate system to generate an environmental visual matrix with spatial depth and semantic color channels. This matrix is ​​used to identify target surfaces, obstacles, and operable areas, providing realistic environmental constraints for subsequent trajectory generation models.

[0018] The robot's current state information is generated by a group of built-in state sensors. This group includes encoders, gyroscopes, accelerometers, and torque sensors, which collect joint angles, end effector poses, and joint torque parameters via a bus communication interface. After time-stamping synchronization and packet loss compensation, the data forms a state information vector. This vector reflects the robot's current spatial attitude and motion state, and serves as the input reference for trajectory prediction and control calculations.

[0019] The execution of timestamp alignment and integrity checks ensures consistency of multi-source information in terms of time dimension and data integrity. Each frame's language instructions, visual information, and status information are matched within the same time window. If missing frames or time drift are found, they are corrected through interpolation algorithms and time resampling. The final generated multi-source data is structurally encapsulated into a unified format to ensure the efficiency and accuracy of subsequent feature extraction and fusion modules.

[0020] Natural language commands can be acquired and converted using various implementation methods. The voice input module can be configured based on an end-to-end speech recognition model (such as a Transformer Encoder-Decoder structure), or it can adopt a traditional architecture that separates the acoustic model from the language model. If the ambient noise is high, an adaptive beamforming algorithm can be integrated to enhance the quality of the speech signal. For the language parsing part, an attention-based intent recognition network can be used, or the parsing effect can be optimized through a small number of supervised learning sessions in a few-sample environment.

[0021] Environmental visual information can be acquired using different types of sensing devices. In structured light camera solutions, depth information can be decoded using infrared light patterns; in Time-of-Flight (ToF) solutions, the depth matrix can be calculated using pulse time-of-flight. To adapt to different lighting conditions, automatic exposure and gamma correction modules can be added. Registration of color and depth images can be solved using an extrinsic calibration matrix, or a vision-inertial fusion algorithm can be used to automatically correct sensor drift.

[0022] Robot state information can be acquired and transmitted in real time via EtherCAT or CAN bus communication, transmitting sensor data. In applications requiring high joint encoder accuracy, high-resolution absolute encoders can be used; in high-speed motion scenarios, inertial measurement units (IMUs) can be combined for fusion estimation to improve the stability of attitude calculations. For the data synchronization module, a message queue-based caching mechanism can be designed to ensure the timing consistency of multimodal data under high-concurrency transmission.

[0023] This embodiment integrates natural language commands, environmental visual information, and robot state information to form multimodal input data, enabling synchronous collaboration of semantic understanding, spatial perception, and state feedback. This allows the robot to generate targeted control responses based on semantic intent understanding, environmental geometry recognition, and its own posture estimation, enhancing the naturalness and environmental adaptability of human-computer interaction and providing an accurate data foundation for subsequent trajectory prediction and force control feedback.

[0024] S20, perform feature fusion processing on the natural language instructions and the environmental visual information to generate visual language fusion features; In this embodiment, the feature fusion process of natural language instructions and environmental visual information is a key step in multimodal representation learning. Its goal is to map linguistic semantics and visual perception information into a unified feature space, enabling the robot to achieve a collaborative understanding of instruction intent and environmental constraints at the perception level. Natural language instructions undergo context modeling via a language encoder, which can employ a Transformer structure based on a self-attention mechanism. This encoder maps the text sequence into a high-dimensional semantic vector through word embedding layers. This semantic vector establishes word dependencies in multi-layer self-attention computation, giving the representation result contextual semantic relevance.

[0025] Environmental visual information is extracted using spatial features by a visual encoder. The visual encoder can employ convolutional neural networks or visual Transformer structures, using multi-scale convolutional kernels to extract features such as local texture, shape, and depth gradients. For RGB images, color and edge information can be extracted; for depth images, surface normals and spatial hierarchy information can be extracted. To improve semantic consistency, visual features are typically incorporating spatial location information through positional encoding, ensuring that semantic and geometric information are represented in a unified feature dimension.

[0026] The fusion of language embedding vectors and visual feature vectors is achieved through a cross-attention mechanism. Cross-attention determines the region of interest for each language unit in visual space by calculating the correlation weight matrix between language vectors and visual vectors. The calculated weights are normalized and then used to weight visual features, causing semantic information to modulate the distribution of visual features in a targeted manner. This process is equivalent to introducing language constraints onto visual features, thereby forming semantically guided visual focusing within the model.

[0027] The fused features undergo dimensionality normalization and feature projection through a feature alignment module, typically using a linear layer or a multi-head attention layer. The resulting visual-language fusion features semantically encode both language intent and spatial structure information, providing high-level semantic and environmental prior support for subsequent trajectory generation.

[0028] Different architectures can be used to implement feature encoding for both language and vision. The language encoder can be based on pre-trained models, such as variants of BERT, RoBERTa, or GPT architectures, and its parameters can be fine-tuned to adapt to the robot's command corpus. Lightweight semantic embedding networks can also be designed to reduce computational latency during embedding generation. The visual encoder can be based on ResNet or Swin Transformer architectures, maintaining high-resolution perception capabilities while enabling multi-scale feature extraction. During the alignment stage, a two-stream Transformer structure can be introduced, allowing the language and visual streams to interact repeatedly through cross-modal attention layers, achieving higher-dimensional semantic-spatial coupling.

[0029] To enhance robustness under varying lighting conditions and diverse linguistic expressions, a modality normalization module can be added to stabilize the numerical distribution of visual and linguistic features. If the target scene has a complex structure (such as a human body surface or non-rigid objects), 3D point cloud features can be introduced as supplementary input to the visual encoder to improve the accuracy of spatial geometric understanding. For semantically ambiguous instructions, contextual memory mechanisms or dialogue history can be combined to enhance the temporal consistency of linguistic representation.

[0030] This embodiment achieves joint perception of task semantics and environmental geometry by fusing features from language and visual information. This enables the trajectory generation process to rely not only on geometric information but also on language guidance. The fused features establish a bidirectional association between human-machine commands and environmental vision at the semantic level, improving the robot's understanding and behavioral consistency during task execution.

[0031] S30, based on the visual language fusion features and the robot's current state information, predict and generate a trajectory sequence using a trajectory generation model; In this embodiment, the joint input of visual-language fusion features and robot current state information provides both semantic and kinematic constraints for trajectory generation. The visual-language fusion features carry task semantics and environmental geometric information, used to determine the target area, relative spatial relationships, and action priorities. The robot's current state information reflects the attitude, velocity, and end-effector pose of the executor, providing dynamic boundary conditions and kinematic constraints for the trajectory generation model. The trajectory generation model establishes a mapping relationship between semantic tasks and physical states by concatenating these two into a conditional input vector, thereby predicting executable trajectory sequences in the continuous time domain.

[0032] The trajectory generation model can employ a deep temporal prediction architecture, with its input layer receiving joint encodings of fused features and state vectors. Internally, the model may contain multiple layers of temporal convolutional networks or Transformer encoding layers to capture the temporal dependencies between instruction semantics, visual geometric features, and state changes. The prediction layer generates a probability distribution of trajectory nodes, describing the set of possible spatial locations and poses of the endpoint in each time slice. To ensure causal consistency in the trajectory generation process, the model introduces a temporal masking mechanism, limiting the direction of prediction dependencies by masking future node information, thereby guaranteeing the dynamic executability and physical plausibility of the trajectory sequence.

[0033] After the probability distribution is generated, the trajectory node sequence is extracted through a sampling mechanism. Sampling strategies may include greedy sampling, temperature-regulated sampling, or distribution sampling based on the Monte Carlo method, to balance the stability and diversity of the trajectory. The generated trajectory sequence maintains continuity in the time dimension and conforms to environmental constraints and attitude boundary conditions in the spatial dimension.

[0034] Trajectory generation models can be implemented based on different network architectures. A conditional variational autoencoder architecture can be used, employing fused features as prior conditions to model trajectory distribution in the latent space, achieving semantic and state-driven trajectory prediction. Alternatively, a Transformer architecture can be used, jointly processing semantic, visual, and state features in a multi-head attention layer to improve the ability to model long-range dependencies between trajectory nodes.

[0035] In a real-time execution environment, a sliding time window mechanism can be used for trajectory updates. This means that while the robot is executing the current trajectory, the model receives new visual-language fusion features and state information and recalculates subsequent trajectory nodes in real time, achieving dynamic trajectory correction. For medical applications requiring high precision, a constraint optimization module can be added to perform inverse kinematics solving after trajectory generation, ensuring that the generated trajectory points meet the robot's kinematic constraints and end-effector posture accuracy requirements.

[0036] The weights of input features can be adjusted for different scenarios. When the quality of visual input is low, the weight of state information can be increased to ensure motion stability. When the semantics of the task instructions are complex or contain spatial constraints, the attention weight of language features in the model input can be increased to enhance task orientation. If the object being manipulated is a non-rigid surface, a pose correction mechanism based on surface normals can be added during the trajectory prediction stage to make the trajectory fit the target surface better.

[0037] This embodiment achieves collaborative modeling of semantics, environment, and dynamic state by inputting visual language fusion features and robot current state information into the trajectory generation model. This enables the generated trajectory to not only conform to operational semantics and spatial environment constraints, but also to adapt to changes in current posture, thereby improving the accuracy of trajectory prediction and execution stability, and reducing repeated corrections and error accumulation.

[0038] S40, acquire the contact force signal generated when the robot end effector comes into contact with the target surface; In this embodiment, the contact force signal generated when the robot's end effector contacts the target surface is the key input for achieving force control closed-loop and contact state determination. The contact force signal is acquired by a multi-dimensional force sensor, typically a six-dimensional force sensor, installed between the robot's end effector and the probe, capable of simultaneously measuring force components in three directions and torque components in three directions. The raw signal acquired by the sensor is an analog or digital voltage output, which is converted into standardized data by the signal acquisition module. After acquisition, the data undergoes zero-point calibration. By recording the static output value in a non-contact state and using it as a reference offset, the initial drift caused by gravity and installation errors is eliminated.

[0039] The calibrated signal undergoes sensitivity correction based on the sensor calibration matrix to ensure a linear correspondence between the measured output of each channel and the actual force value. The corrected contact force data is further filtered to suppress high-frequency interference caused by mechanical vibration and electrical noise. The filter type can be a low-pass filter or an adaptive Kalman filter to ensure the smoothness and stability of force signal changes during contact.

[0040] The force signal is solved by a coordinate transformation algorithm into normal and tangential force components in the end-effector coordinate system. The normal force reflects the magnitude of the probe's indentation force along the surface normal direction, while the tangential force reflects the friction and slippage tendency, serving as an important basis for determining whether there is offset or slippage. All components are recorded in chronological order to form a time-series signal, which is used for subsequent impedance adjustment and trajectory correction calculations.

[0041] The force signal acquisition module can be implemented in different ways. For medical ultrasound robots, a miniature six-dimensional force sensor can be integrated at the probe connection flange to measure stress changes using strain gauge sensing technology. Alternatively, a fiber Bragg grating sensor can be used to improve the resolution of minute contact force changes. In high-frequency dynamic scanning scenarios, piezoelectric sensors can be used to achieve higher response speeds.

[0042] Signal acquisition can be accomplished using the A / D sampling card of the real-time control system. The sampling frequency is set within the range of 500Hz to 2kHz according to application requirements to balance real-time performance and signal stability. In a medical environment, to prevent noise interference, isolation amplification and anti-aliasing filtering modules can be added to the acquisition circuit. The force signal calculation process can employ matrix inversion operations or be updated online based on a dynamic calibration model to compensate for the impact of temperature changes or sensor aging on sensitivity.

[0043] If the system needs to identify the contact state, a threshold judgment mechanism can be set in the force signal processing module. When the normal force exceeds the set threshold, it is determined to be a valid contact, and the trajectory control module is triggered to enter the force control mode. For complex surfaces (such as curved or soft tissue surfaces), a surface normal estimation module can be added to correct the force decomposition direction in real time and improve the consistency between the contact signal and the surface normal.

[0044] This embodiment utilizes a high-precision force sensing and signal calibration mechanism, enabling the system to accurately reflect the contact state between the robot's end effector and the target surface in real time. This provides a reliable basis for subsequent impedance parameter adjustment and trajectory correction. The filtering and processing of the contact force signal effectively improves signal stability and availability, allowing the robot to perceive minute contact changes, avoid excessive pressure or failed contact, and enhance the safety and consistency of the scanning operation.

[0045] S50, adaptively adjust the impedance parameter according to the contact force signal, and generate control commands based on the impedance parameter and the trajectory sequence; In this embodiment, the adaptive adjustment of impedance parameters and the generation of control commands are the core processes for achieving safe and compliant robot operation. The contact force signal provides real-time interactive feedback between the end effector and the target surface. By analyzing the deviation between this signal and the desired contact force, the stiffness and damping parameters in the impedance model can be dynamically updated. The impedance model describes the relationship between external force and displacement, and its mathematical form can be expressed as: the output force equals the weighted combination of the stiffness and damping terms. The purpose of adaptive adjustment is to ensure that the robot maintains constant contact characteristics when the contact environment changes, thereby avoiding excessive pressure or detachment from the target surface.

[0046] In implementation, the system first calculates the error between the actual contact force and the target contact force. This error is then processed by an adaptive law input parameter update module to correct the parameter vector of the current impedance model. Parameter adjustment can be based on a radial basis function neural network (RBFNN) or a gradient descent-based self-learning algorithm, automatically adjusting stiffness and damping through error feedback to ensure the output force converges to the desired value over time. Stiffness parameters primarily affect the system's force response sensitivity, while damping parameters mainly control end-velocity decay and oscillation stability.

[0047] The trajectory sequence, generated from the previous stage, contains a series of end-effector positions and attitude nodes. The system maps the trajectory sequence to the robot joint space and converts it into joint angle target points using forward kinematics equations. The updated impedance parameters and joint path points are input into the dynamic control equations to calculate the corresponding joint torque or moment command. The controller can employ an impedance-based hybrid control structure, ensuring that the output torque simultaneously satisfies the desired trajectory constraints and contact compliance conditions. The resulting control commands are continuous in time and constrained by both force and pose in space, providing directly driveable signals for subsequent execution modules.

[0048] The system can achieve adaptive impedance control in several ways. One approach is Model Reference Adaptive Control (MRAC), which defines ideal contact dynamics through a reference model and updates parameters in real time to approximate the ideal response. Another approach is a neural network-based self-learning impedance adjustment mechanism, which allows the model to automatically learn a suitable parameter distribution under unknown surface conditions.

[0049] When operating on soft tissue or curved surfaces, the desired contact force can be set as a dynamic function, determined by the surface normal and historical force signals, enabling impedance adjustment to possess time-dependent and surface-matching capabilities. To prevent high-frequency oscillations, a variable-gain damping strategy can be introduced, increasing the damping coefficient as the force error increases and decreasing it as the error approaches zero, thus balancing response speed and stability.

[0050] The control command generation section can be implemented based on a discrete-time controller, using updated impedance parameters and trajectory nodes to calculate joint torque inputs. If the robot has a force / position hybrid control architecture, force control and position control weights can be allocated separately on different coordinate axes. For example, force control can be used in the normal direction to maintain a constant contact force, while position control can be used in the tangential direction to ensure trajectory following accuracy.

[0051] This embodiment utilizes adaptive feedback of contact force signals to allow the system to adjust impedance parameters in real time, ensuring the robot maintains stable and compliant motion under various contact environments. By combining trajectory sequence generation with a dynamic control command generation mechanism, coordinated control of force and pose is achieved, improving operational stability and execution accuracy, and significantly reducing trajectory deviations and scanning unevenness caused by external force disturbances.

[0052] S60, the trajectory generation process of the trajectory generation model is corrected by feedback using the contact force signal, and the robot end effector is driven to move according to the control command.

[0053] In this embodiment, the trajectory generation model predicts motion trajectories by fusing visual, linguistic, and state information, while the feedback correction mechanism enables the model to perceive changes in external forces and adjust itself during execution. The system first inputs real-time contact force signals into the trajectory generation model as additional input conditions, allowing the model to consider force feedback information during the trajectory prediction stage, thereby enabling it to perceive changes in contact state.

[0054] The contact force signal reflects the interaction force between the robot's end effector and the target surface. By analyzing its normal and tangential components, the probe's contact state, friction trend, and surface curvature changes can be determined. The trajectory generation model maintains dynamic state variables internally. Feedback signals trigger internal state updates, causing the model to adjust its implicit representation between time steps, thereby correcting the trajectory node distribution. This update process can be implemented based on a gated recurrent unit (GRU) or variational Bayesian inference mechanism, ensuring that the corrected trajectory is physically continuous and conforms to real-time force feedback constraints.

[0055] The updated model generates a corrected trajectory sequence that better conforms to the target surface spatially and more dynamically to the desired contact force variation trend. After trajectory correction, the system sends the control commands generated in the previous step to the joint actuators. Upon receiving the control commands, the joint actuators perform inverse kinematics solving, generating drive signals to control the joint motors, achieving precise end-effector movement in three-dimensional space. During execution, changes in end-effector pose will again cause changes in contact force; the new force signals are then collected and input into the trajectory generation model, forming a closed-loop feedback control that ensures continuous dynamic consistency between trajectory generation and motion execution.

[0056] The feedback correction module can employ various computational mechanisms. Incremental trajectory updates can be used, correcting only local trajectory nodes to reduce computational overhead; alternatively, a global re-prediction approach can be used, regenerating subsequent trajectory nodes after a new contact force signal triggers. Different feedback frequencies can be selected for different application scenarios. If the task involves soft tissue contact (such as medical ultrasound scanning), the feedback frequency can be set to a high frequency (above 100Hz) to ensure real-time response; if the task is mechanical assembly or surface inspection, a lower frequency can be used to balance the computational load.

[0057] The internal state update of the trajectory generation model can be achieved in two ways. One approach is based on latent space correction, which embeds the contact force signal into the latent vector representation of the model and adjusts the trajectory prediction weights through backpropagation. The other approach is based on external control, which calculates the correction vector of the force signal through an independent feedback control module and then superimposes it with the trajectory node position to achieve direct spatial trajectory adjustment.

[0058] The drive control section can employ impedance control, force-position hybrid control, or model predictive control. If the robot's end effector has a force control mode, control commands can be directly mapped to target torque commands to maintain a constant contact force. If it uses a position control architecture, the corrected trajectory is mapped to the end-effector pose target point, and the control system dynamically adjusts the motor input based on the error to achieve smooth motion. In robotic arms with redundant degrees of freedom, zero-space optimization algorithms can be introduced to reduce posture deviations and structural singularities while fulfilling the main task.

[0059] This embodiment introduces contact force signals into the trajectory generation model to form a real-time feedback closed loop, tightly coupling trajectory generation with actual force feedback. The system can sense environmental changes during operation and adaptively adjust the trajectory, improving the stability and compliance of trajectory execution. Combined with the closed-loop execution of control commands, the robot's end effector maintains a suitable contact state throughout its movement, thereby avoiding trajectory drift, posture deviation, or surface detachment, and enhancing motion robustness and execution accuracy in complex environments.

[0060] This invention relates to the field of robotics and intelligent control technology, and discloses a multimodal robot trajectory control method, device, equipment, and medium. The method includes: acquiring natural language commands, environmental visual information, and robot current state information; performing visual and language feature fusion to generate fused features; generating a trajectory sequence based on the fused features and robot state information through a trajectory generation model; acquiring contact force signals generated by the robot's end effector contacting a target surface; adaptively adjusting impedance parameters and generating control commands using the contact force signals; and then performing feedback correction on the trajectory generation model based on the contact force signals to drive the robot's end effector motion. This achieves coordinated control of multimodal perception, dynamic force control, and trajectory updating. This invention introduces a multimodal fusion mechanism into the trajectory generation and force control processes, enabling the robot to generate trajectories with greater semantic understanding by integrating natural language, visual, and state information. Simultaneously, it utilizes contact force signals to form an adaptive feedback closed loop, achieving dynamic correction of the trajectory generation model, thereby maintaining stable contact scanning in complex contact environments and improving the coordination and imaging stability of the robot's autonomous perception and execution.

[0061] In one embodiment, step S10 above includes: S101, uses a color depth camera to acquire color and depth images of the target surface to generate raw environmental visual information; S102, receive natural language instructions in speech form through a speech recognition interface, and perform noise reduction and conversion processing on the natural language instructions in speech form to generate natural language instructions in text format; S103, reads joint angles and end effector poses through robot state sensors to generate current robot state information; S104, perform timestamp alignment and integrity verification on the original environmental visual information, text-formatted natural language instructions and robot current state information to generate verified multi-source data.

[0062] In this embodiment, the acquisition process revolves around three types of inputs and converges to "verified multi-source data" using a unified time base and consistency verification. The first type of input is environmental visual information. A color depth camera simultaneously outputs color and depth images. The color image undergoes distortion correction, white balance, and exposure compensation to stabilize brightness and color. The depth image suppresses random depth noise through missing pixel filling and spatial neighborhood smoothing. Depth and color are geometrically aligned within the camera coordinate system. The intrinsic and extrinsic parameters are determined using a calibration board or structured light identification, generating original environmental visual information with a one-to-one correspondence at the pixel level. To reduce cross-device and cross-scene differences, online brightness histogram constraints and depth confidence thresholds are introduced to eliminate low-confidence depth regions while retaining color edges for subsequent visual-semantic processing. The original environmental visual information is packaged with timestamps, frame numbers, and sensor identifiers to ensure consistent indexing with subsequent language and state data.

[0063] The second type of input is natural language commands. The speech recognition interface receives the speech stream and performs noise reduction and conversion processing. In the noise reduction stage, spectral subtraction and adaptive noise estimation are used in the frequency domain to suppress steady-state noise, and endpoint detection is combined in the time domain to distinguish effective speech from silence, preventing meaningless segments from entering the recognition and decoding process. Acoustic features are constructed using Mel-Cepstral or log-Mel energy to form feature frames, the language model maintains previous semantic cues using a context window, and the decoder outputs stable text-formatted natural language commands. To resist the influence of accents and homonyms, a confidence threshold is set, and local re-decoding is triggered under low confidence conditions, while retaining the original timestamp, segment start and end positions, and channel numbers, so that the text-formatted natural language commands can be aligned with other modalities on the time axis.

[0064] The third type of input is the robot's current state information. Robot state sensors read joint angles and end effector poses. Joint angles are obtained from feedback from joint encoders or actuators, and a sliding window debouncing process is used to suppress minor quantization jitter. The end effector pose is obtained by jointly solving forward and inverse kinematics, and then unified to the working coordinate system using a base-world coordinate system calibration matrix. The pose is represented by translation vectors and direction quaternions or rotation matrices, and a high-precision timestamp corresponding to each joint angle sample is recorded. To improve dynamic consistency, finite difference estimation of velocity and acceleration is added to determine the interpolation direction and validity during the alignment phase.

[0065] Multi-source timing fusion is carried out using a unified clock reference. The system maintains a global clock and periodically compares it with the hardware clock to calculate the relative drift and jitter upper bounds. During timestamp alignment, the visual frame time is used as the reference to find the nearest neighbor text-formatted natural language instruction fragment and the robot's current state information within the corresponding time window. If the cross-modal time difference exceeds the threshold, linear or spline interpolation is used to resample the joint angles and end effector poses in time. At the same time, the text is trimmed by the speech segment boundaries to avoid cross-sentence alignment. Integrity verification includes three aspects: field integrity, numerical range, and physical feasibility. Field integrity requires that the original environmental visual information has paired data of color and depth images, the text-formatted natural language instructions have timestamps and segment indexes, and the robot's current state information has joint angles and end effector poses. Numerical range checks whether the pixel values, depth distances, joint angles, and poses are within the equipment specification range. Physical feasibility uses kinematic coherence and maximum joint velocity constraints to determine whether there are abnormal jumps. Through the above alignment and verification, the three types of inputs are unified under the same time axis and coordinate system, and finally the verified multi-source data is generated. The data items include three sets of interrelated records: original environmental visual information, natural language instructions in text format, and robot current state information. Each record retains the original timestamp, source identifier, and quality label to support the deterministic invocation of subsequent fusion, prediction, and control links.

[0066] This embodiment achieves spatial consistency between the original environmental visual information and the actual scene through geometric alignment and confidence filtering of color and depth images. Through noise reduction, endpoint detection, and confidence constraints, text-formatted natural language instructions are semantically and temporally stable and reliable. By synchronously acquiring joint angles and end effector poses and verifying kinematic consistency, the robot's current state information is continuously solvable at the dynamic level. Through unified clock and timestamp alignment and integrity verification, the three types of data form a consistent mapping within the same time axis. Therefore, the verified multi-source data simultaneously meets the requirements for subsequent processing in terms of temporal consistency, spatial consistency, and semantic consistency, reducing link uncertainties caused by cross-modal mismatches and data gaps, and improving fusion quality, trajectory prediction stability, and control execution repeatability.

[0067] In one embodiment, step S20 above includes: S201, Use a language encoder to perform context encoding on the natural language instruction to generate a language embedding vector; S202, Use a visual encoder to extract multi-scale spatial features from the environmental visual information to generate a visual feature vector; S203, calculate the association weight between the language embedding vector and the visual feature vector through a cross-attention mechanism, and align the language embedding vector and the visual feature vector; S204, Based on the language embedding vector and visual feature vector after the association weight fusion alignment, generate visual language fusion features.

[0068] In this embodiment, the feature fusion chain takes time-aligned data as input and outputs stable and reusable visual-language fusion features. Natural language instructions enter the language encoder for context encoding processing. Text preprocessing preserves punctuation and quantifiers, uses sub-word segmentation to handle out-of-vocabulary words, and maps segment timestamps and intra-sentence positional encodings to the embedding space. Word embeddings, positional embeddings, and segment embeddings are linearly superimposed and normalized, and then stacked in multiple layers of self-attention to form language embedding vectors. To maintain the semantic fine-grainedness of instructions in medical scenarios, semantic boundaries are constrained collaboratively by stop word masks and named entity tags. Self-attention retains dependencies related to actions, positions, and anatomical regions within a defined window, enabling the language embedding vectors to simultaneously carry action semantics and target site semantics in the channel dimension.

[0069] Environmental visual information is fed into the visual encoder in parallel to extract multi-scale spatial features. Color and depth images are first aligned in the camera coordinate system. The color domain undergoes brightness normalization and color correction to stabilize illumination, while the depth domain is filtered by confidence and filled with holes to ensure geometric continuity. Feature extraction employs a hierarchical structure: bottom-level convolution captures edges and textures, middle-level sparse attention aggregates local structures, and high-level cross-window attention converges large-scale associations, resulting in a set of visual feature vectors with complementary resolution and semantic levels. To improve the representation of surface curvature and soft tissue deformation, depth channels and normal vectors are encoded and incorporated into mid-to-high-level features via channel concatenation. Multi-scale features are recalibrated via channels to suppress redundant responses, and scale identifiers are embedded to preserve the scale source, avoiding scale ambiguity in subsequent alignment stages.

[0070] The cross-attention mechanism calculates association weights and performs alignment between the language embedding vector and the visual feature vector. The language side acts as the query, and the visual side as the key and value. First, the embeddings on both sides are dimensionally projected and normalized. Then, a gating mask is used to shield low-confidence depth regions and visual background regions unrelated to the instructions, reducing invalid matches. To handle the correspondence between one-word-multiple-address and one-address-multiple-language, the cross-attention mechanism adopts a multi-head structure. Different attention heads focus on action-related regions, anatomical boundary regions, and image quality-related regions, respectively. Each attention head independently generates association weights and performs temperature self-adjustment between heads to avoid excessive spikes in single-head weights that could lead to information loss. The alignment operation is completed in two stages: the first stage limits the range with a coarse-grained window to ensure that the match falls within the correct anatomical neighborhood; the second stage refines the correspondence within the neighborhood with fine-grained attention within the window to reduce the bias caused by pose changes. The association weights, after normalization, are used to perform weighted convergence on the visual feature vectors. At the same time, the weighting coefficients are written back as saliency markers for the language embedding, so that the language channel retains its salient region references in the visual space.

[0071] The two types of embeddings, aligned by association weights, are fused to generate visual-language fusion features. The fusion employs a dual-branch parallel structure: the visual-to-language branch aggregates the most relevant visual context for each linguistic semantic unit, while the language-to-visual branch introduces the most relevant linguistic control embedding for each visual location. The two branches are concatenated on a shared channel and then subjected to layer normalization and residual transformation to form a unified representation. To ensure the callability of the downstream trajectory generation model, the fusion output is expressed in a fixed-dimensional continuous vector space, accompanied by three types of meta-information labels: a time index for cross-time-step matching, a scale identifier for multi-scale backtracking, and a saliency mask for learnable filtering of downstream attention. The final visual-language fusion features achieve collaborative representation of action semantics, target regions, and image quality cues in the channel dimension, while preserving the local morphology and depth continuity related to the command in the spatial dimension.

[0072] This embodiment employs contextual encoding of natural language commands to obtain action semantics and target region constraints through language embedding vectors. Multi-scale spatial feature extraction from environmental visual information provides texture, shape, and depth continuity through visual feature vectors. A cross-attention mechanism is used to calculate association weights and achieve alignment, establishing a learnable mapping from position to semantics between language and vision, reducing interference from irrelevant regions. Weighted fusion based on association weights outputs visual-language fusion features, integrating semantic constraints and spatial geometry into a unified representation. Consequently, subsequent trajectory prediction can simultaneously access semantic intent and accessible geometric context within the same representation, reducing planning instability caused by path ambiguity and region mismatch, and improving command executability and real-time adjustability of the scanning path.

[0073] In one embodiment, step 30 above includes: S301, the visual language fusion features and the robot's current state information are concatenated into a conditional input vector; S302, Input the conditional input vector into the trajectory generation model to generate a trajectory node probability distribution; S303, apply a temporal masking mechanism to perform causal consistency processing on the probability distribution of the trajectory nodes to generate a consistent probability distribution of trajectory nodes; S304, Sample from the probability distribution of the consistent trajectory nodes to generate a trajectory sequence.

[0074] In this embodiment, the visual-language fusion features and the robot's current state information are first aligned on the same time reference. The time index is represented by a monotonically increasing sampling sequence, and missing moments are filled with the hold-out values ​​of the previous valid moment and marked as valid bits. The visual-language fusion features include two types of sub-vectors: semantic channels and spatial channels. The semantic channels carry action semantics and target region constraints, while the spatial channels carry reachable surface geometry and image quality cues. The robot's current state information includes joint angles, joint angular velocities, end-effector pose, contact allowable direction markers, current safety limits, and soft and hard tissue contact priority markers. Both types of inputs undergo layer normalization and linear projection to unify the channel dimensions to a preset dimension, while retaining temporal position embedding and source embedding to distinguish between semantic sources and sensor sources. The conditional input vectors are concatenated temporally in the sequence dimension and semantically in the channel dimension, ultimately forming a fixed-dimensional temporal tensor. To avoid any source dominating a channel, channel attention gating uses learnable weights to suppress redundant channels and boost task-related channels, ensuring that the conditional input vectors remain numerically stable and distinctive.

[0075] The trajectory generation model receives a conditional input vector and outputs a probability distribution of trajectory nodes. Trajectory nodes are represented in the discrete-time domain, with each node consisting of the end-effector pose increment and the desired propulsion direction. The end-effector pose is jointly represented by a 3D translation vector and a 3D rotation vector, while the propulsion direction is jointly represented by a unit vector and a propulsion amplitude, with the amplitude boundary constrained by the current safety limit. Internally, the model uses a self-attention structure to aggregate historical node context and conditional cross-attention to read semantic and spatial channel information from the conditional input vector. The output header is a parameterized probability header, providing the distribution parameters of translation, rotation, and propulsion amplitudes, as well as a discrete selection distribution related to the allowed contact direction. All distributions are numerically stabilized using temperature and variance lower bounds to prevent excessive spikes or degradation into a uniform distribution. To suppress out-of-bounds candidates, boundary projection clips the high-confidence intervals of the distribution to the feasible region, while simultaneously feeding the clipping factor back to the upstream normalization layer to maintain end-to-end numerical consistency.

[0076] The temporal masking mechanism implements causal consistency processing on the probability distribution of trajectory nodes. Based on a strict causal structure, the masking shields the influence of future moments on the current moment. Building upon this, a view-motion coupling mask is introduced, mapping image frame time to action frame time to prevent leakage across batches or patient samples. In soft tissue contact scenarios with slow deformation, attenuated references to the distribution of previous moments are allowed within a finite hysteresis window to retain memory of low-frequency deformations, while attenuation kernels limit the dominant effect of long-term information. Causal consistency processing is performed in the probability space. First, intra-head consistency is achieved for the multi-head distribution at each moment, then Mahalanobis distance constraints are applied for smoothing in the temporal dimension, ensuring that changes in distribution parameters between adjacent moments fall within the restricted range. Abnormal spikes are truncated to the stable range of the previous moment, and the mask simultaneously records the truncated moments for unified processing during training and inference phases, forming a consistent probability distribution of trajectory nodes.

[0077] The sampling process generates a trajectory sequence based on the probability distribution of consistent trajectory nodes. At each time step, the discrete selection distribution is first sampled to determine the propulsion direction category, and then temperature-adaptive sampling is performed on the continuous distribution. The temperature is jointly adjusted by image quality cues and safety limits; the lower the image quality, the lower the temperature to converge to a conservative solution. The sampling results undergo feasibility testing, including joint limits, collision distance thresholds, and consistency checks of the allowed contact direction. Samples that fail the test are replaced with the conditional mode of the distribution, and a replacement mark is recorded. The samples are reconstructed in the reference coordinate system using pose accumulation to obtain a seamless temporal trajectory sequence. At the same time, the confidence level and replacement mark of each node are retained for use by downstream impedance and force control loops. To reduce cumulative drift, zero-drift correction is performed at the end of the sequence, amortizing the cumulative error of overall translation and rotation across the entire sequence using least squares, ensuring that the closed loop can absorb the error during execution.

[0078] This embodiment introduces semantics and geometry synchronously with a unified dimension and time reference through conditional input vectors. Both types of sources are numerically stabilized and gating to suppress redundancy before and after splicing, reducing distribution oscillations caused by multi-source mismatch. The trajectory generation model expresses node uncertainty in parameterized probability form, and avoids infeasible candidates with boundary projection, outputting a stable cover of the executable solution set. The temporal masking mechanism performs causal consistency processing in the probability space, combined with smoothing and truncation, significantly reducing the disruption of planning continuity caused by future information leakage and short-period spikes. The sampling stage introduces temperature adaptation and feasibility verification of image quality and safety limits, further suppressing risky candidates and outputting trajectory sequences with confidence.

[0079] In one embodiment, step S40 above includes: S401 uses a six-dimensional force sensor to collect the original contact force signal generated when the robot's end effector comes into contact with the target surface in real time. S402, perform zero-point calibration and sensitivity calibration on the six-dimensional force sensor, and generate calibrated sensor parameters; S403, use the calibrated sensor parameters to correct the original contact force signal and generate a corrected contact force signal; S404, The corrected contact force signal is filtered to remove high-frequency noise components and generate a filtered contact force signal. S405, the filtered contact force signal is solved into normal contact force component and tangential contact force component; S406, record the time series data of the normal contact force component and the tangential contact force component to generate a complete contact force signal.

[0080] In this embodiment, a six-dimensional force sensor is installed between the end effector and the probe mount. The installation orientation is aligned with the robotic arm flange coordinate system during assembly using a tooling. Installation errors are recorded using an installation matrix for subsequent coordinate transformations. The raw contact force signal enters the acquisition buffer in a synchronous sampling format of timestamps, triaxial force, and triaxial torque. Sampling triggering uses the same hardware clock as visual frame and end-effector pose sampling. The buffer has a ring structure, supporting frame loss marking and saturation marking to prevent latency accumulation caused by data congestion. Contact triggering is jointly identified by a threshold and rate of change. Upon triggering, data from a short time window prior to contact is retained in the buffer to ensure a reference window for subsequent calibration and correction.

[0081] Zero-point calibration and sensitivity calibration are performed in two stages. Zero-point calibration is performed under no-load contact conditions, acquiring raw contact force signals within a short time window, calculating the static bias and writing it into the zero-point parameter table, while simultaneously recording temperature and humidity readings to form a lookup table relating the bias to environmental factors for online correction of environmental drift. Sensitivity calibration is performed under multi-pose, known loading conditions, with loading combinations covering the coupling of triaxial forces and triaxial moments, obtaining a set of mapping coefficients between the measured output and the actual load. These mapping coefficients are robustly fitted to obtain the calibrated sensor parameters, which are stored in matrix form with batch numbers and timestamps for easy traceability. To avoid long-term drift outside of shutdown calibration, online micro-calibration is performed intermittently within the non-contact window, rapidly updating the zero-point parameters through the zero-load window, while sensitivity parameters are only slightly adjusted when a persistent deviation trend is observed, and adjustment records are retained.

[0082] The corrected contact force signal is generated jointly from the original contact force signal and the calibrated sensor parameters. First, static bias is eliminated using zero-point parameters. Then, the scaling factor and cross-coupling of each channel are decoupled and scaled according to the sensitivity parameters. Subsequently, the sensor coordinate system data is transformed to the end effector coordinate system using an installation matrix. The end effector coordinate system is then associated with the current end effector pose and can be further transformed to the probe contact coordinate system when needed, ensuring consistency with the surface normal definition. If channel saturation or sample loss is detected, the correction process indicates invalid segments with a marker bit and calls short-time prediction compensation to generate placeholder values ​​to prevent filter state collapse and ensure that invalid segments do not enter the downstream force determination.

[0083] The filtered contact force signal employs a two-stage structure to balance disturbance suppression and contact edge preservation. The first stage is an anti-aliasing filter, with a cutoff band covering high-frequency noise within the sampling bandwidth. The filter coefficients are stored in pre-calculated form and switched according to the sampling frequency. The second stage is an adaptive smoothing process that suppresses low-amplitude fluctuations caused by slow drift and electrical interference. The smoothing intensity is dynamically adjusted with the contact rate of change and end-effector velocity, and the smoothing intensity is reduced during contact edge and rapid attitude adjustment to avoid edge blunting. A confidence score is retained at each time point during the filtering process. The score is determined by the signal-to-noise ratio, the sampling loss ratio, and the saturation flag, and is included in subsequent calculations along with the data.

[0084] The calculation of the normal and tangential contact force components is performed based on the contact coordinate system. The contact coordinate system is defined with the surface normal as its axis. The normal acquisition path supports multiple sources: 1) surface normal maps reconstructed from environmental vision and interpolated near the contact point; 2) the probe axis established based on the end effector attitude and probe geometry; and 3) a weighted fusion of the two, with contact stability as the weighting criterion. The filtered contact force signal is decomposed in the contact coordinate system and projected onto the normal and tangential planes to obtain the normal and tangential contact force components. To suppress spurious tangential components caused by contact coordinate system jitter, a short-term attitude smoothing is performed on the contact coordinate system before projection. The smoothing window is jointly set by the end effector velocity and image stability. The decomposition results simultaneously generate quality labels, including coordinate system confidence, filter confidence, and projection residuals, for quality weighting in downstream control stages.

[0085] Time series data is recorded according to a unified log format. Each sampling point includes fields such as timestamp, normal contact force component, tangential contact force component, end-effector pose summary, quality label, missing sample and saturation markers, and data source marker. A double-buffering strategy is used for writing, decoupling the computation thread from the disk writing thread to avoid blocking real-time acquisition. A complete contact force signal is defined as a time series set covering the entire contact event process, including continuous records of the three stages of contact establishment, stable contact, and contact release, ensuring that the timestamp is monotonic and the fields are complete. When a cross-segment splicing or system reset is detected, a boundary marker is inserted at the splicing point in the log to remind downstream modules to perform intra-segment consistency checks. To improve alignment quality with other channels, a global synchronization count is appended to the time series during writing, allowing the trajectory and vision modules to perform cross-channel alignment and error attribution.

[0086] In this embodiment, the original contact force signal is stably acquired through a unified clock, a ring buffer, and a trigger retention window, providing a complete context for subsequent processing. Zero-point calibration and sensitivity calibration separate the processing of bias, proportional error, and coupling error, and combine environmental factors and online micro-calibration to suppress long-term drift. The corrected contact force signal maintains a consistent scale across different operating conditions. Two-stage filtering suppresses noise while preserving the contact edge, reducing the loss of deformation information during the contact establishment and dissolution stages. Based on the decomposition of the contact coordinate system, the force information is mapped into normal and tangential contact force components. Combined with quality label output, this facilitates downstream impedance adjustment and path replanning for quality-weighted decision-making. Structured time-series recording improves cross-channel alignment accuracy and data traceability through double buffering and synchronous counting.

[0087] In one embodiment, step S50 above includes: S501, Calculate the error value between the contact force signal and the desired contact force; S502, dynamically adjust the stiffness and damping parameters in the impedance model according to the error value; S503, convert the trajectory sequence into robot joint space path points; S504, combining the adjusted stiffness and damping parameters with the joint space path points, calculates the joint torque command; S505, generate control commands based on the joint torque commands.

[0088] In this embodiment, the error between the contact force signal and the desired contact force is calculated with reference to the contact coordinate system. The normal and tangential directions of the contact coordinate system are provided by the preceding force decomposition and surface normal estimation link, and the force unit and direction convention are consistent with the contact force signal. The desired contact force is given as a scalar or vector, and its source can be a fixed force setting, a force setting self-adjusted according to image quality indicators, or a force template mapping based on the anatomical region. The mapping result is expressed in the contact point coordinate system. The error value is calculated after time alignment, using the contact force signal sampled at the same frequency and the desired contact force to calculate the difference point by point, while adding a confidence weight to suppress the influence of low-quality readings; a short window consistency judgment is performed on the error value sequence to remove transient spikes and maintain the stability of the error input.

[0089] In the impedance model, stiffness and damping parameters are maintained separately via normal and tangential channels, and their form can be a diagonal matrix or a symmetric positive definite matrix with a small number of coupling terms. Parameter adaptive updates are driven by the error value, introducing a joint regulation law for error amplitude, error rate of change, and end velocity. Stiffness parameters are updated incrementally along the error direction; the update step size is reduced in small oscillation intervals to avoid overfitting and contact noise, while the step size is increased in persistent deviation intervals to accelerate convergence. Damping parameters are linked to the error rate of change and end velocity; the damping parameter is increased when the error rate of change or end velocity increases to enhance dynamic steady-state suppression, and appropriately reduced when the error rate of change decreases and the velocity is low to improve compliance. To constrain parameter numerical stability, bounded projection and minimum positive definiteness constraints are introduced to ensure that the updated stiffness and damping parameters remain positive definite and within a preset range. At the end of each control cycle, the version number and timestamp of parameter updates are recorded for subsequent backtracking and consistency checks.

[0090] The conversion from trajectory sequences to robot joint space pathpoints is performed using the current robot kinematics model, including calls to forward and inverse kinematics and singularity avoidance. The trajectory sequence provides the expected pose sequence and timestamps of the end effector in the workspace. The conversion process starts with timestamp alignment. When multiple solutions are obtained through inverse kinematics, the solution with the smallest distance from the current joint state is selected first. When approaching singular poses or joint limits, bias correction is performed through redundancy decomposition and joint weight adjustment. For unanalyzable pose segments, numerical iteration is used with a limited number of iterations. If convergence fails, transition pathpoints are generated by approximating the last feasible solution and the end effector's micro-displacement to maintain path continuity. The pathpoint set is expressed as a time-ordered queue of joint vectors, and the corresponding workspace pose summary is retained for verification.

[0091] The joint torque command is calculated using an impedance outer loop plus joint dynamics compensation. First, the impedance force is calculated in the workspace. The normal force is obtained by multiplying the stiffness parameter by the normal error and the damping parameter by the normal velocity in the normal direction. The tangential force vector is obtained in the same form in the tangential plane. These two are combined to form the desired end-effector force. Then, the desired end-effector force is mapped to the joint space equivalent torque through Jacobian matrix transpose. Gravity compensation, Coriolis and centrifugal term compensation, and friction compensation generated by the joint dynamics model are then superimposed. The parameters of the dynamics terms are taken from the calibration database and slightly corrected according to temperature or load conditions. The obtained joint torque command is subject to safety constraints and rate limits to prevent mechanical shocks caused by sudden changes, while maintaining an output beat consistent with the path point timestamp. If contact loss or sensor saturation is detected, the joint torque command is switched to low-gain hold to wait for effective contact recovery before being output through the normal channel.

[0092] Control commands are encapsulated from joint torque commands and corresponding timing tags, saturation flags, and quality tags, conforming to the message format of the fieldbus or controller interface. Before packaging, control commands undergo integrity and time synchronization checks to ensure consistency with the pace of the pathpoint queue. In multi-axis controller scenarios, control commands use synchronous trigger bits to ensure simultaneous updates across all axes. In remote delivery scenarios, control commands utilize dual-queue buffering and packet loss retransmission flags to mitigate link jitter. After command delivery, acknowledgments are received and execution times are recorded for subsequent error attribution analysis and parameter update strategy evaluation.

[0093] This embodiment calculates the error value in the contact coordinate system and couples it with the confidence weight, ensuring that the input deviation is consistent with the physical meaning of the surface normal and tangential, thus avoiding false corrections caused by coordinate jitter. The stiffness and damping parameters are jointly driven by the error amplitude, error rate of change, and end-effector velocity and are subject to positive definiteness and range constraints, forming a convergent, controllable, and non-oscillating adaptive impedance update. The trajectory sequence is transformed by inverse kinematics and combined with singularity avoidance and solution selection rules to generate continuously executable robot joint space path points, reducing path interruptions and posture flips. The joint torque command maintains compliance with external contact and the ability to compensate for the gravity and inertia of the mechanism under the superposition of impedance outer loop and dynamic compensation. The command-level rate limit and safety constraints reduce the risk of instantaneous impact. The control command is issued with time and quality tags and closed loops with the execution acknowledgment, improving the control of timing consistency and post-event traceability.

[0094] In one embodiment, step S60 above includes: S601, The contact force signal is input into the trajectory generation model as an additional input condition; S602, Update the internal state of the trajectory generation model according to the contact force signal; S603, using the updated trajectory generation model to generate the corrected trajectory sequence; S604, the control command is sent to the robot joint driver, and the robot end effector is driven to move through the robot joint driver.

[0095] In this embodiment, the additional input condition for the contact force signal to enter the trajectory generation model is based on maintaining consistency in time and coordinates with the previously generated visual language fusion features and the robot's current state information. The contact force signal includes normal and tangential contact force components, as well as contact quality labels and saturation markers synchronized with them. Time alignment uses a unified clock or timestamps aligned to the trajectory sequence. Coordinate consistency is achieved by mapping the normal and tangential components to the end-effector pose coordinate system used by the trajectory sequence. The surface normal required for mapping is provided by the preceding perception link, and any coordinate offset is compensated for with a small rotation correction matrix. To improve usability, the contact force signal is encapsulated as a fixed-length condition vector, which includes the current normal and tangential contact force components, short-time mean and short-time variance, contact stability score, and a binary indication of the saturation marker. This condition vector is concatenated with the visual language fusion features and the robot's current state information in the feature dimension to form a joint input that matches the input dimension of the trajectory generation model. To avoid abnormal contact from having an excessive impact on the generation, a force gating coefficient is introduced to scale the condition vector dimension by dimension. The gating coefficient is determined by the contact stability score and the saturation mark, and it approaches one when there is stable contact and approaches zero when there is noise or saturation.

[0096] The trajectory generation model's internal state update is driven by joint input and uses statistics of the contact force signal as correction parameters. The internal state comprises two parts: a short-term motion memory vector (SMMB) and a wandering offset vector. The former represents the trajectory trend over several recent time slices, while the latter represents the slowly drifting attitude trend. Internal state updates are executed along two paths: one for the SMMB, which performs gated reset based on the deviation between the contact force signal and the desired contact conditions. When the deviation increases, the reset ratio is increased to quickly discard mismatched motion trends; when the deviation is small, exponential sliding aggregation is used to maintain trajectory continuity. The other path for the wandering offset vector introduces low-frequency correction, mapping the median offset of the normal contact force component to a slow attitude fine-tuning amount, limiting the maximum change per cycle to prevent large attitude swaying. To suppress state oscillations caused by input jitter, consistency checks are performed before and after internal state updates. If the contact stability score is below a threshold for several consecutive sampling windows, the state update is frozen and resumed only after stable contact is restored. After the update, a state version number and timestamp are generated and written to the model's state cache to ensure that subsequent generation stages read a state consistent with the current input.

[0097] The corrected trajectory sequence is generated by the updated trajectory generation model on a time grid consistent with the original sampling strategy. First, the probability distribution of trajectory nodes is output at each time grid, and a previously defined causal mask constrains the scope of future information participation, ensuring that the generation process relies only on observed contact force signals and generated nodes. Then, sampling or modulo operation is performed on each grid according to the probability distribution to obtain trajectory nodes containing position and orientation. After sampling, two types of consistency checks are performed on the trajectory nodes: one is a kinematic reachability check, which uses the current robot joint boundaries and singularity criteria to eliminate unreachable nodes and replace them with the nearest reachable node; the other is a contact consistency check, which uses a threshold range of the normal contact force component to ensure that nodes are near the surface. If a node causes disengagement or excessive indentation, fine-tuning is performed along the surface normal direction while maintaining a smooth tangential transition. To maintain temporal continuity with the original trajectory sequence, the corrected trajectory sequence inherits the original timestamps, and spline smoothing is applied to the start and end nodes crossing the modified segment to avoid velocity discontinuities.

[0098] The control command issuance and robot end-effector motion are driven by previously calculated control commands as the data source. The transmission link includes three stages: packetization, synchronization, and acknowledgment. During packetization, the control command, target timestamp, quality tag, and rate limit flag are merged into a control message, conforming to the joint driver interface constraints and including a verification field. In multi-axis scenarios, synchronization trigger bits or centralized clock edges ensure simultaneous updates; in network link scenarios, double-buffered queues and retransmission counters are enabled to ensure consistent timing and reliable delivery. In the acknowledgment stage, the driver returns the execution time and execution status flags. The control layer records the difference between the acknowledgment and the target timestamp for subsequent statistics. When a late or failed acknowledgment is detected, a degradation strategy is triggered, switching subsequent messages to conservative gain and shortening the prediction window. Once the acknowledgment stabilizes, the original strategy is automatically returned. After the driver completes the update, the robot end-effector motion proceeds according to the message rhythm. The trajectory nodes and contact conditions remain consistent throughout the execution link. If contact loss is detected, advancement is paused, and minimal end-effector disturbance is maintained until contact is restored.

[0099] Example Description: In a complete operational instance of an automated ultrasound scanning task, the system first acquires color and depth images of the target area of ​​the patient using a color depth camera to generate environmental visual information, and receives voice commands from the doctor via a speech recognition interface. The voice data undergoes noise reduction and speech-to-text conversion to form structured natural language instructions, which, together with joint angle and end-effector pose data acquired by the robot's state sensors, constitute a multi-source input. All input data, after timestamp alignment and integrity verification, forms a synchronized multimodal input set, providing the foundational data for subsequent fusion processing.

[0100] In the multimodal input fusion stage, the language encoder models the context of natural language instructions, extracting semantic embedding vectors that describe the scanning region, operation, and termination condition. The visual encoder extracts multi-scale spatial features from the color and depth images, generating visual feature vectors that include surface texture, depth variations, and spatial topology. A cross-attention mechanism is used to calculate the correlation weights between language and visual features, locating the target region mentioned in the language instruction in the image space, thus aligning semantics and space. The fusion module merges the aligned features based on the correlation weights to form visual-language fusion features, which describe the high-level task intent of "where and what operation to perform".

[0101] The trajectory generation stage uses visual-language fusion features and the robot's current state information as joint inputs, concatenating them into a conditional input vector for the trajectory generation model. The model outputs a probability distribution of trajectory nodes, describing the temporal distribution trend of the end effector's possible path points in space. To prevent the model from using future information and causing prediction drift, a temporal masking mechanism is introduced to constrain predictions to rely only on past information, and causal consistency processing is performed on the probability distribution of trajectory nodes. After consistency constraints, the model samples the probability distribution in chronological order to generate a trajectory sequence, obtaining the target motion path of the robot's end effector during task execution, ensuring the consistency between the motion trajectory and the semantic target and scene geometry.

[0102] During trajectory execution, when the robot's end effector contacts the target surface, a six-dimensional force sensor installed at the end effector collects contact force data in real time. Before data acquisition, the sensor undergoes zero-point calibration and sensitivity calibration. A correction parameter matrix is ​​generated using multi-pose loading test data to perform linear and nonlinear corrections on the original contact force signal, removing bias and scaling distortion. The corrected contact force signal then undergoes a two-stage filtering process. The first stage filters out high-frequency noise, and the second stage adaptively adjusts the smoothness based on the contact state. The filtered signal is decomposed into normal and tangential components in the contact coordinate system, forming a contact force signal dataset containing a time series, providing continuous feedback for subsequent force control adjustment and trajectory feedback.

[0103] In the force control adjustment phase, the system first calculates the error between the contact force signal and the preset desired contact force, and then dynamically adjusts the stiffness and damping parameters of the impedance model based on the error amplitude and rate of change. The stiffness parameter determines the relative displacement response between the end effector and the surface, while the damping parameter is used to suppress rapid oscillations and hysteresis. The model incrementally updates the parameters while maintaining positive definiteness and convergence, enabling the robot to be both compliant and maintain stable contact with soft tissue. The trajectory sequence is converted into joint space path points, and a set of joint angles is generated through inverse kinematics solution. The joint torque command is calculated by combining the adjusted stiffness and damping parameters, and gravity, friction, and inertia compensation terms are superimposed to form a complete control command.

[0104] In the feedback correction phase, the contact force signal is input into the trajectory generation model as an additional input condition to correct the trajectory generation process in real time. The model's internal state is updated based on contact stability and force deviation. The short-term memory is reset when errors are significant to eliminate erroneous trends, while the long-term offset is corrected with low-frequency adjustment to reduce attitude drift. The updated model regenerates the trajectory sequence, correcting contact consistency while ensuring temporal continuity, making the trajectory more closely conform to the target surface. The generated control commands are encapsulated and sent to the robot joint actuators. The actuators execute torque commands sequentially to control the robot's end effector movement, and real-time monitoring of feedback is used to evaluate execution latency and state consistency. When contact loss or exceeding limits is detected, the system automatically reduces the output gain and pauses trajectory advancement until contact is restored, at which point normal control is reactivated.

[0105] Throughout the process, visual language fusion and trajectory generation form a semantically driven path planning layer, contact force signals provide a physical feedback channel, and impedance adjustment and feedback correction constitute a dynamic adaptive closed loop, enabling the robot to maintain stable contact and uniform scanning in complex soft tissue environments. This example demonstrates a multimodal collaborative mechanism of perception, decision-making, and execution: the perception layer integrates semantic and spatial information, the decision-making layer achieves trajectory prediction and parameter adjustment, and the execution layer relies on real-time feedback for trajectory re-optimization and posture correction, realizing seamless control from the semantic to the physical layer. Through this link, the system can maintain stable force control and highly consistent motion trajectories under different body surface morphologies, different scanning directions, and different operational requirements, providing a high-quality, repeatable imaging path for automated ultrasound scanning.

[0106] This embodiment normalizes the contact force signal into additional conditions aligned with visual language fusion features and the robot's current state information, and then enters the generation channel through gating scaling, reducing the spread of abnormal triggers on path generation. Internal state updates employ a dual-path approach: the fast path resets short-term memory to block erroneous trends when contact deviation increases, while the slow path uses low-frequency correction to reduce attitude drift. The combination of these two approaches reduces long-term accumulated errors and improves response to contact changes. The corrected trajectory sequence is ensured to be executable and continuous after triple checks of causal constraints, reachability, and contact consistency, reducing the probability of contact loss and over-intrusion. Control commands are transmitted to the joint actuators through encapsulation with timing and quality tags, synchronous triggering, and a closed-loop acknowledgment mechanism. Timing consistency and failure degradation mechanisms on the execution side improve the stability and traceability of the end effector motion. Thus, force feedback is not only used for end effector compliance adjustment but also directly participates in trajectory generation and resampling, forming a closed-loop correction link from perception, generation to execution. Under the same perception conditions, it is easier to obtain a scan path with stable contact and uniform coverage.

[0107] In one embodiment, a multimodal robot trajectory control device is provided, which corresponds one-to-one with the multimodal robot trajectory control method described in the above embodiments. (Refer to...) Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the multimodal robot trajectory control device of the present invention. The modules include a multimodal input acquisition module 10, a visual-language feature fusion module 20, a trajectory prediction and generation module 30, a contact force signal acquisition module 40, an impedance adaptive control module 50, and a feedback correction and end effector module 60. Detailed descriptions of each functional module are as follows: The multimodal input acquisition module 10 is used to acquire natural language commands, environmental visual information, and the robot's current state information; The visual language feature fusion module 20 is used to perform feature fusion processing on the natural language instructions and the environmental visual information to generate visual language fusion features; The trajectory prediction and generation module 30 is used to predict and generate a trajectory sequence based on the visual language fusion features and the robot's current state information using a trajectory generation model. The contact force signal acquisition module 40 is used to acquire the contact force signal generated when the robot end effector comes into contact with the target surface; The impedance adaptive control module 50 is used to adaptively adjust the impedance parameters according to the contact force signal, and generate control commands based on the impedance parameters and the trajectory sequence. The feedback correction and end effector module 60 is used to use the contact force signal to perform feedback correction on the trajectory generation process of the trajectory generation model, and drive the robot end effector to move according to the control command.

[0108] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a server-side multimodal robot trajectory control method.

[0109] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 5As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a client-side multimodal robot trajectory control method.

[0110] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: Acquire natural language commands, environmental visual information, and the robot's current state information; The natural language instructions and the environmental visual information are subjected to feature fusion processing to generate visual-language fusion features; Based on the visual language fusion features and the robot's current state information, a trajectory generation model is used to predict and generate a trajectory sequence. Acquire the contact force signal generated when the robot end effector comes into contact with the target surface; The impedance parameters are adaptively adjusted based on the contact force signal, and control commands are generated based on the impedance parameters and the trajectory sequence. The contact force signal is used to correct the trajectory generation process of the trajectory generation model, and the robot end effector is driven to move according to the control command.

[0111] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor: Acquire natural language commands, environmental visual information, and the robot's current state information; The natural language instructions and the environmental visual information are subjected to feature fusion processing to generate visual-language fusion features; Based on the visual language fusion features and the robot's current state information, a trajectory generation model is used to predict and generate a trajectory sequence. Acquire the contact force signal generated when the robot end effector comes into contact with the target surface; The impedance parameters are adaptively adjusted based on the contact force signal, and control commands are generated based on the impedance parameters and the trajectory sequence. The contact force signal is used to correct the trajectory generation process of the trajectory generation model, and the robot end effector is driven to move according to the control command.

Claims

1. A multimodal robot trajectory control method, characterized in that, Includes the following steps: Acquire natural language commands, environmental visual information, and the robot's current state information; The natural language instructions and the environmental visual information are subjected to feature fusion processing to generate visual-language fusion features; Based on the visual language fusion features and the robot's current state information, a trajectory generation model is used to predict and generate a trajectory sequence. Acquire the contact force signal generated when the robot end effector comes into contact with the target surface; The impedance parameters are adaptively adjusted based on the contact force signal, and control commands are generated based on the impedance parameters and the trajectory sequence. The contact force signal is used to correct the trajectory generation process of the trajectory generation model, and the robot end effector is driven to move according to the control command.

2. The robot trajectory control method based on multimodal operation as described in claim 1, characterized in that, Acquire natural language commands, environmental visual information, and the robot's current state information, including: Color and depth images of the target surface are acquired using a color depth camera to generate raw environmental visual information; The system receives natural language commands in speech form through a speech recognition interface, performs noise reduction and conversion processing on the speech commands, and generates natural language commands in text form. The robot's current state information is generated by reading joint angles and end effector poses using robot state sensors. The original environmental visual information, text-formatted natural language instructions, and robot current state information are timestamped and their integrity is verified to generate verified multi-source data.

3. The robot trajectory control method based on multimodal operation as described in claim 1, characterized in that, The natural language instructions and the environmental visual information are subjected to feature fusion processing to generate visual-language fusion features, including: The natural language instructions are context-encoded using a language encoder to generate language embedding vectors. A visual encoder is used to extract multi-scale spatial features from the environmental visual information to generate a visual feature vector. The association weights between the language embedding vector and the visual feature vector are calculated using a cross-attention mechanism, and the language embedding vector and the visual feature vector are aligned. Based on the language embedding vector and visual feature vector after the association weight fusion and alignment, visual language fusion features are generated.

4. The multimodal robot trajectory control method as described in claim 1, characterized in that, Based on the visual language fusion features and the robot's current state information, a trajectory generation model is used to predict and generate a trajectory sequence, including: The visual language fusion features and the robot's current state information are concatenated into a conditional input vector; The conditional input vector is input into the trajectory generation model to generate a probability distribution of trajectory nodes; A temporal masking mechanism is applied to perform causal consistency processing on the probability distribution of the trajectory nodes to generate a consistent probability distribution of trajectory nodes. A trajectory sequence is generated by sampling from the probability distribution of the consistent trajectory nodes.

5. The multimodal robot trajectory control method as described in claim 1, characterized in that, Acquiring the contact force signal generated when the robot end effector contacts the target surface includes: The original contact force signal generated when the robot's end effector comes into contact with the target surface is collected in real time using a six-dimensional force sensor. Zero-point calibration and sensitivity calibration are performed on the six-dimensional force sensor to generate calibrated sensor parameters; The original contact force signal is corrected using the calibrated sensor parameters to generate a corrected contact force signal; The corrected contact force signal is filtered to remove high-frequency noise components, generating a filtered contact force signal. The filtered contact force signal is solved into normal contact force components and tangential contact force components; Record the time series data of the normal contact force component and the tangential contact force component to generate a complete contact force signal.

6. The multimodal robot trajectory control method as described in claim 1, characterized in that, The impedance parameter is adaptively adjusted based on the contact force signal, and control commands are generated based on the impedance parameter and the trajectory sequence, including: Calculate the error value between the contact force signal and the desired contact force; The stiffness and damping parameters in the impedance model are dynamically adjusted based on the error value. Convert the trajectory sequence into robot joint space path points; By combining the adjusted stiffness and damping parameters with the joint space path points, the joint torque command is calculated; Control commands are generated based on the joint torque commands.

7. The robot trajectory control method based on multimodal operation as described in claim 1, characterized in that, The trajectory generation process of the trajectory generation model is corrected by feedback using the contact force signal, and the robot end effector is driven to move according to the control command, including: The contact force signal is input into the trajectory generation model as an additional input condition; The internal state of the trajectory generation model is updated based on the contact force signal; Use the updated trajectory generation model to generate the corrected trajectory sequence; The control command is sent to the robot joint actuator, which then drives the robot end effector to move.

8. A robot trajectory control device based on multimodal operation, characterized in that, The multimodal robot trajectory control device includes: The multimodal input acquisition module is used to acquire natural language commands, environmental visual information, and the robot's current state information; The visual language feature fusion module is used to perform feature fusion processing on the natural language instructions and the environmental visual information to generate visual language fusion features; The trajectory prediction and generation module is used to predict and generate a trajectory sequence based on the visual language fusion features and the robot's current state information using a trajectory generation model. The contact force signal acquisition module is used to acquire the contact force signal generated when the robot end effector comes into contact with the target surface; An impedance adaptive control module is used to adaptively adjust the impedance parameters according to the contact force signal, and generate control commands based on the impedance parameters and the trajectory sequence; The feedback correction and end effector module is used to use the contact force signal to perform feedback correction on the trajectory generation process of the trajectory generation model, and drive the robot end effector to move according to the control command.

9. A computer device, characterized in that, The computer device includes a memory, a processor, and a multimodal robot trajectory control program stored in the memory and executable on the processor. When executed by the processor, the multimodal robot trajectory control program implements the steps of the multimodal robot trajectory control method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The storage medium stores a multimodal robot trajectory control program, which, when executed by a processor, implements the steps of the multimodal robot trajectory control method as described in any one of claims 1-7.

Citation Information

Cited By

  • Self-adaptive grabbing and force control adjusting system of cooperative arm

    CN121821408A

  • Multi-modal fusion perception and control method and device for robot dexterous hand

    CN122185244A

  • A vision-tactile fusion control method and system for embodied intelligent robots for biochemical experimental tasks

    CN122353625A