A method, system, device, and medium for generating robot motion based on a large-scale multimodal fusion model of visual, tactile, and language perception.

CN122560059APending Publication Date: 2026-08-14JIANGSU YUNMU ZHIZAO TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611040534.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-14
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0009]为解决现有机器人动作生成中,面对密集型接触任务仅依靠视觉、语言信号难以预估连续、稳定、安全的动作、且简单的拼接多模态信号不足以提取上下文高层表征的问题,本申请提供一种基于视觉触觉语言多模态融合大模型的机器人动作生成方法、系统、设备和介质

Benefits of technology

[0025] This invention incorporates tactile signals into robot motion generation tasks, which can enhance the model's understanding of the task through tactile semantic supervision and correct deviations in robot actions when performing intensive contact tasks through tactile signal distribution prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122560059A_ABST
    Figure CN122560059A_ABST
Patent Text Reader

Abstract

This application belongs to the field of robot intelligent control technology, and discloses a method, system, device, and medium for generating robot actions based on a large-scale multimodal fusion model of vision, tactile feedback, and language. The method includes: acquiring observation data containing tactile signals, visual signals, language signals, and state signals; inputting the observation data into a large-scale multimodal fusion model of vision, tactile feedback, and language to obtain predicted robot action commands and predicted tactile distributions; calculating the closed-loop guidance strength based on the deviation statistics between the current tactile signal and the predicted current tactile distribution; correcting the predicted action commands based on the closed-loop guidance strength to obtain corrected action commands; and generating robot actions based on the corrected action commands. This method, by introducing tactile signals and utilizing the semantic constraints and closed-loop guidance of tactile signals, enables tactile signals to better assist in action generation and improves the accuracy of the actions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of robot intelligent control technology, and in particular to a method, system, device and medium for generating robot motion based on a large model of visual, tactile and language multimodal fusion. Background Technology

[0002] As robots are increasingly used in contact-intensive operations such as assembly, pressing, plugging, twisting, wiping, opening doors, and object placement, robot control systems not only need to understand visual images and verbal task commands, but also need to generate continuous, stable, and safe action sequences based on the contact state between the end effector and the environment. In these tasks, physical interaction tactile information such as contact force, friction, slippage, collision, impedance changes, and sudden force changes is often difficult to reliably obtain from a single frame image or low-dimensional pose state. Existing research on generative motion strategies, visual-tactile representation learning, and multi-sensor fusion has shown that continuous motion generation, tactile feedback, and cross-modal representation learning can all improve the robot's contact operation capabilities.

[0003] In existing technologies, common robot motion generation methods typically use images, verbal commands, and low-dimensional robot states as input, and generate a sequence of actions over a future period using neural networks. While these methods achieve good results in visual semantic understanding and continuous motion generation, they suffer from the following limitations in contact-based operation scenarios:

[0004] First, visual and verbal inputs cannot uniquely determine the actual contact force. For example, under the same visual appearance, the tactile signals received by the robot may be significantly different depending on the object's stiffness, coefficient of friction, contact point state, or sensor bias.

[0005] Second, simply concatenating the current tactile information directly into the state vector can increase the sensor dimension, but it also makes the network treat the tactile signal as a low-level numerical condition, making it difficult to form a high-level representation oriented towards contact semantics.

[0006] Third, directly using long-term tactile information as input to the visual-language network may lead to a strong dependence on tactile modalities during the training phase. When historical tactile information is unavailable, noisy, or its distribution changes during the inference phase, inconsistencies between training and inference can easily occur.

[0007] Fourth, existing motion generation networks typically only supervise motion errors and lack explicit modeling of "what kind of tactile response the motion will cause," making it difficult to perform closed-loop correction of the motion sampling process based on anomalies in real tactile signals during the inference stage.

[0008] Therefore, a new robot motion generation network is needed, which enables tactile signals to not only serve as low-dimensional inputs, but also to play the roles of contact semantic supervision during training, tactile distribution modeling at the action layer, and closed-loop guidance of real tactile information during inference, thereby improving the stability and safety of motion in contact operations. Summary of the Invention

[0009] To address the challenges in existing robot motion generation methods, where relying solely on visual and linguistic signals is insufficient to predict continuous, stable, and safe movements for intensive contact tasks, and where simply splicing multimodal signals is inadequate for extracting high-level contextual representations, this application provides a robot motion generation method, system, device, and medium based on a large-scale model that integrates visual, tactile, and linguistic multimodal signals.

[0010] This application provides a robot motion generation method based on a large-scale visual-tactile-language multimodal fusion model in a first aspect, comprising: acquiring observation data including tactile signals from long-historical time points, tactile signals from short-historical time points, tactile signals from the current time point, visual signals from the current time point, language signals from the current time point, and state signals from the current time point; inputting the observation data into a large-scale visual-tactile-language multimodal fusion model to obtain a predicted next-cycle motion command 1 and a predicted next-cycle tactile distribution; and calculating the closed-loop guidance intensity based on the deviation statistics between the current-time tactile signal and the predicted current-time tactile distribution. According to the closed-loop guiding strength The predicted next-cycle action command 1 is corrected to obtain the corrected next-cycle action command 2; the robot's action is generated based on the next-cycle action command 2.

[0011] Furthermore, the large-scale model based on visual-tactile-language multimodal fusion includes, during the training phase, a visual feature extraction network, a language feature extraction network, a tactile semantic teacher extraction network, a state feature extraction network, a tactile semantic alignment network, an action-level tactile distribution prediction network, a task understanding network, a tactile feature extraction network, and an action generation expert network; and during the inference phase, the large-scale model based on visual-tactile-language multimodal fusion includes, during the reasoning phase, a visual feature extraction network, a language feature extraction network, a state feature extraction network, an action-level tactile distribution prediction network, a task understanding network, a tactile feature extraction network, and an action generation expert network.

[0012] Further, the visual feature extraction network uses the current visual signal as input to obtain visual feature labels; the language feature extraction network uses the current language signal as input to obtain language feature labels; the tactile semantic teacher extraction network uses the long-historical tactile signals as input to obtain historical tactile semantic teacher features; the state feature extraction network uses the current state signal as input to obtain state feature labels; the tactile semantic alignment network is used to constrain the semantic consistency between the historical tactile semantic query features output by the task understanding network and the historical tactile semantic teacher features; the action-level tactile distribution pre- The prediction network takes the tactile latent features output by the action generation expert network as input to obtain the predicted tactile distribution for the next cycle; the task understanding network takes the language feature tags and visual feature tags as input to obtain the historical tactile semantic query features and tactile language visual semantic features; the tactile feature extraction network takes the tactile signals at short historical moments as input to obtain short historical tactile conditional feature tags; the action generation expert network takes the tactile language visual semantic features, the state feature tags, and the short historical tactile conditional feature tags as input to obtain the predicted next cycle action instruction 1 and the tactile latent features.

[0013] Furthermore, the predicted tactile distribution at the current moment is the predicted tactile distribution for the next cycle output by the large model based on visual-tactile language multimodal fusion at the previous moment; under initial conditions, the predicted tactile distribution at the current moment is the distribution of the tactile signal at the current moment, and the closed-loop guidance intensity... .

[0014] Furthermore, the training phase includes action generation loss, action-level haptic distribution prediction loss, and haptic semantic alignment loss, and the total loss can be expressed as: ,in, This indicates the loss generated by the action. This represents the prediction loss for action-level haptic distribution. This represents the haptic semantic alignment loss. and These are the weighting coefficients;

[0015] The action generation loss is used to constrain the consistency between the predicted next-cycle action instruction 1 and the target action instruction; the action-level tactile distribution prediction loss uses the next-cycle real tactile signal as the teacher value to constrain the output of the action-level tactile distribution prediction network, so that the tactile latent features explicitly carry the ability to determine what kind of tactile response a given action will result in; the tactile semantic alignment loss is used to constrain the task understanding network output of the tactile language visual semantic features with contact semantics.

[0016] In a second aspect, this application provides a robot motion generation system based on a large-scale multimodal fusion model of visual, tactile, and language perception, comprising:

[0017] The observation data input module is used to acquire observation data including tactile signals from long historical moments, tactile signals from short historical moments, visual signals from the current moment, speech signals from the current moment, and state signals from the current moment.

[0018] The action generation module is used to input the observation data into a large model based on visual-tactile language multimodal fusion to obtain the predicted next-cycle action command 1 and the predicted next-cycle tactile distribution.

[0019] The tactile distribution storage module is used to store and update the predicted tactile distribution for the next cycle, and outputs the predicted tactile distribution at the current moment in real time.

[0020] The motion correction module is used to calculate the closed-loop guidance intensity based on the deviation statistics between the current tactile signal and the predicted current tactile distribution. Based on the closed-loop guiding strength The predicted next cycle action command 1 is corrected to obtain the corrected next cycle action command 2;

[0021] The motion execution module is used to generate the robot's motion according to the next cycle motion instruction 2.

[0022] In a third aspect, this application provides an electronic device, the device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the aforementioned method for generating robot motion based on a large model of visual-tactile language multimodal fusion.

[0023] In a fourth aspect, this application provides a computer-readable storage medium storing at least one instruction or at least one program, which, when loaded and executed by a processor, implements the robot motion generation method based on a large model of visual-tactile language multimodal fusion.

[0024] The above technical solution has the following advantages compared to the existing technology:

[0025] This invention incorporates tactile signals into robot motion generation tasks, which can enhance the model's understanding of the task through tactile semantic supervision and correct deviations in robot actions when performing intensive contact tasks through tactile signal distribution prediction.

[0026] This invention functionally divides tactile signals into long-history tactile semantic supervision, short-history tactile local action conditions, future tactile distribution supervision, and current real tactile closed-loop guidance, avoiding the representational confusion caused by simply splicing tactile signals with other multimodal signals as input.

[0027] This invention enables task understanding networks to learn tactile semantic query tags, thereby avoiding reliance on long-historical tactile inputs during both the training and inference phases.

[0028] This invention uses an action-level tactile distribution prediction head to explicitly include the mapping relationship between action and tactile response in the latent features of action, providing a differentiable target for closed-loop correction during inference.

[0029] This invention uses real-time measured tactile signals directly as closed-loop guidance signals during the inference phase, instead of relying on virtual tactile targets predicted by the task understanding network. This provides a more direct physical basis and makes action correction more accurate.

[0030] This invention can be used for contact-intensive tasks, helping to reduce the risk of failure caused by motion jitter, abnormal contact, and contact instability. At the same time, this invention is applicable to embodied robots with multimodal perception and end effector, including at least one type of robotic arm among single-arm robotic arms, dual-arm collaborative robots, mobile manipulation robots, humanoid robots, and quadruped robots. Attached Figure Description

[0031] Figure 1 A flowchart illustrating a robot motion generation method based on a large-scale model of visual-tactile-language multimodal fusion, provided for an embodiment of this application;

[0032] Figure 2 A schematic diagram of the architecture of a robot motion generation system based on a large model of visual-tactile-language multimodal fusion is provided for embodiments of this application;

[0033] Figure 3 A schematic diagram of a network structure for the training stage of a large model based on visual-tactile-language multimodal fusion, provided in an embodiment of this application;

[0034] Figure 4 A schematic diagram of a network structure for the inference stage of a large model based on visual-tactile-language multimodal fusion is provided in this application embodiment;

[0035] Figure 5 A schematic diagram illustrating the training process of a large-scale model based on visual-tactile-language multimodal fusion, provided for an embodiment of this application;

[0036] Figure 6 A schematic diagram of the reasoning process of a large model based on visual-tactile-language multimodal fusion provided for embodiments of this application;

[0037] Figure 7 A comparison chart of training set losses for the large model and the base model provided in the embodiments of this application;

[0038] Figure 8 A comparison chart of validation set losses for the large model and the base model provided in the embodiments of this application;

[0039] Figure 9 A comparison chart of task success rates for the large model and the base model provided in this application embodiment;

[0040] Figure 10 This application provides a schematic diagram of robot action execution.

[0041] Explanation of reference numerals in the attached figures:

[0042] Observational data input module; 110, Action generation module; 120, Action correction module; 130, Action execution module; 140, Tactile distribution storage module; 200, Visual feature extraction network; 210, Language feature extraction network; 220, Tactile semantic teacher extraction network; 230, State feature extraction network; 240, Tactile semantic alignment network; 250, Action-level tactile distribution prediction network; 260, Task understanding network; 270, Tactile feature extraction network; 280, Action generation expert network; 300, Visual feature labeling; 310, Language feature labeling; 320, Historical tactile semantic teacher features; 330, State feature labeling; 340, Tactile semantic features; 350, Predicted next cycle. Tactile distribution; 360. Historical tactile semantic query features; 370. Tactile language visual semantic features; 380. Short historical tactile conditional feature labeling; 390. Tactile latent features; 400. Acquisition of training dataset; 410. Construction of training samples; 420. Forward propagation of training samples; 430. Backpropagation of loss function; 440. Update of network parameters; 700. Robot body; 710. Voice acquisition device or voice receiving device; 720. Visual acquisition device or visual receiving device; 730. State acquisition device or state receiving device; 740. Tactile acquisition device or tactile receiving device; 750. Processor; 760. Memory; 770. Motion execution mechanism; 780. End effector. Detailed Implementation

[0043] The embodiments of this application will now be described with reference to the accompanying drawings. It should be understood that the following embodiments are used to illustrate the technical solutions of this application, and not to limit the scope of protection of this application. Where there is no conflict, the technical features in the embodiments of this application can be combined with each other. For those skilled in the art, various substitutions, modifications, or equivalent adjustments can be made to the following embodiments without departing from the technical concept of this application.

[0044] In the description of this application, unless otherwise expressly defined, terms such as "comprising," "including," and "having" should be understood as open-ended expressions used to indicate the presence of the described features, steps, units, modules, components, or combinations thereof, but do not exclude the possibility of the presence or addition of other features, steps, units, modules, components, or combinations thereof. "At least one" means one or more; "multiple" means two or more. "And / or" is used to indicate any one, any multiple, or all combinations of the objects listed before and after it. For example, "A and / or B" can indicate only A, only B, or both A and B. For ordinal numbers such as "first," "second," "1," and "2" used in this application, they are only used to distinguish objects of the same or similar types and should not be construed as indicating an order of importance, chronological order, spatial order, or quantity limitation. The terms "connection," "coupling," or "linked" used in this application, unless otherwise expressly defined, can be a direct connection or an indirect connection implemented through intermediate components, communication links, signal lines, buses, networks, or other structures; for connections between data processing modules or functional units, they can also be understood as the transmission relationship of data, control commands, or parameters. The "unit," "module," and "logic" described in this application can be implemented by hardware, software, firmware, or a combination thereof. For example, they can be implemented by a processor executing a program in memory, or by a dedicated circuit, controller, computing module, or multiple functional components working together. The term "preset" in this application can refer to pre-setting, calibrating, training, storing, or configuring before executing a corresponding step, or it can refer to updating or selecting based on the type of object to be driven, facial expression style, voice state, or control requirements. Unless the execution order of steps is explicitly defined, the step numbers and flow directions in the accompanying drawings in the embodiments of this application are mainly for illustrative purposes and should not be construed as an absolute restriction on the execution order of steps; some steps can be executed in parallel, alternately, or with adjusted execution order without affecting the data dependencies and processing logic.

[0045] In this embodiment, the object to be driven can be at least one of the following: a single-arm robotic arm, a dual-arm collaborative robot, a mobile manipulation robot, a humanoid robot, a quadruped robot, or other embodied robot objects capable of generating actions based on control data. The actions of the object to be controlled can be achieved through a physical end effector or through an animation model.

[0046] In this embodiment of the application, a long-history tactile signal refers to a signal that includes the current moment and previous moments within the sliding observation window, with a total length of [missing information]. The tactile sensing data collected continuously across multiple time points includes one or more of the following: six-dimensional force and torque at the end of each historical moment, pressure distribution on the contact surface, and contact deformation feedback information. This data is used to characterize the continuous interactive changes in the robot's contact process with the environment. Short-history tactile signals refer to signals containing the current moment and previous moments within a sliding observation window, with a total length of [missing information]. The tactile sensor data collected at multiple consecutive moments only reflects the contact state at the few moments closest to the robot's current position. .

[0047] In this embodiment of the application, the visual signal at the current moment refers to one or more image data obtained by instantaneous sampling, such as the environmental image captured by the global scene camera, the local image of the workpiece captured by the wrist palm eye, the object outline frame from the end view, etc., which only reflects the robot's current instantaneous spatial observation state; the image data includes, but is not limited to, RGB image, depth image, point cloud image.

[0048] In the embodiments of this application, the current moment language signal refers to the instruction text or speech encoding sequence sampled synchronously with the visual and tactile signals, such as task operation instructions, human-computer interaction speech, natural language assembly description, etc., used to convey the current task objectives and constraints. The language signal includes, but is not limited to, text embedding vectors and speech Mel-spectrum features.

[0049] In this embodiment, the current state signal refers to the motion parameters output by the robot body at the sampling instant, such as joint angles, end effector pose, joint velocity, joint torque, chassis coordinates, etc. It only represents the robot's current instantaneous motion condition. The state signal includes, but is not limited to, joint coding data and six-dimensional pose information in Cartesian space.

[0050] Figure 1 This is a flowchart illustrating a robot motion generation method based on a large-scale multimodal fusion model of visual, tactile, and language embodied in this application. Figure 1 As shown, the method in this embodiment may include the following steps.

[0051] S101, acquire observation data including tactile signals from long historical moments, tactile signals from short historical moments, tactile signals from the current moment, visual signals from the current moment, speech signals from the current moment, and state signals from the current moment.

[0052] In this embodiment, tactile signals from long historical timeframes, the current tactile signal, and the short historical timeframe can all be obtained by a tactile sensor. The tactile signals continuously acquired by the tactile sensor are categorized into long historical timeframe tactile signals, current tactile signals, and short historical timeframe tactile signals based on the length of their period. The tactile signals may include one or more combinations of six-dimensional force signals, contact surface pressure distribution, and contact deformation feedback information. The tactile signals may undergo preprocessing such as resampling, noise reduction, endpoint detection, and loudness normalization to facilitate subsequent extraction of tactile features.

[0053] S102, input the observation data into the large model based on visual-tactile language multimodal fusion to obtain the predicted next cycle action command 1 and the predicted next cycle tactile distribution.

[0054] In this embodiment, the action command 1 includes at least one robot action command among the following: action velocity field integral, actuator pose, joint angle, joint acceleration, and joint velocity; the tactile distribution includes one or more statistical features that can fully reflect the distribution of tactile signals, such as mean, variance, standard deviation, peak value, extreme value, contact area, pressure gradient, and temporal change slope; the next cycle is the sequence length of the action predicted by the large model of visual-tactile language multimodal fusion based on the current cycle.

[0055] S103, calculate the closed-loop guidance intensity based on the deviation statistic between the current tactile signal and the predicted current tactile distribution. .

[0056] In this embodiment, the predicted tactile distribution at the current moment is derived from the tactile distribution output by the large-scale model based on visual-tactile language multimodal fusion from the previous cycle; under initial conditions, the closed-loop guidance intensity... =0.

[0057] In one implementation, the tactile signal at the current moment can be represented as real force. The predicted tactile distribution at the current moment can be represented as the predicted force distribution at the current moment. Based on current true strength Calculate the statistic of the deviation between the current force distribution and the predicted current force distribution. This statistic considers both the magnitude and variance of the force deviation. When the actual force differs significantly from the predicted force distribution mean, and the predicted force distribution variance is small, A large deviation indicates that the current contact state may be abnormal or deviate from expectations. However, when the predicted force distribution itself has high uncertainty, the same force deviation will not produce an excessively strong correction.

[0058] In one implementation, the closed-loop guidance strength is calculated based on the deviation statistic. It can be represented as: .in, For maximum guiding strength, The slope coefficient, This is the trigger threshold. In this implementation, 0.2 can be taken. You can use 1.0. A value of 6.0 can be used. Closed-loop guidance strength. It can maintain a weak correction when the tactile signal is normal, and enhance closed-loop guidance when the tactile signal deviates significantly from the expectation.

[0059] S104, based on closed-loop guidance strength The predicted next cycle action instruction 1 is corrected to obtain the corrected next cycle action instruction 2.

[0060] In one implementation, action instruction 1 is used It indicates that the intensity is guided by a closed loop. The revised action instruction 2 uses It can be represented as To guide the direction of the action command at the current moment, a correction term is used to incorporate real physical feedback during the action command generation process. That is, after the robot executes the generated action, it collects the tactile signal caused by this action through tactile sensors. The difference between this tactile signal and the tactile signal predicted by the model is used to close the loop and guide the generation of the action. Under initial conditions... .

[0061] S105, Generate the robot's action according to the next cycle action instruction 2.

[0062] In this embodiment, the next cycle action command is converted into robot control commands at the current moment, such as joint angle commands, end effector six-dimensional pose commands, target compliance torque commands, etc. The robot controller receives the control commands and drives the robot body to perform the corresponding actions.

[0063] Figure 2 This is a schematic diagram of the architecture of a robot motion generation system based on a large-scale multimodal fusion model of visual, tactile, and language embodied in this application. Figure 2 As shown, the system may include an observation data input module 100, an action generation module 110, an action correction module 120, an action execution module 130, and a tactile distribution storage module 140.

[0064] The observation data input module 100 is used to acquire observation data including tactile signals from long-history time, tactile signals from short-history time, tactile signals from the current time, visual signals from the current time, language signals from the current time, and state signals from the current time. The action generation module 110 is used to input the observation data acquired by the observation data input module 100 into a large-scale model based on visual-tactile-language multimodal fusion to obtain the predicted next-cycle action command 1 and the predicted next-cycle tactile distribution. The action correction module 120 is used to calculate the closed-loop guidance intensity based on the deviation statistics between the current-cycle tactile signal and the predicted current-cycle tactile distribution. And based on closed-loop guidance strength The predicted next-cycle action instruction 1 is corrected to obtain the corrected next-cycle action instruction 2; the action execution module 130 is used to generate the action that the robot will execute in the next cycle according to the next-cycle action instruction 2; the tactile distribution storage module 140 is used to store and update the predicted tactile distribution of the next cycle, and provide the action correction module 120 with the predicted tactile distribution at the current moment in real time.

[0065] Figure 3 This is a schematic diagram of a network structure for the training phase of a large model based on visual-tactile-language multimodal fusion, as provided in an embodiment of this application. Figure 3 As shown, the training stage of a large model based on visual-tactile-language multimodal fusion includes a visual feature extraction network 200, a language feature extraction network 210, a tactile semantic teacher extraction network 220, a state feature extraction network 230, a tactile semantic alignment network 240, an action-level tactile distribution prediction network 250, a task understanding network 260, a tactile feature extraction network 270, and an action generation expert network 280.

[0066] In one specific embodiment, the visual feature extraction network 200 can be represented as any one of a visual feature encoder, an image convolutional coding network, a multi-scale visual mapping network, or a palm-eye global image fusion encoder; the language feature extraction network 210 can be represented as a text embedding encoder, a speech Mel-spectrum coding network, or a natural language Transformer coding module; the tactile semantic teacher extraction network 220 can be represented as a six-dimensional force temporal coding network, a tactile array pressure distribution encoder, or a tactile feature extraction backbone network with a temporal window; the state feature extraction network 230 can be represented as a robot body pose coding network, a joint motion parameter mapping encoder, or a multi-degree-of-freedom motion state temporal coding module; the tactile semantic alignment network 240 and the action-level tactile distribution prediction network 250 can both be composed of a multi-layer MLP network, a one-dimensional temporal convolutional network, a Transformer coding layer, or a multi-head attention mapping network. The task understanding network 260 can be composed of at least one of the following: a visual language network, a multimodal task encoding network, a natural language instruction parsing network, a cross-modal task semantic understanding Transformer network, and a scene task association mapping network, with a learnable tactile semantic query label added on top of at least one of these networks; the tactile feature extraction network 270 can be composed of a stacked causal dilated one-dimensional convolutional module, a multi-scale dilated causal temporal encoding network, a causal dilated one-dimensional convolutional backbone network with residual connections, a hybrid network of one-dimensional causal convolution and multi-head temporal attention, or a temporal extraction unit with causal dilated convolution and gated activation; the action generation expert network 280 can be composed of a stream matching Transformer action decoder with tactile feature labels, a tactile perception action block generation network with stacked causal dilated one-dimensional convolution, or a bidirectional mask action integrating tactile cross-attention. The network consists of an Expert backbone network, a conditional flow matching temporal prediction module that integrates long-historical tactile window features, a causal extended Conv1D residual layer and a multi-head tactile cross-attention hybrid π0 action expert network, and a block action diffusion generation unit that embeds tactile distribution statistics.

[0067] In another specific embodiment, the visual feature extraction network 200 outputs visual feature labels 300; the language feature extraction network 210 outputs language feature labels 310; the tactile semantic teacher extraction network 220 outputs historical tactile semantic teacher features 320; the state feature extraction network 230 outputs state feature labels 330; the tactile semantic alignment network 240 outputs tactile semantic features 340; the action-level tactile distribution prediction network 250 outputs the predicted next-cycle tactile distribution 350; the task understanding network 260 outputs historical tactile semantic query features 360 and tactile language visual semantic features 370; the tactile feature extraction network 270 outputs short-historical tactile conditional feature labels 380; and the action generation expert network 280 outputs tactile latent features 390 and the next-cycle action instruction 1.

[0068] The visual input of the visual feature extraction network 200 can be processed by data augmentation, such as image visual input being processed by image augmentation preprocessing; the action input of the action generation expert network 280 can be a demonstration action containing noise, and the output can be an executable next-cycle action instruction 1 obtained by multi-step integration of the action velocity field action instruction, and the deviation between the next-cycle action instruction 1 and the demonstration action serves as the closed-loop feedback of the network.

[0069] Figure 4 This is a schematic diagram of a network structure for the inference stage of a large model based on visual-tactile-language multimodal fusion, as provided in an embodiment of this application. Figure 4 As shown, a large-scale model inference stage based on visual-tactile-language multimodal fusion includes a visual feature extraction network 200, a language feature extraction network 210, a state feature extraction network 230, an action-level tactile distribution prediction network 250, a task understanding network 260, a tactile feature extraction network 270, and an action generation expert network 280. The action input of the action generation expert network 280 can be one or more noises that satisfy Gaussian or Laplace distribution.

[0070] In one specific embodiment, during the multimodal fusion large model inference stage, the visual feature extraction network 200 receives the visual signal at the current moment and outputs a visual feature label 300; the language feature extraction network 210 receives the language signal at the current moment and outputs a language feature label 310; the tactile feature extraction network 270 receives the tactile signal at a short historical moment and outputs a short historical tactile conditional feature label 380; the state feature extraction network 230 receives the state signal at the current moment and outputs a state feature label 330; and the task understanding network 260 receives the visual feature label 300 and the language feature label 310, and outputs tactile-language-visual-semantic features 370. The action generation expert network 280 receives random noise distribution, tactile language visual semantic features 370, short history tactile condition feature labels 380 and state feature labels 330, and outputs tactile latent features 390 and the next cycle action instruction 1; the action-level tactile distribution prediction network 250 receives tactile latent features 390 and outputs the predicted next cycle tactile distribution 350; the predicted next cycle tactile distribution 350 is stored and updated for the correction of the next action instruction; the next cycle action instruction 1 is corrected according to the deviation between the tactile signal at the current moment and the next cycle tactile distribution 350 predicted in the previous cycle, to obtain the next cycle action instruction 2.

[0071] Figure 5 This is a schematic diagram illustrating the training process of a large-scale model based on visual-tactile-language multimodal fusion, as provided in an embodiment of this application. Figure 5As shown, the training process includes obtaining the training dataset 400, constructing training samples 410, forward propagation of training samples 420, backpropagation of the loss function 430, and updating network parameters 440.

[0072] In a specific embodiment, the training data includes various observation data such as vision, language, state, tactile sensation, and demonstrated actions. Specifically, the training data can consist of images, text, robot body state, force / torque, and motion velocity field. The construction of training samples includes cleaning, denoising, and processing of time-series cycles of the training data. The intermediate features output after the network forward propagation include visual feature labels 300, language feature labels 310, historical tactile semantic teacher features 320, state feature labels 330, historical tactile semantic query features 360, tactile language visual semantic features 370, short historical tactile conditional feature labels 380, and tactile latent features 390. The final output of the network forward propagation includes tactile semantic features 340, the predicted tactile distribution for the next cycle 350, and the action instruction for the next cycle 1.

[0073] In another specific embodiment, action generation loss This is used to train the network to progressively recover from noisy actions to the model action. During training, noisy actions are sampled from a standard noise distribution. and sampling time steps Let the sequence of demonstration actions be... Intermediate states can be constructed through linear interpolation. ,in Corresponding noise action, Corresponding to the actual demonstration action, the corresponding target action speed is: The network forward propagation outputs the predicted motion velocity field. The action generation loss is used to constrain the predicted action velocity field to be close to the target action velocity, and can be expressed as: Action-level haptic distribution prediction loss The output of the action-level haptic distribution prediction network is constrained, enabling latent haptic features to explicitly carry the ability to determine the haptic response based on a given action. Future haptic label sequences are employed. As a supervisory signal, the network outputs the corresponding mean value. and logarithmic standard deviation If the diagonal Gaussian negative log-likelihood is used as the tactile distribution prediction loss, then... It can be represented as: ; Haptic semantic alignment loss To make the tactile semantic query features of Task Understanding Network 260 approximate real-world tactile teacher features, 340 represents the semantic features of touch. The tactile semantic teacher feature 320 represents the tactile semantics of the same task scenario at the same moment, but from different sources: the former comes from visual and linguistic context, while the latter comes from real long-term historical tactile sequences. The tactile semantic alignment loss... The dot product similarity can be used to calculate it, and it can be expressed as: ,in, This represents the dot product similarity.

[0074] Specifically, using the ForceVLA dataset as a benchmark, the ForceVLA dataset was divided into a training set and a validation set in an 8:2 ratio, and the training set was trained using both this training method and the original π0 model training method. For example... Figure 6 The diagram shows a comparison of the loss function values ​​of the two training methods. The training set loss of both models gradually decreases during training. In the initial training phase, the introduction of an additional haptic modeling module and training objective increases the optimization difficulty of the model. The training loss of this model is higher than that of the original π0 model. However, as training progresses, the loss of this model gradually approaches and falls below that of the original π0 model, indicating that the addition of haptic information and the corresponding constraint of the loss function can gradually achieve effective optimization without significantly affecting the original motion generation capability.

[0075] In addition, such as Figure 7 The diagram shows a comparison of the loss of this model and the original π0 model on the validation set. As the training phase progresses, the trend of the loss on the validation set closely follows the change in the loss on the training set. After 6.5k iterations, the loss of this model on the validation set is significantly lower than that of the original π0 model on the validation set.

[0076] Figure 8 This is a schematic diagram illustrating the reasoning process of a large-scale model based on visual-tactile-language multimodal fusion, provided as an embodiment of this application. For example... Figure 8 As shown, the inference phase receives the observation data at the current moment. Under initial conditions, the observation data does not contain the tactile distribution prediction result for the next cycle. Under non-initial conditions, the observation data contains the tactile distribution prediction result for the next cycle, visual, linguistic, and state observations. The observation data is sent to the inference network to obtain the network output. The system determines whether it is an initial condition based on whether there is a tactile distribution prediction for the previous cycle. If it is an initial condition, closed-loop guidance is not enabled, and the network output is directly converted into robot actions. If it is not an initial condition, closed-loop guidance is calculated, and the action instructions are corrected according to the closed-loop guidance before being converted into robot actions. The system saves the tactile prediction result for the next cycle output of the network and updates the historical window data.

[0077] In the inference phase, the performance of the proposed method, the π0 model, OpenVLA, FACTR, and ForceVLA was compared on five typical contact tasks: pressing, inserting, wiping, flipping, and long-range contact. The probability of success for each task was calculated after 10 repeated attempts. Figure 9 As shown, the success rate of the method corresponding to this invention in all five tasks is higher than that of other VLA methods that do not contain tactile information or simply combine tactile information. The overall average success rate of the tasks is also the highest, reaching 93.33%, far exceeding the success rate of ForceVLA, which ranks second at 85.33%.

[0078] Figure 10 This application provides a schematic diagram of a robot motion generation and task execution structure based on a large-scale multimodal fusion model of visual, tactile, and language perception. (See diagram below.) Figure 10 As shown, the robot may include a robot body 700, a voice acquisition device or voice receiving device 710, a vision acquisition device or vision receiving device 720, a status acquisition device or status receiving device 730, a tactile acquisition device or tactile receiving device 740, a processor 750, a memory 760, a motion execution mechanism 770, and an end effector 780.

[0079] The voice acquisition device or voice receiver 710 can be installed on the robot body 700 or connected to the robot body 700 via a communication interface; the vision acquisition device or vision receiver 720 can be installed on the robot body 700 or connected to the robot body 700 via a communication interface; the status acquisition device or status receiver 730 can be installed on the robot body 700 or connected to the robot body 700 via a communication interface; the tactile acquisition device or tactile receiver 740 can be installed on the robot body 700 or connected to the robot body 700 via a communication interface; the voice acquisition device or voice receiver 710 is used to acquire voice signals and send them to the processor 750; the vision acquisition device or vision receiver 720 is used to acquire visual signals and send them to the processor 750; the status acquisition device or status receiver 730 is used to acquire status signals and send them to the processor 750; the tactile acquisition device or tactile receiver 740 is used to acquire tactile signals and send them to the processor 750. The memory 760 stores a computer program, and when the processor 750 executes the computer program, it can implement the previous day action generation method in any of the foregoing embodiments.

[0080] The processor 750 can generate end effector motion control data based on voice, vision, status, and tactile signals, and send the end effector motion control data to the motion actuator 770. The motion actuator 770 can drive the robot end effector 780 to perform actions corresponding to tasks such as sorting, insertion, and gripping based on the end effector motion control data.

[0081] In one embodiment, the processor 750 can be deployed inside the robot body 700, or it can be located in a server, edge computing device, or control terminal that communicates with the robot. The memory 760 can be a read-only memory, random access memory, flash memory, solid-state drive, or other storage media. When the computer program is executed by the processor 750, it can process observation data such as vision, language, touch, and state, generate and correct robot actions, and perform robot actions through the robot body.

[0082] Although the above embodiments use physical robots as typical examples for illustration, the action generation scheme based on a large model of visual, tactile, and linguistic multimodal fusion proposed in this application can also be extended to various virtual interactive carriers that have multimodal perception and need to generate continuous motion behaviors, such as augmented reality virtual characters, virtual digital anchors, game interactive characters, online meeting virtual avatars, and human-computer emotional interaction terminals. As long as the carrier adopts the implementation idea of ​​using visual, tactile, linguistic, and state multimodal signals as inputs and realizing action generation through tactile semantic constraints, tactile distribution prediction, and tactile closed-loop guidance, it falls within the protection scope of this application.

[0083] In the above embodiments, although a robot was used as an example for illustration, the method of this application can also be applied to augmented reality characters, virtual anchors, game characters, virtual avatars for online meetings, emotional interaction terminals, or other objects that need to generate facial expressions and actions based on speech. As long as it adopts the method of using macroscopic semantic actions driven by speech as the target equilibrium state and microscopic physiological actions as the perturbation force and integrating them through a virtual mass-spring-damping system, it can fall within the scope of the technical concept of this application.

[0084] The above embodiments are only for illustrating the technical concept and features of this application, and are intended to enable those skilled in the art to understand the content of this application and implement it accordingly. They should not be used to limit the scope of protection of this application. It is obvious to those skilled in the art that this application is not limited to the details of the above exemplary embodiments, and that this application can be implemented in other specific forms without departing from the spirit or basic characteristics of this application. Therefore, all embodiments should be considered exemplary and non-limiting, and all changes within the meaning and scope of equivalent elements of the technical features of this application are included within this application.

Claims

1. A method for generating robot actions based on a large-scale multimodal fusion model of visual, tactile, and language senses, characterized in that: include: Acquire observational data that includes tactile signals from long-history time, tactile signals from short-history time, tactile signals from the current time, visual signals from the current time, speech signals from the current time, and state signals from the current time; The observation data is input into a large model based on visual-tactile language multimodal fusion to obtain the predicted next cycle action command 1 and the predicted next cycle tactile distribution. The closed-loop guidance intensity is calculated based on the deviation statistics between the current tactile signal and the predicted current tactile distribution. ; Based on the closed-loop guiding strength The predicted next cycle action command 1 is corrected to obtain the corrected next cycle action command 2; The robot's actions are generated according to the next cycle action instruction 2.

2. The robot motion generation method according to claim 1, characterized in that, The large-scale model based on visual-tactile-language multimodal fusion includes, during the training phase, a visual feature extraction network, a language feature extraction network, a tactile semantic teacher extraction network, a state feature extraction network, a tactile semantic alignment network, an action-level tactile distribution prediction network, a task understanding network, a tactile feature extraction network, and an action generation expert network; and during the inference phase, the large-scale model based on visual-tactile-language multimodal fusion includes, during the inference phase, a visual feature extraction network, a language feature extraction network, a state feature extraction network, an action-level tactile distribution prediction network, a task understanding network, a tactile feature extraction network, and an action generation expert network.

3. The robot motion generation method according to claim 2, characterized in that, The visual feature extraction network obtains visual feature labels by taking the visual signal at the current time as input. The language feature extraction network uses the language signal at the current moment as input to obtain language feature labels; The tactile semantic teacher extraction network uses the tactile signals from the long historical moments as input to obtain historical tactile semantic teacher features; The state feature extraction network obtains state feature labels by taking the current state signal as input. The tactile semantic alignment network is used to constrain the semantic consistency between the historical tactile semantic query features output by the task understanding network and the historical tactile semantic teacher features; The action-level tactile distribution prediction network uses the tactile latent features output by the action generation expert network as input to obtain the predicted tactile distribution for the next period. The task understanding network takes the language feature tags and visual feature tags as input to obtain the historical tactile semantic query features and tactile language visual semantic features; The tactile feature extraction network takes the tactile signals from the short historical moments as input to obtain short historical tactile conditional feature labels; The action generation expert network takes the tactile language visual semantic features, the state feature labels, and the short history tactile condition feature labels as inputs to obtain the predicted next cycle action instruction 1 and the tactile latent features.

4. The robot motion generation method according to claim 1, characterized in that, The predicted tactile distribution at the current moment is the predicted tactile distribution for the next cycle output by the large model based on visual-tactile language multimodal fusion at the previous moment; under initial conditions, the predicted tactile distribution at the current moment is taken as the distribution of the tactile signal at the current moment, and the closed-loop guidance intensity... .

5. The robot motion generation method according to claim 2, characterized in that, The training phase includes action generation loss, action-level haptic distribution prediction loss, and haptic semantic alignment loss. The total loss can be expressed as: in, This indicates the loss generated by the action. This represents the prediction loss for action-level haptic distribution. This represents the haptic semantic alignment loss. and These are the weighting coefficients.

6. A robot motion generation system based on a large-scale multimodal fusion model of vision, touch, and language, characterized in that, To implement the robot motion generation method as described in any one of claims 1-5, comprising: The observation data input module is used to acquire observation data including tactile signals from long historical moments, tactile signals from short historical moments, tactile signals from the current moment, visual signals from the current moment, speech signals from the current moment, and state signals from the current moment. The action generation module is used to input the observation data into a large model based on visual-tactile language multimodal fusion to obtain the predicted next-cycle action command 1 and the predicted next-cycle tactile distribution. The tactile distribution storage module is used to store and update the predicted tactile distribution for the next cycle, and outputs the predicted tactile distribution at the current moment in real time. The motion correction module is used to calculate the closed-loop guidance intensity based on the deviation statistics between the current tactile signal and the predicted current tactile distribution. Based on the closed-loop guiding strength The predicted next cycle action command 1 is corrected to obtain the corrected next cycle action command 2; The motion execution module is used to generate the robot's motion according to the next cycle motion instruction 2.

7. An electronic device, the device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the robot motion generation method based on a large model of visual-tactile language multimodal fusion as described in any one of claims 1 to 5.

8. A computer-readable storage medium storing at least one instruction or at least one program, characterized in that, When the at least one instruction or the at least one program segment is loaded and executed by the processor, the robot motion generation method based on a large model of visual-tactile language multimodal fusion as described in any one of claims 1 to 5 is implemented.