Training method of multi-modal action mapping model, robot arm control method and device

CN122657643APending Publication Date: 2026-08-28XIAN JIAOTONG LIVERPOOL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610785073.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-02
Publication Date
2026-08-28

AI Technical Summary

Benefits of technology

[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122657643A_ABST
    Figure CN122657643A_ABST
Patent Text Reader

Abstract

The embodiment of the application discloses a kind of multi-modal action mapping model training method, robot control method and device.The multi-modal action mapping model training method includes: obtaining candidate demonstration data sequence, and determining sequence global time length and event detection data;According to event detection data, determine semantic event, sequence score and end event start time;According to sequence global time length, sequence score, end event start time, truncation proportion threshold and preset retention time, determine sequence truncation time;According to sequence truncation time, obtain reference demonstration data sequence, and post-processing is carried out to reference demonstration data sequence, and obtain target demonstration data sequence;According to target demonstration data sequence, multi-modal action mapping model is trained, and the trained multi-modal action mapping model is obtained.The performance of the trained multi-modal action mapping model is improved, and then the accuracy of controlling robot is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of robot imitation learning technology, and in particular to a training method for a multimodal motion mapping model, a robotic arm control method, and a device. Background Technology

[0002] In industrial manufacturing, robotic arms are commonly used to perform handling and sorting tasks to improve production efficiency and reduce costs. Therefore, improving the accuracy of controlling robotic arms is crucial. Summary of the Invention

[0003] This invention provides a training method for a multimodal motion mapping model, a robotic arm control method, and a device to improve the performance of the trained multimodal motion mapping model, thereby improving the accuracy of controlling the robotic arm by increasing the joint angles corresponding to the predicted poses determined based on the trained multimodal motion mapping model.

[0004] According to one aspect of the present invention, a method for training a multimodal action mapping model is provided, comprising: Obtain candidate demonstration data sequences of the target handle corresponding to the target robotic arm under the demonstration training task, and determine the corresponding global duration of the sequence and event detection data based on the candidate demonstration data sequences; wherein, the candidate demonstration data sequences include the demonstration end pose and the demonstration visual vector; Based on the event detection data, determine the semantic events corresponding to the corresponding candidate demonstration data sequences, and based on the candidate demonstration data sequences and the semantic events, determine the sequence score and the start time of the terminal event corresponding to the corresponding candidate demonstration data sequences. The sequence truncation time of the corresponding candidate demonstration data sequence is determined based on the global duration of the sequence, the sequence score, the start time of the terminal event, the preset truncation ratio threshold, and the preset retention duration. Based on the sequence truncation time, the candidate demonstration data sequence is truncated to obtain a reference demonstration data sequence, and the event data under different semantic events in the reference demonstration data sequence are post-processed to obtain the target demonstration data sequence; wherein, the post-processing includes intra-segment resampling and mask padding; The constructed multimodal action mapping model is trained based on the target demonstration data sequence to obtain the trained multimodal action mapping model.

[0005] According to another aspect of the present invention, a robotic arm control method is provided, comprising: Acquire the current joint data and current visual image corresponding to each target joint in the target robotic arm at the current moment; wherein, the current joint data includes the current joint angle, the current joint current, and the current joint acceleration; Based on the current joint current and the current joint acceleration, the current end-effector contact force of the target robotic arm is determined, and based on the current visual image, the current visual contact probability is determined. Based on the current end contact force and the current visual contact probability, determine whether the current moment is the moment of object contact; If not, then when the current time is the model prediction time, determine the current end-effector pose of the target robotic arm based on the current joint angle, and determine the current visual vector based on the current visual image; The current end-effector pose and the current visual vector are input into a trained multimodal motion mapping model corresponding to the target robotic arm type to obtain candidate predicted poses corresponding to candidate prediction times within a preset prediction period. Based on the candidate predicted poses, the predicted joint angles at the corresponding candidate prediction times are determined. The multimodal motion mapping model is trained using a multimodal motion mapping model training method. Based on the predicted joint angle, the operation of the target robotic arm is controlled at the corresponding candidate predicted time.

[0006] According to another aspect of the present invention, a training apparatus for a multimodal action mapping model is provided, comprising: The demonstration data acquisition module is used to acquire candidate demonstration data sequences of the target handle corresponding to the target robotic arm under the demonstration training task, and determine the corresponding global duration of the sequence and event detection data based on the candidate demonstration data sequences; wherein, the candidate demonstration data sequences include the demonstration end pose and the demonstration visual vector; The sequence scoring determination module is used to determine the semantic event corresponding to the corresponding candidate demonstration data sequence based on the event detection data, and to determine the sequence score and end event start time corresponding to the corresponding candidate demonstration data sequence based on the candidate demonstration data sequence and the semantic event. The sequence truncation time determination module is used to determine the sequence truncation time of the corresponding candidate demonstration data sequence based on the global duration of the sequence, the sequence score, the start time of the terminal event, a preset truncation ratio threshold, and a preset retention time. The target demonstration data sequence determination module is used to truncate the candidate demonstration data sequence according to the sequence truncation time to obtain a reference demonstration data sequence, and to post-process the event data under different semantic events in the reference demonstration data sequence to obtain the target demonstration data sequence; wherein, the post-processing includes intra-segment resampling and mask padding; The multimodal action mapping model training module is used to train the constructed multimodal action mapping model based on the target demonstration data sequence to obtain the trained multimodal action mapping model.

[0007] According to another aspect of the present invention, a robotic arm control device is provided, comprising: The current data acquisition module is used to acquire the current joint data and current visual image corresponding to each target joint in the target robotic arm at the current moment; wherein, the current joint data includes the current joint angle, the current joint current and the current joint acceleration; The current visual contact probability determination module is used to determine the current end contact force of the target robotic arm based on the current joint current and the current joint acceleration, and to determine the current visual contact probability based on the current visual image. The current moment determination module is used to determine whether the current moment is the moment of object contact based on the current end contact force and the current visual contact probability. The current visual vector determination module is used to determine the current end-effector pose of the target robotic arm based on the current joint angle, and to determine the current visual vector based on the current visual image, if no, when the current time is the model prediction time. The joint angle prediction module is used to input the current end-effector pose and the current visual vector into a trained multimodal motion mapping model corresponding to the target robotic arm type, obtain the candidate prediction pose corresponding to the candidate prediction time within a preset prediction period, and determine the predicted joint angle at the corresponding candidate prediction time based on the candidate prediction pose; wherein, the multimodal motion mapping model is trained using a multimodal motion mapping model training method. The robotic arm control module is used to control the operation of the target robotic arm at the corresponding candidate prediction time according to the predicted joint angle.

[0008] According to another aspect of the present invention, an electronic device is provided, comprising: One or more processors; Memory, used to store one or more programs; When one or more programs are executed by one or more processors, the one or more processors are able to execute any of the multimodal motion mapping model training methods or robotic arm control methods provided in the embodiments of the present invention.

[0009] According to another aspect of the present invention, a computer-readable storage medium is provided, which stores computer instructions for causing a processor to execute and implement any of the multimodal motion mapping model training methods or robotic arm control methods provided in the embodiments of the present invention.

[0010] This invention provides a training scheme for a multimodal motion mapping model. It involves acquiring candidate demonstration data sequences of the target handle corresponding to the target robotic arm under a demonstration training task, and determining the corresponding global sequence duration and event detection data based on these sequences. The candidate demonstration data sequences include the demonstration end-effector pose and the demonstration visual vector. Semantic events corresponding to the candidate demonstration data sequences are determined based on the event detection data, and a sequence score and end-effector event start time are determined based on the candidate demonstration data sequences and semantic events. The sequence truncation time of the corresponding candidate demonstration data sequences is determined based on the global sequence duration, sequence score, end-effector event start time, a preset truncation ratio threshold, and a preset retention time. The candidate demonstration data sequences are truncated according to the truncation time to obtain a reference demonstration data sequence. Event data under different semantic events in the reference demonstration data sequence are post-processed to obtain the target demonstration data sequence. The post-processing includes intra-segment resampling and mask filling. The constructed multimodal motion mapping model is trained based on the target demonstration data sequence to obtain the trained multimodal motion mapping model. The above scheme determines the truncation time of the corresponding candidate demonstration data sequence based on the truncation ratio threshold, preset retention time, sequence score, global sequence duration, and end event start time of the candidate demonstration data sequence. Then, it adaptively truncates the corresponding candidate demonstration data sequence according to the truncation time to obtain the reference demonstration data sequence. Finally, it performs intra-segment resampling and mask filling on the reference demonstration data sequence to obtain the target demonstration data sequence. This improves the accuracy of the determined target demonstration data sequence, thereby improving the accuracy of training the multimodal motion mapping model based on the target demonstration data sequence. This improves the performance of the trained multimodal motion mapping model, and further improves the accuracy of subsequent control of the robotic arm based on the joint angles corresponding to the predicted poses determined by the trained multimodal motion mapping model.

[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0012] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 This is a flowchart of a training method for a multimodal action mapping model provided in Embodiment 1 of the present invention; Figure 2 This is a flowchart of a training method for a multimodal action mapping model provided in Embodiment 2 of the present invention; Figure 3 This is a flowchart of a robotic arm control method provided in Embodiment 3 of the present invention; Figure 4 This is a schematic diagram of the structure of a training device for a multimodal action mapping model provided in Embodiment 5 of the present invention; Figure 5 This is a schematic diagram of the structure of a robotic arm control device provided in Embodiment Six of the present invention; Figure 6 This is a schematic diagram of the structure of an electronic device that implements a training method for a multimodal motion mapping model or a robotic arm control method, as provided in Embodiment 7 of the present invention. Detailed Implementation

[0014] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present invention, and not all of the structures.

[0015] Example 1 Figure 1 This is a flowchart of a training method for a multimodal motion mapping model provided in Embodiment 1 of the present invention. This embodiment is applicable to the training of a multimodal motion mapping model for a robotic arm based on multimodal demonstration data. The method can be executed by a training device for the multimodal motion mapping model, which can be implemented in software and / or hardware and can be configured in an electronic device that carries the training function of the multimodal motion mapping model.

[0016] See Figure 1 The training method for the multimodal action mapping model shown includes: S110. Obtain the candidate demonstration data sequence of the target handle corresponding to the target robotic arm under the demonstration training task, and determine the corresponding global duration of the sequence and event detection data based on the candidate demonstration data sequence.

[0017] In this context, the target robotic arm refers to the robotic arm that needs to be controlled. The target handle refers to the handle corresponding to the target robotic arm, which can be used to control the operation of the target robotic arm. The demonstration training task refers to the task that needs to be performed to collect demonstration data. For example, a demonstration training task could be a sorting task or a handling task.

[0018] Here, the candidate demonstration data sequence refers to the sequence of demonstration data collected by the target handle during the demonstration training task. There can be multiple candidate demonstration data sequences, and one candidate demonstration data sequence can correspond to one running trajectory of the target robotic arm. The global duration of the sequence refers to the global duration corresponding to any candidate demonstration data sequence.

[0019] Event detection data refers to data used to detect semantic events contained in candidate demonstration data sequences. For example, event detection data may include at least one of the following: the demonstration speed of the target handle, the demonstration relative distance, and the demonstration visual contact probability at different data acquisition times under any running trajectory.

[0020] Here, "demonstration speed" refers to the operating speed of the target handle at the moment of data acquisition. "Demonstration relative distance" refers to the relative distance between the target handle and the goods to be processed at the moment of data acquisition. "Demonstration visual contact probability" refers to the probability of visual contact between the target handle and the goods at the moment of data acquisition.

[0021] For example, a candidate demonstration data sequence includes demonstration end-effector pose, demonstration visual vector, demonstration velocity, demonstration relative distance, demonstration visual contact probability, demonstration acceleration, and the corresponding data acquisition time. Here, demonstration end-effector pose refers to the end-effector pose of the target handle. Demonstration visual vector refers to the visual vector extracted from the visual image acquired by the visual sensor on the target handle. Demonstration acceleration refers to the acceleration of the target handle at the data acquisition time. A candidate demonstration data sequence includes demonstration end-effector poses and demonstration visual vectors at multiple data acquisition times.

[0022] In an optional embodiment, the demonstration end-effector pose is determined as follows: for any data acquisition time, the visual end-effector pose and inertial end-effector pose of the target handle at that data acquisition time are acquired; based on the visual end-effector pose, inertial end-effector pose, preset visual weight matrix, and preset inertial weight matrix at that data acquisition time, the demonstration end-effector pose at that data acquisition time is determined.

[0023] Here, "data acquisition time" refers to the moment when the demonstration data is acquired. "Visual end-effector pose" refers to the end-effector pose of the target handle determined by the image acquired by the visual sensor on the target handle. "Inertial end-effector pose" refers to the end-effector pose of the target handle acquired by the inertial measurement unit on the target handle.

[0024] The preset visual weight matrix refers to the pre-set weight matrix corresponding to the visual end pose. The preset inertial weight matrix refers to the pre-set weight matrix corresponding to the inertial end pose. This embodiment of the invention does not impose any limitations on the setting of the preset visual weight matrix and the preset inertial weight matrix; they can be set by technicians based on experience or needs, or determined through repeated experiments. It should be noted that the preset visual weight matrix and the preset inertial weight matrix may be different at different data acquisition times, and this embodiment of the invention does not impose specific limitations on this.

[0025] For example, for any data acquisition moment, the visual end pose and inertial end pose of the target handle at that data acquisition moment are obtained; based on the preset visual weight matrix and preset inertial weight matrix at that data acquisition moment, the visual end pose and inertial end pose are fused with confidence to determine the demonstration end pose at that data acquisition moment.

[0026] For example, the exemplary end-effector pose of the target handle at any given data acquisition moment can be determined using the following formula: ; in, The demonstration end pose is represented at the data acquisition time tg. That is, the demonstration end pose represents the end pose estimate obtained after fusing visual and teleoperated handle observations at the data acquisition time tg, which includes position and attitude components. The uncertainty covariance matrix of visual observations at time tg represents the data acquisition time, i.e., the preset visual weight matrix, which characterizes the confidence level of visual pose estimation. This indicates the visual end-effector pose at time tg, where the data was acquired. The uncertainty covariance matrix of the handle observation at time tg (consistent with the visual covariance dimension) represents the data acquisition time, i.e., the preset inertial weight matrix; This represents the inertial end-effector pose at time tg, where the data was acquired. It should be noted that matrix inversion, addition, and weighting in the formula must be performed in the same coordinate system and with the same dimensions. If the observations are in different reference systems, a coordinate transformation should be performed before fusion to ensure dimensional consistency.

[0027] Understandably, by determining the demonstration end pose at any given data acquisition moment based on the visual end pose, inertial end pose, preset visual weight matrix, and preset inertial weight matrix, the fusion of the visual end pose and inertial end pose at any given data acquisition moment is achieved, thereby improving the accuracy of the determined demonstration end pose at that given data acquisition moment.

[0028] S120. Determine the semantic events corresponding to the corresponding candidate demonstration data sequences based on the event detection data, and determine the sequence score and terminal event start time corresponding to the corresponding candidate demonstration data sequences based on the candidate demonstration data sequences and semantic events.

[0029] Semantic events refer to the key semantic events included in the candidate demonstration data sequence. Sequence scores can be used to quantify the quality of the candidate demonstration data sequence. The start time of the last event refers to the timestamp of the last semantic event occurring in the candidate demonstration data sequence.

[0030] For example, for any candidate demonstration data sequence, based on the preset semantic event detection strategy corresponding to the demonstration training task, the semantic events corresponding to the candidate demonstration data sequence are determined according to the event detection data of the candidate demonstration data sequence. Here, the preset semantic event detection strategy refers to a pre-set strategy used to determine the semantic events involved in the candidate demonstration data sequence.

[0031] S130. Determine the sequence truncation time of the corresponding candidate demonstration data sequence based on the global sequence duration, sequence score, start time of the terminal event, preset truncation ratio threshold, and preset retention duration.

[0032] In this embodiment of the invention, the size of the truncation ratio threshold is not limited. It can be set by a technician based on experience or needs, or determined through repeated experiments. For example, the truncation ratio threshold may include a preset upper limit value and a preset lower limit value for the truncation ratio.

[0033] In this embodiment of the invention, the preset retention time is not limited in any way. It can be set by technicians based on experience or needs, or determined repeatedly through a large number of experiments.

[0034] The sequence truncation time refers to the moment when the candidate demonstration data sequence is truncated. The sequence truncation time can be understood as the end moment of the reference demonstration data sequence. For example, at least a portion of the candidate demonstration data located after the sequence truncation time is removed to obtain the reference demonstration data sequence.

[0035] S140. Based on the sequence truncation time, the candidate demonstration data sequence is truncated to obtain the reference demonstration data sequence. The event data under different semantic events in the reference demonstration data sequence are then post-processed to obtain the target demonstration data sequence.

[0036] The reference demonstration data sequence refers to the sequence of demonstration data retained after pruning the candidate demonstration data sequence. At least a portion of the demonstration data after the truncation point in the candidate demonstration data sequence is deleted to obtain the reference demonstration data sequence.

[0037] For example, event data refers to the sample data corresponding to each semantic event in the reference sample data sequence. For example, post-processing includes intra-segment resampling and mask padding.

[0038] For example, intra-segment resampling of event data under different semantic events in any candidate demonstration data sequence means extracting a corresponding number of data points from the event data under the corresponding semantic event in the candidate demonstration data sequence, based on the extraction quantity corresponding to different semantic events. The extraction quantity refers to the number of event data points corresponding to any semantic event in the target demonstration data sequence. The extraction quantity corresponding to different semantic events can be different.

[0039] Continuing from the previous example, if the number of event data extracted under any semantic event is less than the number of extractions corresponding to that semantic event (i.e., the number of event data under that semantic event is less than the number of extractions corresponding to that semantic event), then the missing event data under that semantic event is filled based on the last event data under that semantic event, which is called mask filling.

[0040] The target demonstration data sequence refers to the demonstration data sequence obtained by resampling within segments and masking the reference demonstration data sequence.

[0041] S150. Train the constructed multimodal action mapping model based on the target demonstration data sequence to obtain the trained multimodal action mapping model.

[0042] The multimodal motion mapping model can be used to predict the end-effector pose of a target robotic arm at different prediction times. In other words, it can estimate the end-effector pose of a target robotic arm at different times within a future period. The multimodal motion mapping model can be understood as a cross-platform motion mapping / redirection model, or a pose prediction model.

[0043] In one optional embodiment, the constructed multimodal action mapping model is trained based on the target demonstration data sequence to obtain a trained multimodal action mapping model, including: determining the demonstration input pose and demonstration label pose from the target end pose of the target demonstration data sequence, and determining the demonstration input vector from the target visual vector; inputting the demonstration input pose and demonstration input vector into the constructed multimodal action mapping model to obtain the demonstration prediction pose; determining the model loss value based on the demonstration prediction pose, demonstration label pose, demonstration input pose, target demonstration data sequence, intermediate semantic vector of the multimodal action mapping model, and preset standard semantic vector, and training the constructed multimodal action mapping model based on the model loss value to obtain a trained multimodal action mapping model.

[0044] Here, the target end pose refers to the demonstration end pose in the target demonstration data sequence. The demonstration input pose refers to the target end pose input into the multimodal action mapping model during training. The demonstration label pose refers to the label used by the multimodal action mapping model for pose prediction during model training; that is, the demonstration label pose refers to the target end pose used as the label.

[0045] Here, the target visual vector refers to the demonstration visual vector in the target demonstration data sequence. The demonstration input vector refers to the target visual vector input into the multimodal action mapping model during training. The demonstration predicted pose refers to the output of the multimodal action mapping model during training.

[0046] For example, the demonstration input time and demonstration prediction time are determined from the data acquisition time corresponding to the target demonstration data sequence; the target end pose at the demonstration input time is used as the demonstration input pose; the target end pose at the demonstration prediction time is used as the demonstration label pose; the target visual vector at the demonstration input time is used as the demonstration input vector; the demonstration input pose and demonstration input vector are input into the constructed multimodal motion mapping model to obtain the demonstration prediction pose at the demonstration prediction time.

[0047] In this invention, the demonstration input moment refers to the data acquisition moment during training, used to provide demonstration data as input to the multimodal motion mapping model. The demonstration prediction moment refers to the moment during training when pose prediction is performed based on the multimodal motion mapping model. This embodiment of the invention does not impose any limitations on the setting of the demonstration input moment and the demonstration prediction moment; these can be set by technicians based on experience or needs, or determined through extensive experimentation. For example, the demonstration input moment and the demonstration prediction moment can be determined based on a preset prediction duration. The number of demonstration prediction moments between any two adjacent demonstration input moments can be at least one.

[0048] For example, a multimodal action mapping model includes an encoder and an adapter. The encoder can be used to determine intermediate semantic vectors. The adapter can be used to determine the demonstration predicted pose. The intermediate semantic vectors can represent the latent semantic representation output by the encoder. The standard semantic vectors can represent the latent semantic representation output by the encoder corresponding to the standard demonstration data sequence of the target handle under the demonstration training task. The standard demonstration data sequence refers to the pre-set demonstration data sequence corresponding to the standard running trajectory of the target handle under the demonstration training task. The model loss value refers to the loss function value of the multimodal action mapping model.

[0049] Understandably, by inputting the demonstration input pose and demonstration input vector into the constructed multimodal action mapping model, the demonstration predicted pose is obtained. Based on the demonstration predicted pose, demonstration label pose, demonstration input pose, target demonstration data sequence, intermediate semantic vectors of the multimodal action mapping model, and preset standard semantic vectors, the model loss value is determined. Finally, the constructed multimodal action mapping model is trained based on the model loss value to obtain a trained multimodal action mapping model, thereby improving the performance of the trained multimodal action mapping model.

[0050] In one optional embodiment, the model loss value is determined based on the demonstration predicted pose, demonstration label pose, demonstration input pose, target demonstration data sequence, intermediate semantic vector of the multimodal action mapping model, and preset standard semantic vector. This includes: determining a pose loss value based on the demonstration predicted pose and demonstration label pose; determining a predicted demonstration trajectory based on the demonstration predicted pose and demonstration input pose, and determining a trajectory loss value based on the predicted demonstration trajectory and the target demonstration trajectory corresponding to the target demonstration data sequence; determining a semantic loss value based on the intermediate semantic vector and standard semantic vector; wherein the intermediate semantic vector is the latent semantic representation output by the encoder in the multimodal action mapping model; and determining a model loss value based on the pose loss value, trajectory loss value, and semantic loss value.

[0051] The pose loss value quantifies the difference between the predicted and labeled poses of the demonstration model. The predicted demonstration trajectory is the trajectory determined by the predicted and input poses associated with the same target demonstration data sequence. The target demonstration trajectory is the trajectory determined based on the target end pose in the target demonstration data sequence. There is a one-to-one correspondence between the predicted and target demonstration trajectories. The trajectory loss value quantifies the difference between the predicted and target demonstration trajectories. The semantic loss value quantifies the difference between the intermediate semantic vector and the standard semantic vector.

[0052] For example, the model loss value is determined based on the pose loss value, trajectory loss value, semantic loss value, preset pose loss weights, preset trajectory loss weights, and preset semantic loss weights. The pose loss weights can be used to quantify the importance of the pose loss value in the model loss value. The trajectory loss weights can be used to quantify the importance of the trajectory loss value in the model loss value. The semantic loss weights can be used to quantify the importance of the semantic loss value in the model loss value. This embodiment of the invention does not impose any limitations on the settings of the pose loss weights, trajectory loss weights, and semantic loss weights; these can be set by technicians based on experience or needs, or determined through extensive experimentation.

[0053] Understandably, by determining the model loss value based on the given pose loss value, trajectory loss value, and semantic loss value, the model loss value is determined from multiple dimensions, thus improving the accuracy of the determined model loss value.

[0054] For example, when the model loss value is less than a preset loss threshold, training of the multimodal action mapping model is stopped, resulting in a trained multimodal action mapping model. This embodiment of the invention does not impose any limitation on the size of the preset loss threshold; it can be set by technicians based on experience or needs, or determined repeatedly through numerous experiments.

[0055] It should be noted that different types of target robotic arms require different multimodal motion mapping models.

[0056] This invention provides a training scheme for a multimodal motion mapping model. It involves acquiring candidate demonstration data sequences of the target handle corresponding to the target robotic arm under a demonstration training task, and determining the corresponding global sequence duration and event detection data based on these sequences. The candidate demonstration data sequences include the demonstration end-effector pose and the demonstration visual vector. Semantic events corresponding to the candidate demonstration data sequences are determined based on the event detection data, and a sequence score and end-effector event start time are determined based on the candidate demonstration data sequences and semantic events. The sequence truncation time of the corresponding candidate demonstration data sequences is determined based on the global sequence duration, sequence score, end-effector event start time, a preset truncation ratio threshold, and a preset retention time. The candidate demonstration data sequences are truncated according to the truncation time to obtain a reference demonstration data sequence. Event data under different semantic events in the reference demonstration data sequence are post-processed to obtain the target demonstration data sequence. The post-processing includes intra-segment resampling and mask filling. The constructed multimodal motion mapping model is trained based on the target demonstration data sequence to obtain the trained multimodal motion mapping model. The above scheme determines the truncation time of the corresponding candidate demonstration data sequence based on the truncation ratio threshold, preset retention time, sequence score, global sequence duration, and end event start time of the candidate demonstration data sequence. Then, it adaptively truncates the corresponding candidate demonstration data sequence according to the truncation time to obtain the reference demonstration data sequence. Finally, it performs intra-segment resampling and mask filling on the reference demonstration data sequence to obtain the target demonstration data sequence. This improves the accuracy of the determined target demonstration data sequence, thereby improving the accuracy of training the multimodal motion mapping model based on the target demonstration data sequence. This improves the performance of the trained multimodal motion mapping model, and in turn, improves the accuracy of subsequent control of the robotic arm based on the joint angles corresponding to the predicted pose determined by the trained multimodal motion mapping model.

[0057] Example 2 Figure 2 This is a flowchart of a training method for a multimodal action mapping model provided in Embodiment 2 of the present invention. Based on the above embodiments, this embodiment further refines the operation of "determining the sequence score and terminal event start time corresponding to the corresponding candidate demonstration data sequence based on the candidate demonstration data sequence and semantic events" into "determining the evaluation data and terminal event start time corresponding to the corresponding candidate demonstration data sequence based on the candidate demonstration data sequence and semantic events; determining the sequence score corresponding to the corresponding candidate demonstration data sequence based on the evaluation data," thereby improving the sequence score determination mechanism. It should be noted that for parts not detailed in this embodiment, please refer to the descriptions in other embodiments.

[0058] See Figure 2 The training method for the multimodal action mapping model shown includes: S210. Obtain the candidate demonstration data sequence of the target handle corresponding to the target robotic arm under the demonstration training task, and determine the corresponding global duration of the sequence and event detection data based on the candidate demonstration data sequence.

[0059] The candidate demonstration data sequence includes the demonstration end pose and the demonstration visual vector.

[0060] S220. Determine the semantic events corresponding to the corresponding candidate demonstration data sequences based on the event detection data, and determine the evaluation data and the start time of the terminal events corresponding to the corresponding candidate demonstration data sequences based on the candidate demonstration data sequences and semantic events.

[0061] Evaluation data refers to data that can be used to determine the sequence score of candidate demonstration data sequences. For example, evaluation data may include the demonstration velocity and acceleration of the target handle, as well as the event duration of each semantic event corresponding to the candidate demonstration data sequence.

[0062] S230. Based on the evaluation data, determine the sequence score corresponding to the candidate demonstration data sequence.

[0063] For example, for any candidate demonstration data sequence, based on a preset smoothness scoring strategy, the sequence smoothness score of the candidate demonstration data sequence is determined according to the demonstration speed and demonstration acceleration of the candidate demonstration data sequence; based on a preset consistency scoring strategy, the semantic event alignment consistency score of the candidate demonstration data sequence is determined according to the event duration of each semantic event corresponding to the candidate demonstration data sequence; and based on the sequence smoothness score, the semantic event alignment consistency score, the preset smoothness weight, and the preset consistency weight, the sequence score of the candidate demonstration data sequence is determined.

[0064] Among them, the preset smoothness scoring strategy refers to the pre-set strategy used to determine the smoothness score of the sequence. The preset consistency scoring strategy refers to the pre-set strategy used to determine the semantic event alignment consistency score.

[0065] The sequence smoothness score quantifies the smoothness of the candidate demonstration data sequence. The semantic event alignment consistency score quantifies the consistency between the event duration of different semantic events in the candidate demonstration data sequence and the corresponding preset event standard duration. The event standard duration refers to the preset average duration of a semantic event. Different semantic events may correspond to different event standard durations.

[0066] The preset smoothing weight can be used to quantify the importance of sequence smoothness score in sequence scoring. The preset consistency weight can be used to quantify the importance of semantic event alignment consistency score in sequence scoring. This embodiment of the invention does not impose any limitations on the magnitude of the preset smoothing weight and the preset consistency weight; they can be set by technicians based on experience or needs, or determined repeatedly through numerous experiments.

[0067] For example, the sequence score can be determined using the following formula: ; in, Let represent the sequence score of the i-th candidate demonstration data sequence. The sequence score is dimensionless, and Belongs to [0,1]; Indicates the preset smoothing weight; Represents the sequence smoothness score of the i-th candidate demonstration data sequence; This indicates a preset consistency weight; This represents the semantic event alignment consistency score of the i-th candidate sample data sequence. It should be noted that here... + =1.

[0068] S240. Determine the sequence truncation time of the corresponding candidate demonstration data sequence based on the global duration of the sequence, the sequence score, the start time of the terminal event, the preset truncation ratio threshold, and the preset retention time.

[0069] In one optional embodiment, determining the sequence truncation time of the corresponding candidate demonstration data sequence based on the global sequence duration, sequence score, start time of the terminal event, preset truncation ratio threshold, and preset retention duration includes: for any candidate demonstration data sequence, determining the sequence truncation ratio of the candidate demonstration data sequence based on the sequence score and truncation ratio threshold; and determining the sequence truncation time of the candidate demonstration data sequence based on the sequence truncation ratio, global sequence duration, preset retention duration, and start time of the terminal event.

[0070] The sequence truncation ratio refers to the ratio between the portion of the candidate demonstration data sequence that needs to be pruned and the original total.

[0071] For example, the sequence truncation ratio can be determined using the following formula: ; in, It represents the sequence truncation ratio of the i-th candidate demonstration data sequence. It is a dimensionless scalar used to trim the redundant interval at the end of the candidate demonstration data sequence by time or frame number. This indicates the preset lower limit of the cutoff ratio; This indicates the preset upper limit of the cutoff ratio.

[0072] For example, the sequence truncation time can be determined using the following formula: ; in, This represents the sequence truncation time of the i-th candidate demonstration data sequence; Indicates the start time of the last event of the i-th candidate demonstration data sequence (in seconds), used for event protection; This indicates the preset retention period, which is an additional retention time after the last semantic event is retained, in order to avoid accidentally deleting key post-processing actions; This represents the global duration of the i-th candidate demonstration data sequence. It should be noted that the sequence truncation time should be the maximum of the position calculated using the sequence truncation ratio and the event protection position to ensure semantic integrity.

[0073] Understandably, determining the sequence truncation ratio of the corresponding candidate demonstration data sequence based on the sequence score and truncation ratio threshold improves the accuracy of the determined sequence truncation ratio. At the same time, determining the sequence truncation time of the corresponding candidate demonstration data sequence based on the sequence truncation ratio, the global duration of the sequence, the preset retention duration, and the start time of the end event achieves event anchor point protection and improves the accuracy of the determined sequence truncation time.

[0074] The embodiments of the present invention employ an adaptive sequence truncation ratio, which differs from the fixed truncation ratio method used in the prior art. The embodiments of the present invention retain more tail information in high-quality examples and truncate more in low-quality examples to reduce noise, thus balancing information preservation and noise reduction.

[0075] S250. Based on the sequence truncation time, the candidate demonstration data sequence is truncated to obtain the reference demonstration data sequence. The event data under different semantic events in the reference demonstration data sequence are then post-processed to obtain the target demonstration data sequence.

[0076] Post-processing includes intra-segment resampling and mask padding.

[0077] S260. Train the constructed multimodal action mapping model based on the target demonstration data sequence to obtain the trained multimodal action mapping model.

[0078] This invention provides a training scheme for a multimodal action mapping model. By refining the process of determining the sequence score and terminal event start time of a candidate demonstration data sequence based on the candidate demonstration data sequence and semantic events, it further refines this process to determining the evaluation data and terminal event start time of the candidate demonstration data sequence based on the candidate demonstration data sequence and semantic events; and then determining the sequence score of the corresponding candidate demonstration data sequence based on the evaluation data, thus improving the sequence score determination mechanism. This scheme, for any candidate demonstration sequence data, determines the sequence score based on the evaluation data of that candidate demonstration sequence data, improving the accuracy of the determined sequence score.

[0079] Example 3 Figure 3 This is a flowchart of a robotic arm control method provided in Embodiment 3 of the present invention. This embodiment is applicable to situations where the operation of a robotic arm is controlled. The method can be executed by a robotic arm control device, which can be implemented in software and / or hardware and can be configured in an electronic device that carries the robotic arm control function.

[0080] See Figure 3 The robotic arm control method shown includes: S310. Obtain the current joint data and current visual image corresponding to each target joint in the target robotic arm at the current moment; wherein, the current joint data includes the current joint angle, the current joint current and the current joint acceleration.

[0081] Here, "current moment" refers to the moment when the target robotic arm is currently being controlled. For example, the current moment could be the moment when the target robotic arm is performing a target task. For instance, the target task could be a sorting or handling task performed by the target robotic arm. The target robotic arm is the robotic arm that needs to be controlled.

[0082] Here, "target joint" refers to the joint on the target robotic arm. "Current joint data" refers to the data associated with the target joint. "Current visual image" refers to the visual image acquired by the vision sensors on the target robotic arm at the current moment.

[0083] For example, the current visual image may include the end effector of the target robotic arm and / or the object it is in contact with. The object in contact refers to an object that the target robotic arm is in contact with, such as goods to be processed or obstacles. For instance, if the target robotic arm is currently performing a handling task, the current visual image may include the gripper of the target robotic arm and / or the goods to be handled.

[0084] For example, the current joint data includes the current joint angle, current joint current, and current joint acceleration. The current joint angle refers to the joint angle of each target joint in the target robotic arm at the current moment. The current joint current refers to the joint current of each target joint in the target robotic arm at the current moment. The current joint acceleration refers to the joint acceleration of each target joint in the target robotic arm at the current moment.

[0085] Specifically, it acquires the current visual image of the target robotic arm at the current moment, as well as the current joint data corresponding to each target joint.

[0086] S320. Based on the current joint current and current joint acceleration, determine the current end contact force of the target robotic arm, and based on the current visual image, determine the current visual contact probability.

[0087] Here, the current end-effector contact force refers to the end-effector contact force of the target robotic arm at the current moment. For example, the current end-effector contact force is a vector with two attributes: magnitude and direction.

[0088] For example, for any target joint, the current joint torque of the target joint is determined based on the preset joint association data, the current joint acceleration, and the current joint current; the current end contact force is determined based on the current joint torque of each target joint. Here, the preset joint association data refers to the pre-set data corresponding to the target joint. The preset joint association data may include the motor torque constant of the target joint, the torque estimation compensation of the target joint, and the joint inertia matrix or inertia term of the target joint.

[0089] The current visual contact probability refers to the probability of visual contact with the target robotic arm at the current moment. For example, in a scenario where goods are sorted or transported based on a target robotic arm, the visual contact probability can be understood as the likelihood that the target robotic arm will successfully identify the goods to be transported or sorted in a complex environment and establish an effective visual association through visual sensors.

[0090] S330. Based on the current end contact force and the current visual contact probability, determine whether the current moment is the moment of object contact.

[0091] The moment of object contact refers to the moment when the target robotic arm comes into contact with the object. For example, if the target robotic arm is currently performing a handling task, the object in contact with the target robotic arm can be the goods to be processed or an obstacle, etc.

[0092] In one optional embodiment, the current end contact force value is determined based on the current end contact force; if the current end contact force value is greater than a preset end contact force threshold and the current visual contact probability is greater than a preset contact probability threshold, then the current moment is determined to be the moment of object contact.

[0093] Here, the current end-effector contact force value refers to the magnitude of the current end-effector contact force. For example, the L2 norm of the current end-effector contact force can be used as the current end-effector contact force value. This embodiment of the invention does not impose any limitations on the magnitude of the preset end-effector contact force threshold and the preset contact probability threshold; these can be set by technicians based on experience or needs, or determined through repeated experiments.

[0094] Understandably, when the current end contact force value is greater than the preset end contact force threshold, and the current visual contact probability is greater than the preset contact probability threshold, the current moment is determined to be the moment of object contact, thus improving the accuracy of determining the current moment as the moment of object contact.

[0095] In another optional embodiment, determining whether the current moment is the moment of object contact based on the current end contact force and the current visual contact probability includes: if the current end contact force value is less than or equal to a preset end contact force threshold, and / or the current visual contact probability is less than or equal to a preset contact probability threshold, then prohibiting the current moment from being used as the moment of object contact.

[0096] S340. If not, then when the current time is the model prediction time, determine the current end pose of the target robotic arm based on the current joint angle, and determine the current visual vector based on the current visual image.

[0097] Here, the model prediction time refers to the moment when the candidate predicted pose is determined based on the multimodal motion mapping model. For example, if the current moment is a non-object contact moment and there is no corresponding candidate predicted pose in the next moment adjacent to the current moment, the current moment is determined as the model prediction time.

[0098] Here, the current end-effector pose refers to the end-effector pose of the target robotic arm at the current moment. For example, the current end-effector pose can be determined based on forward kinematics and the current joint angles.

[0099] Here, the current visual vector refers to the feature vector in the current visual image. For example, the current visual vector can represent local texture / depth information of visual perception. The current visual vector can be extracted from the current visual image.

[0100] For example, if the current time is neither the time of object contact nor the time of model prediction, then no further processing is required.

[0101] S350. Input the current end-effector pose and the current visual vector into the trained multimodal motion mapping model corresponding to the target robotic arm type to obtain the candidate prediction pose corresponding to the candidate prediction time within the preset prediction period, and determine the prediction joint angle at the corresponding candidate prediction time based on the candidate prediction pose.

[0102] Here, "target robotic arm type" refers to the variety of target robotic arms. Different robotic arm types can correspond to different multimodal motion mapping models.

[0103] The multimodal motion mapping model can be used to predict the end-effector pose of a target robotic arm at different prediction times; that is, it can be used to estimate the end-effector pose of a target robotic arm at different times within a future period. For example, the multimodal motion mapping model can be trained using the training methods for multimodal motion mapping models.

[0104] The preset prediction time period refers to the pre-set time period for end-effector pose prediction based on the multimodal motion mapping model. The candidate prediction time refers to the end-effector pose prediction time within the preset prediction time period. The candidate predicted pose refers to the end-effector pose of the target robotic arm predicted based on the multimodal motion mapping model at the candidate prediction time. The predicted joint angles refer to the predicted joint angles of each target joint of the target robotic arm at the candidate prediction time.

[0105] For example, the preset prediction period can be determined based on the current time and a preset prediction duration. The preset prediction duration refers to the prediction duration set in advance. This embodiment of the invention does not limit the size of the preset prediction duration; it can be set by a technician based on experience or needs, or determined repeatedly through numerous experiments.

[0106] For example, the embodiments of the present invention do not impose any limitations on the setting of candidate prediction times within the preset prediction period, and can be set by technicians based on experience or needs.

[0107] For example, based on inverse kinematics, the predicted joint angle at any candidate prediction time can be determined according to the candidate predicted pose at any candidate prediction time.

[0108] S360. Based on the predicted joint angle, control the operation of the target robotic arm at the corresponding candidate predicted time.

[0109] For example, for any candidate prediction time, the operation control command at the candidate prediction time can be determined based on the predicted joint angle at the candidate prediction time; and the operation of the target robotic arm at the candidate prediction time can be controlled according to the operation control command.

[0110] Among them, operation control commands refer to the commands that control the operation of the target robotic arm. For example, operation control commands may include joint torque or joint current.

[0111] This invention provides a robotic arm control scheme. It acquires current joint data and a current visual image corresponding to each target joint in the target robotic arm at the current moment. The current joint data includes the current joint angle, current joint current, and current joint acceleration. Based on the current joint current and acceleration, the current end-effector contact force of the target robotic arm is determined, and based on the current visual image, the current visual contact probability is determined. Based on the current end-effector contact force and the current visual contact probability, it is determined whether the current moment is an object contact moment. If not, when the current moment is a model prediction moment, the current end-effector pose of the target robotic arm is determined based on the current joint angle, and the current visual vector is determined based on the current visual image. The current end-effector pose and the current visual vector are input to a trained multimodal motion mapping model corresponding to the target robotic arm type to obtain candidate prediction poses corresponding to candidate prediction moments within a preset prediction time period. Based on the candidate prediction poses, the predicted joint angles at the corresponding candidate prediction moments are determined. The multimodal motion mapping model is trained using a multimodal motion mapping model training method. Based on the predicted joint angles, the operation of the target robotic arm at the corresponding candidate prediction moments is controlled. The above scheme determines whether the current moment is an object contact moment based on the current end-effector contact force and the current visual contact probability. If not, when the current moment is a model prediction moment, the current end-effector pose and the current visual vector are input into the trained multimodal motion mapping model to output candidate predicted poses. Based on the candidate predicted poses, predicted joint angles are determined. Finally, the target robotic arm is controlled to run at the corresponding candidate prediction moment based on the predicted joint angles. This improves the accuracy of the candidate predicted poses determined by the trained multimodal motion mapping model, and improves the accuracy of controlling the target robotic arm based on the predicted joint angles determined by the candidate predicted poses. In other words, it improves the accuracy of controlling the robotic arm based on the joint angles corresponding to the predicted poses determined by the trained multimodal motion mapping model.

[0112] Based on the above technical solution, an optional example is also provided. For example, after determining whether the current moment is the moment of object contact, the method further includes: if so, determining the reference predicted joint torque corresponding to the next moment based on the current end contact force and the current joint angle of the target robotic arm; and controlling the operation of the target robotic arm at the next moment based on the reference predicted joint torque.

[0113] Here, "next moment" refers to the moment immediately following the current moment. "Reference predicted joint torque" refers to the predicted joint torque corresponding to the next moment.

[0114] For example, the current impedance control component is determined based on the current end contact force and the preset impedance parameters; the current position control component is determined based on the current joint angle; and the reference predicted joint torque corresponding to the next moment is determined based on the current impedance control component, the current position control component, and the preset component weights.

[0115] Here, the current impedance control component refers to the impedance control component at the current moment. The current position control component refers to the position control component at the current moment. For example, both the position control component and the impedance control component can be the joint torque.

[0116] Here, the preset impedance parameters refer to impedance parameters that are set in advance. For example, preset impedance parameters may include stiffness and damping. This embodiment of the invention does not limit the magnitude of the preset component weights; these can be set by technicians based on experience or needs, or determined through extensive experimentation.

[0117] It should be noted that if the current moment is the moment of object contact, and there is a pre-determined predicted joint angle for the next moment adjacent to the current moment, after determining the reference predicted joint torque for the next moment, the target robotic arm is controlled to run at the next moment based on the reference predicted joint torque, and the predicted joint angle for the next moment can be deleted.

[0118] Understandably, when the object is in contact at the current moment, the reference predicted joint torque for the next moment is determined based on the current end contact force and the current joint angle of the target robotic arm. Finally, the operation of the target robotic arm is controlled according to the reference predicted joint torque at the next moment, which improves the accuracy of the determined reference predicted joint torque for the next moment, thereby improving the accuracy of controlling the operation of the target robotic arm according to the reference predicted joint torque at the next moment and improving the smoothness of the target robotic arm's operation.

[0119] Example 4 This invention provides an optional example based on the above embodiments. It should be noted that for parts not described in detail in this invention's embodiments, please refer to the descriptions in other embodiments.

[0120] This invention relates to robotic arm teleoperation, embodied perception, multimodal data fusion, imitation learning, and industrial handling and sorting tasks. The focus is on achieving robust motion redirection between actuators and cross-platform demonstration data acquisition on a low-cost hardware platform. Technical aspects include multi-source signal fusion based on RGB-D (color depth image) vision and handheld teleoperation handles; data preprocessing driven by demonstration quality through adaptive time-step truncation and mask filling; intra-segment alignment based on semantic events; few-sample cross-platform mapping using latent semantic encoders and lightweight adapters; and task-oriented impedance switching control combining current-kinematic estimation and visual redundancy determination under sensorless conditions. This field covers the entire technology chain from data acquisition and representation learning to control strategies and system engineering implementation, aiming to improve demonstration sample efficiency, contact robustness, and industrial deployment feasibility under low-cost conditions.

[0121] Overview of Existing Technologies: 1) Low-cost / Replicable Teleoperation and Data Acquisition Platforms. Existing technologies propose low-cost, replicable teleoperation hardware and data acquisition pipelines. These works focus on hardware replicability, handle / gripper interfaces, and force feedback for dataset construction and learning research. 2) Methods for Multimodal Data Preprocessing and Feature Extraction. Existing solutions describe time-step truncation and padding / masking as core strategies for multimodal sequence preprocessing, aiming to remove meaningless time steps at the end of the demonstration sequence to reduce training noise. 3) Research on Vision, Body / Handle Fusion, and Cross-Platform Mapping. Systems such as ACE (Vision-Exoskeleton / Handle Integration) emphasize high-precision hand semantic capture through the fusion of vision and operator-worn sensors to drive relocalization of multiple actuators.

[0122] Challenges and shortcomings of existing technologies: 1) Insufficient multimodal alignment and semantic consistency: Alignment based solely on timestamps or processing by a fixed window makes it difficult to guarantee consistent mapping of semantic events such as "contact / insertion / release" across different demonstrations, thus affecting the learning quality of event-driven segments (self-supervised image feature methods such as DINO (Self-Distillation with No Labels) cannot independently solve the temporal semantic alignment problem). 2) Low-cost platforms lack proprioceptive force perception but rely on contact determination: Many low-cost systems (to save costs) do not equip themselves with high-precision end force sensors, making it difficult to reliably sense and switch impedance / compliance control when contact / jamming occurs; existing approaches either require expensive sensors or rely solely on visual determination, resulting in insufficient robustness. 3) Rigidity of truncation strategies and risk of information loss: Existing fixed truncation ratios (e.g., fixed a=10%) may incorrectly delete key "post-processing" actions (e.g., small adjustments after insertion / removal, stabilization actions) in some demonstrations, leading to loss of temporal integrity of training data and a decrease in policy generalization ability. 4) Sample efficiency issues in cross-platform mapping: Different actuators (human-like hand, two-finger gripper, parallel gripper) have significant differences in action semantics and joint degrees of freedom, making direct mapping or simple linear transformation difficult to generalize. A latent semantic layer and adapter mechanism need to be designed to improve the adaptability with few samples (Pika data pipeline and ACT (Action Chunking with Transformers) model provide data support, but do not cover the complete adaptive truncation + adapter combination (i.e., adaptive stage adapter)). Based on the above problems, this invention proposes a low-cost, cross-platform vision-teleoperated handle fusion multimodal demonstration acquisition and action mapping system for industrial handling and sorting.

[0123] The technical problems solved by the embodiments of this invention include: 1) Multimodal data redundancy and noise suppression: The fixed truncation and padding methods of existing demonstration data are prone to accidentally deleting key actions or causing a large amount of invalid data to enter the training. The embodiments of this invention fundamentally reduce redundant noise at the tail of the demonstration by using an adaptive truncation mechanism driven by demonstration quality, combined with the protection of key moments of semantic events and intra-segment resampling methods, thereby increasing the proportion of effective data and making the training process more stable and reliable. 2) Cross-modal temporal alignment and semantic consistency: Multimodal signals such as vision, handle, joint, and current are not synchronized in time, which can easily lead to semantic misalignment. The embodiments of this invention use semantic events as anchor points for segment alignment and perform unified time normalization on each segment to ensure that different modalities are consistent in key behaviors (such as approach, contact, stabilization, and release), thereby improving the model learning quality. 3) Cross-platform action retargeting efficiency under low sample conditions: Traditional action mapping methods usually require a large amount of demonstration data when facing new robotic arm platforms. This invention employs a structure combining "latent hand semantic representation and lightweight adapter network," keeping the backbone representation unchanged. Only slight adjustments to the adapter module on the target robotic arm side are needed to complete motion redirection, significantly reducing the number of demonstration samples required for adaptation and shortening deployment time. 4) Contact sensing and safety control under sensorless conditions: Low-cost robotic arm platforms generally lack high-precision force sensors, making it difficult to accurately identify contact states. This invention combines estimation methods based on motor current and kinematic changes, and utilizes visual information as an auxiliary judgment criterion to construct a reliable contact detection mechanism. Based on this, a task-oriented impedance switching strategy is triggered to achieve safe and reliable physical interaction. 5) Engineering-grade, low-cost, and cross-platform system integration: To address the problems of complex deployment and difficult process reuse across different robotic arm platforms, this invention standardizes and engineers each module, forming a stable data acquisition, training, and deployment pipeline. This pipeline can be quickly reused and expanded under various hardware conditions and industrial scenarios, significantly reducing system integration costs.

[0124] This invention provides a low-cost, cross-platform vision-teleoperation handle fusion teleoperation system and its multimodal data acquisition method, achieving efficient acquisition, unified mapping, and task-oriented control of high-quality demonstration data. Through semantic event-driven adaptive time-step truncation and intra-segment resampling mechanisms related to demonstration quality, tail-end redundancy is effectively removed and demonstration noise is reduced. The system employs vision-handle confidence fusion, latent semantic encoding, and a lightweight adapter network to achieve cross-actuator motion redirection under limited sample conditions. Simultaneously, based on robotic arm current feedback—kinematic contact estimation and visual redundancy determination, task-oriented impedance switching is completed, enabling the low-cost platform to maintain reliable contact recognition and compliant control even without sensors. The overall solution combines the advantages of data-driven and modular approaches, improving demonstration utilization efficiency and training stability, and accelerating the deployment of heterogeneous robotic arms.

[0125] For example, visual observation ∈SE(3) (pose / attitude estimation obtained from RGB-D+PnP), and handle measurement The end pose is obtained by confidence-weighted fusion of ∈SE(3) (estimated by the handle IMU / encoder). Semantic event detection is performed on the candidate demonstration data sequence to obtain the semantic events included in the corresponding candidate demonstration data sequence.

[0126] For example, if the demonstration training task is a transport task, the event set can include approach, initial contact, stable contact, and release. Event detection can be based on multi-source signals, such as demonstration speed, demonstration relative distance, and demonstration visual contact probability. Example criteria: Under the initial contact event: the demonstration visual contact probability is greater than a preset contact probability threshold, or the demonstration relative distance gradually decreases; Under the stable contact event: after the initial contact event, the demonstration speed is constant; Under the release event: the demonstration relative distance gradually increases; Under the initial contact event...

[0127] For example, if the object used for the demonstration training task is a target robotic arm, the event detection data may also include the demonstration feedback current change rate. If the demonstration sequence task is a handling task, then at the initial contact event: the absolute value of the demonstration feedback current change rate is greater than a preset current change rate threshold. Here, the demonstration feedback current change rate refers to the change rate of the feedback current of the target robotic arm at the data acquisition moment. This embodiment of the invention does not limit the size of the preset current change rate threshold; it can be set by technicians based on experience or needs, or determined through numerous experiments. In response to control commands sent by the remote control handle, the target robotic arm is controlled to perform the demonstration training task. The control commands can be used to control the target robotic arm to perform the demonstration training task.

[0128] For example, if the object used for the demonstration training task is a target robotic arm, the sequence score of the candidate demonstration data sequence can be determined based on the following formula: ; in, This indicates a preset stable weight; This represents the contact stability score of the i-th candidate demonstration data sequence. It should be noted that here... + + =1.

[0129] The contact stability score can be used to quantify the contact stability of the target robotic arm under candidate demonstration data sequences. The preset stability weight can be used to quantify the importance of the contact stability score in the sequence score. This embodiment of the invention does not impose any limitation on the magnitude of the preset stability weight; it can be set by technicians based on experience or needs, or determined repeatedly through numerous experiments.

[0130] The sequence smoothness score can be calculated based on the normalized inverse of jerk (jerk acceleration) or other smoothness metrics. The semantic event alignment consistency score measures the degree of consistency between the event duration and the standard event duration of different semantic events in a candidate demonstration data sequence. A higher semantic event alignment consistency score indicates that the temporal rhythm of the corresponding candidate demonstration data sequence is more consistent with the standard rhythm. For example, calculating the time difference (or time warping distance) between the event duration and the corresponding standard event duration of different semantic events in a candidate demonstration data sequence typically uses a negative exponential function or normalized mapping to convert this time difference into a score between [0,1], thus obtaining the semantic event alignment consistency score. The contact stability score measures the force / torque variance or the number of contact exceedances in the contact segment. The sequence score can be used for sample weighted training, adaptive truncation ratios, and demonstration data sequence selection.

[0131] For example, the specific calculation methods for sequence smoothness score, contact stability score, and semantic event alignment consistency score may include Jerk's calculation window, alignment distance normalization method, and contact stability statistic.

[0132] For example, intra-segment resampling and mask padding (to reduce padding noise) are performed: Each semantic segment in the extracted reference demonstration data sequence is locally time-normalized and resampled to a fixed number of frames (segment parameters can be preset, such as proximity events: 16, contact events: 32, release events: 12); finally, the segments are concatenated and padded to a fixed batch length. The padded portion is marked with a mask and ignored in the loss function to avoid introducing gradient noise through zero padding. This approach differs structurally from simple zero padding and can work in conjunction with semantic alignment mechanisms to improve demonstration consistency and training efficiency.

[0133] For example, a latent semantic representation and cross-platform mapping adapter (i.e., a multimodal action mapping model) includes an encoder and an adapter. The encoder's expression can be: ; in, This represents the intermediate semantic vector at time ts, which is the input of the example. This represents the sample input pose at time ts. This represents the sample input vector at time ts. Indicates parameters The multimodal semantic coding network (encoder) takes pose sequences and visual local features as inputs and outputs a time series or latent semantic vector at time ts. That is, the encoder's input data includes the example input pose and example input vector at time ts, and the output data is the intermediate semantic vector at time ts, which is a dimensionless vector.

[0134] For example, the expression for the adapter can be: ; in, This represents the demonstration predicted pose at time ts+k. This indicates a target-oriented robotic arm with parameters. Lightweight adapter networks. The adapter is a lightweight network, such as an MLP (Multilayer Perceptron) or a Transformer (self-attention mechanism), which supports few-sample fine-tuning: only N calibration examples (suggested N belongs to [5,20]) are needed for rapid adaptation on the target robotic arm.

[0135] For example, the model loss value can be determined using the following formula: ; Where L represents the model loss value; Indicates the pose loss weights; This represents the pose loss value; Indicates the trajectory loss weights; This represents the trajectory loss value; Indicates semantic loss weights; This represents the semantic loss value.

[0136] For example, demonstration data sequence weights can be introduced to either increase or decrease the influence of data from a specific target demonstration data sequence on determining the corresponding sub-loss value. The sub-loss value can include pose loss, semantic loss, and trajectory loss. Demonstration data sequence weights can quantify the importance of the target demonstration data sequence in determining the sub-loss value. These weights can be calculated by combining baseline weights with sequence scores, and are used to increase the influence of high-quality target demonstration data sequences during training. The loss calculation should specify whether the filled region is ignored using a mask and provide the normalization calculation method.

[0137] For example, a feedback current-kinematic contact estimation method can be used to determine the current end-effector contact force. For example, the current joint torque can first be determined using the following formula: ; in, This represents the current joint torque of the target joint m at time t, which is the approximate value of the joint output torque estimated by the current and dynamics model (unit: Nm). The estimated value needs to take into account the compensation of friction and inertia terms; t can represent the current time. The torque constant of the motor at the target joint m (unit: Nm / A) is used to approximate the current as the motor output torque. This represents the current joint current of target joint m at time t, i.e., the motor feedback current of target joint m at time t (unit: A), which is usually the measured value for each target joint. The torque estimation compensation (unit: Nm) caused by non-ideal terms such as friction / backlash of the target joint m can be estimated through offline calibration or online observer; Represents the joint inertia matrix or inertia term of the target joint m (unit: (or corresponding dimensions). This represents the current joint acceleration of target joint m at time t; The product value represents the inertial component.

[0138] For example, the current end contact force can be determined using the following formula: ; in, This represents the current end contact force at time t, which is the estimated end environmental contact force (unit: N). In singular configurations or irreversible cases, pseudo-inverse should be used, and a pseudo-inverse fault tolerance threshold should be preset. represents the Jacobian matrix, used to map joint torque estimates to end-effector force estimates; T represents the transpose.

[0139] For example, to avoid misjudgment in feedback current-force mapping, a dual criterion is adopted: the current moment is determined to be the moment of object contact when the current end-effector contact force value is greater than a preset end-effector contact force threshold and the current visual contact probability is greater than a preset contact probability threshold. For instance, the L2 norm of the current end-effector contact force is used as the current end-effector contact force value. The current visual contact probability is dimensionless, with a specific value ranging from [0,1], serving as redundant evidence for contact determination.

[0140] For example, to avoid abrupt changes and ensure the stability of the target robotic arm's operation, smoothing weights (i.e., preset component weights) can be introduced. The reference predicted joint torque for the next moment is determined using the following formula: ; in, This represents the reference predicted joint torque at time t+1; This represents the preset component weights, with a value range of [0,1]. This represents the position control component at time t; This represents the impedance control component at time t. The reference predicted joint torque at the next time step is obtained by smoothing and interpolating the position control component and impedance control component at the current time step according to preset component weights.

[0141] The impedance control component is the compliant control quantity required by the robotic arm after physical contact occurs (i.e., after visual redundancy detection and current threshold triggering). Its determination relies on the impedance control law, which treats the robotic arm's end effector as a dynamic system composed of mass, springs, and damping. Specifically, it uses the end-effector contact force and desired impedance parameters (including stiffness and damping) to calculate the control command required to allow the robotic arm to produce a certain positional deviation to buffer the collision torque. The position control component is the default control quantity when the robotic arm is in free space without contact. It is typically determined using a classic position closed-loop controller (such as a PID controller).

[0142] The inventiveness and key protection points of the technical solutions provided in this invention include: Adaptive time-step truncation mechanism: the proposal and implementation of an adaptive truncation ratio (i.e., sequence truncation ratio) driven by sequence scoring, and its combined process with event anchor protection (a substantial improvement to the fixed truncation method). Event anchor protection, intra-segment resampling, and mask filling joint process: implementing local time normalization / resampling using semantic events as anchors, and employing m(t)(t)-aware filling to ensure batch consistency and avoid training noise. Vision-handle confidence weighted fusion mechanism: based on dynamically estimated covariance. , Perform weighted fusion to obtain robustness. And strategies for weight adjustment in task semantics. Latent semantic representation, cross-platform rapid adaptation scheme for configuration adapters: including encoders. With lightweight adapter Architecture, training / fine-tuning (few-shot) process, and loss design (demonstrating the use of quality weights in demonstration data sequences). Impedance switching mechanism for joint feedback current-kinematic contact estimation and visual judgment: contact recognition based on feedback current / kinetic estimation under conditions without external force sensors, using visual probability as redundant verification, combined with a smooth impedance switching strategy. Includes current-force estimation equations, threshold judgment logic, and smooth switching implementation. System integration for industrial handling and sorting tasks: a systematic combination of low-cost remote control handles, multimodal data acquisition, adaptive preprocessing, cross-platform adapter mapping, and impedance switching without external force sensors, and its specific implementation in industrial application scenarios.

[0143] For example, this embodiment takes handling and sorting tasks in an industrial manufacturing scenario as its application background. Addressing the shortcomings in existing robotic arm teleoperation demonstration data acquisition, such as insufficient multimodal alignment and semantic consistency, low-cost platforms lacking proprioceptive force perception but relying on contact judgment, rigidity of truncation strategies and risk of information loss, and sample efficiency issues in cross-platform mapping, this embodiment proposes a low-cost, cross-platform vision-teleoperation handle fusion multimodal demonstration acquisition and motion mapping system. Through semantic event-driven adaptive time-step truncation and intra-segment resampling mechanisms related to demonstration quality, it effectively removes tail-end redundancy and reduces demonstration noise. The system employs vision-handle confidence fusion, latent semantic encoding, and a lightweight adapter network to achieve cross-actuator motion redirection under limited sample conditions. Simultaneously, based on robotic arm current feedback—kinematic contact estimation and visual redundancy judgment, it completes task-specific impedance switching, enabling the low-cost platform to still possess reliable contact recognition and compliant control even without sensors. The overall solution combines the advantages of data-driven and modular approaches, improving demonstration utilization efficiency and training stability, and accelerating the deployment of heterogeneous robotic arms.

[0144] For example, a multimodal demonstration acquisition and semantic-driven data preprocessing method based on vision-teleoperation handle fusion is proposed. This method uses a teleoperation handle to acquire multiple candidate demonstration data sequences for handling and sorting tasks, and determines the demonstration end-effector pose and demonstration visual vector at the time of data acquisition based on these candidate demonstration data sequences. Specifically, the method includes the following steps: Step S1: The robotic arm system and the teleoperation system receive the same task instruction set; wherein, the task instruction set includes target object identification information, grasping position, handling path, and target placement area information. Step S2: The target robotic arm is controlled to perform task actions using teleoperation. The operator controls the target robotic arm through the teleoperation handle, enabling the target robotic arm to complete the grasping, moving, and placing actions, forming a complete candidate demonstration data sequence. Step S3: Multimodal demonstration data is acquired synchronously. During the teleoperation process, the following data are acquired synchronously: target robotic arm end-effector pose data, target robotic arm joint position and joint velocity data, target robotic arm drive motor current data, and visual image data of the working environment; and the above data are uniformly timestamped to form a synchronous multimodal demonstration sequence. Step S4: Perform semantic event detection, analyze the candidate demonstration data sequence, and identify semantic event nodes in the task process, including key stages such as approaching the target, initial contact, stable contact, and releasing the target. Step S5: Perform semantic-driven adaptive time step truncation. Based on the overall quality and semantic event distribution of the candidate demonstration data sequence, adaptively truncate redundant time steps at the end of the candidate demonstration data sequence, while retaining the necessary task completion time period after the last semantic event to avoid accidental deletion of key action information. Step S6: Perform segmented resampling and time alignment. Using semantic events as segment boundaries, perform time normalization and resampling on the event data within each semantic segment, ensuring that different candidate demonstration data sequences have a consistent data length structure at the same semantic stage. Step S7: Perform padding and masking. Pad data segments that are insufficient in length and set masking marks on the padded parts to ignore the corresponding data during subsequent training, thereby avoiding interference with the model learning process. Step S8: Construct a structured demonstration dataset. Store the processed multimodal demonstration data in a unified data structure for subsequent model training, task learning, and action mapping.

[0145] For example, a cross-platform motion mapping method based on action semantic representation and a lightweight adaptation network is proposed, including the following steps: establishing a target robotic arm adaptation network model (i.e., a multimodal motion mapping model). Specifically, for different models or structures of target robotic arms, a corresponding lightweight motion adaptation network model is constructed to realize the mapping relationship between action semantic representation and target robotic arm control commands; performing few-shot adaptation training. Specifically, a small amount of demonstration data is collected on the target robotic arm platform, and only the parameters of the adaptation network module are trained and adjusted, while the parameters of the action semantic encoding module remain unchanged; generating demonstration predicted poses. Specifically, the action semantic representation is input into the adaptation network module, and the demonstration predicted poses corresponding to the target robotic arm are output.

[0146] An exemplary contact estimation and task-oriented impedance switching method based on current-kinematic fusion includes the following steps: First, collecting target robotic arm operating state data, specifically, real-time collection of robotic arm joint angle data, joint acceleration data, and drive motor current data; second, constructing an end-effector contact state estimation model, specifically, estimating and judging the external contact state of the robotic arm end based on the robotic arm kinematic model and motor current change characteristics; third, fusing visual information for redundancy judgment, specifically, combining the recognition results of the contact state between the robotic arm end and the target object by the visual perception unit to perform multi-source information fusion judgment, improving the reliability of contact recognition; fourth, triggering control mode switching, specifically, when the system determines that a contact event has occurred, automatically switching the robotic arm control mode from position control to impedance control mode; fifth, performing smooth transition control, specifically, during the control mode switching process, generating continuous control commands through smooth interpolation to avoid system instability caused by abrupt changes in control commands; and sixth, restoring the control mode, specifically, when the contact ends and the restoration conditions are met, restoring the robotic arm control mode to the original control mode.

[0147] This invention provides a low-cost, cross-platform, and scalable industrial robotic arm handling and sorting task execution solution through the systematic integration of remote operation demonstration acquisition, multimodal data processing, motion semantic mapping, and contact safety control. It is suitable for the rapid deployment and engineering application of various types of robotic arm platforms.

[0148] Example 5 Figure 4 This is a schematic diagram of a training device for a multimodal motion mapping model provided in Embodiment 5 of the present invention. This embodiment is applicable to the training of a multimodal motion mapping model for a robotic arm based on multimodal demonstration data. The method can be executed by a training device for the multimodal motion mapping model, which can be implemented in software and / or hardware and can be configured in an electronic device that carries the training function of the multimodal motion mapping model.

[0149] like Figure 4 As shown, the device includes: a demonstration data acquisition module 410, a sequence scoring determination module 420, a sequence truncation time determination module 430, a target demonstration data sequence determination module 440, and a multimodal action mapping model training module 450. Among them, The demonstration data acquisition module 410 is used to acquire candidate demonstration data sequences of the target handle corresponding to the target robotic arm under the demonstration training task, and determine the corresponding global duration of the sequence and event detection data based on the candidate demonstration data sequences; wherein, the candidate demonstration data sequences include demonstration end pose and demonstration visual vector; The sequence scoring determination module 420 is used to determine the semantic event corresponding to the corresponding candidate demonstration data sequence based on the event detection data, and to determine the sequence score and end event start time corresponding to the corresponding candidate demonstration data sequence based on the candidate demonstration data sequence and the semantic event. The sequence truncation time determination module 430 is used to determine the sequence truncation time of the corresponding candidate demonstration data sequence based on the global duration of the sequence, the sequence score, the start time of the terminal event, the preset truncation ratio threshold, and the preset retention time. The target demonstration data sequence determination module 440 is used to truncate the candidate demonstration data sequence according to the sequence truncation time to obtain a reference demonstration data sequence, and to post-process the event data under different semantic events in the reference demonstration data sequence to obtain the target demonstration data sequence; wherein, the post-processing includes intra-segment resampling and mask padding; The multimodal action mapping model training module 450 is used to train the constructed multimodal action mapping model based on the target demonstration data sequence to obtain the trained multimodal action mapping model.

[0150] This invention provides a training scheme for a multimodal motion mapping model. It involves acquiring candidate demonstration data sequences of the target handle corresponding to the target robotic arm under a demonstration training task, and determining the corresponding global sequence duration and event detection data based on these sequences. The candidate demonstration data sequences include the demonstration end-effector pose and the demonstration visual vector. Semantic events corresponding to the candidate demonstration data sequences are determined based on the event detection data, and a sequence score and end-effector event start time are determined based on the candidate demonstration data sequences and semantic events. The sequence truncation time of the corresponding candidate demonstration data sequences is determined based on the global sequence duration, sequence score, end-effector event start time, a preset truncation ratio threshold, and a preset retention time. The candidate demonstration data sequences are truncated according to the truncation time to obtain a reference demonstration data sequence. Event data under different semantic events in the reference demonstration data sequence are post-processed to obtain the target demonstration data sequence. The post-processing includes intra-segment resampling and mask filling. The constructed multimodal motion mapping model is trained based on the target demonstration data sequence to obtain the trained multimodal motion mapping model. The above scheme determines the truncation time of the corresponding candidate demonstration data sequence based on the truncation ratio threshold, preset retention time, sequence score, global sequence duration, and end event start time of the candidate demonstration data sequence. Then, it adaptively truncates the corresponding candidate demonstration data sequence according to the truncation time to obtain the reference demonstration data sequence. Finally, it performs intra-segment resampling and mask filling on the reference demonstration data sequence to obtain the target demonstration data sequence. This improves the accuracy of the determined target demonstration data sequence, thereby improving the accuracy of training the multimodal motion mapping model based on the target demonstration data sequence. This improves the performance of the trained multimodal motion mapping model, and in turn, improves the accuracy of subsequent control of the robotic arm based on the joint angles corresponding to the predicted pose determined by the trained multimodal motion mapping model.

[0151] Optionally, the sequence scoring determination module 420 includes: The evaluation data determination unit is used to determine the evaluation data and the start time of the terminal event corresponding to the candidate demonstration data sequence based on the candidate demonstration data sequence and the semantic event; The sequence scoring determination unit is used to determine the sequence score corresponding to the corresponding candidate demonstration data sequence based on the evaluation data.

[0152] Optionally, the sequence truncation time determination module 430 includes: The sequence truncation ratio determination unit is used to determine the sequence truncation ratio of any candidate demonstration data sequence based on the sequence score of the candidate demonstration data sequence and the truncation ratio threshold. The sequence truncation time determination unit is used to determine the sequence truncation time of the candidate demonstration data sequence based on the sequence truncation ratio, the global duration of the sequence, the preset retention duration, and the start time of the end event.

[0153] Optionally, the demonstration end-effector pose is determined based on the following device: The end-effector pose determination unit is used to obtain the visual end-effector pose and inertial end-effector pose of the target handle at any given data acquisition time. The demonstration end-effector pose determination unit is used to determine the demonstration end-effector pose at the time of data acquisition based on the visual end-effector pose, inertial end-effector pose, preset visual weight matrix, and preset inertial weight matrix at the time of data acquisition.

[0154] Optionally, the multimodal action mapping model training module 450 includes: The demonstration input data determination unit is used to determine the demonstration input pose and demonstration label pose from the target end pose of the target demonstration data sequence, and to determine the demonstration input vector from the target visual vector; The demonstration prediction pose determination unit is used to input the demonstration input pose and the demonstration input vector into the constructed multimodal motion mapping model to obtain the demonstration prediction pose; The model loss value determination unit is used to determine the model loss value based on the demonstration predicted pose, the demonstration label pose, the demonstration input pose, the target demonstration data sequence, the intermediate semantic vector of the multimodal action mapping model, and the preset standard semantic vector, and to train the constructed multimodal action mapping model based on the model loss value to obtain the trained multimodal action mapping model.

[0155] Optionally, the model loss value determination unit is specifically used for: The pose loss value is determined based on the predicted pose and the labeled pose. The predicted demonstration trajectory is determined based on the predicted demonstration pose and the predicted demonstration input pose, and the trajectory loss value is determined based on the predicted demonstration trajectory and the target demonstration trajectory corresponding to the target demonstration data sequence. The semantic loss value is determined based on the intermediate semantic vector and the standard semantic vector; wherein the intermediate semantic vector is the latent semantic representation output by the encoder in the multimodal action mapping model; The model loss value is determined based on the pose loss value, the trajectory loss value, and the semantic loss value.

[0156] The training apparatus for the multimodal action mapping model provided in this embodiment of the invention can execute the training method of the multimodal action mapping model provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects for executing the training method of each multimodal action mapping model.

[0157] In the technical solution of this invention, the collection, storage, use, processing, transmission, provision and disclosure of candidate demonstration data sequences, etc., all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0158] Example 6 Figure 5 This is a schematic diagram of a robotic arm control device provided in Embodiment Six of the present invention. This embodiment is applicable to situations where the operation of a robotic arm is controlled. The method can be executed by a robotic arm control device, which can be implemented in software and / or hardware and can be configured in an electronic device that carries the robotic arm control function.

[0159] like Figure 5 As shown, the device includes: a current data acquisition module 510, a current visual contact probability determination module 520, a current time determination module 530, a current visual vector determination module 540, a predicted joint angle determination module 550, and a robotic arm control module 560. Among them, The current data acquisition module 510 is used to acquire the current joint data and current visual image corresponding to each target joint in the target robotic arm at the current moment; wherein, the current joint data includes the current joint angle, the current joint current and the current joint acceleration; The current visual contact probability determination module 520 is used to determine the current end contact force of the target robotic arm based on the current joint current and the current joint acceleration, and to determine the current visual contact probability based on the current visual image. The current moment determination module 530 is used to determine whether the current moment is the moment of object contact based on the current end contact force and the current visual contact probability. The current visual vector determination module 540 is used to determine the current end pose of the target robotic arm based on the current joint angle and determine the current visual vector based on the current visual image if no, when the current time is the model prediction time. The joint angle prediction module 550 is used to input the current end-effector pose and the current visual vector into a trained multimodal motion mapping model corresponding to the target robotic arm type, to obtain the candidate prediction pose corresponding to the candidate prediction time within a preset prediction period, and to determine the predicted joint angle at the corresponding candidate prediction time based on the candidate prediction pose; wherein, the multimodal motion mapping model is trained using a multimodal motion mapping model training method. The robotic arm control module 560 is used to control the operation of the target robotic arm at the corresponding candidate prediction time according to the predicted joint angle.

[0160] This invention provides a robotic arm control scheme. It acquires current joint data and a current visual image corresponding to each target joint in the target robotic arm at the current moment. The current joint data includes the current joint angle, current joint current, and current joint acceleration. Based on the current joint current and acceleration, the current end-effector contact force of the target robotic arm is determined, and based on the current visual image, the current visual contact probability is determined. Based on the current end-effector contact force and the current visual contact probability, it is determined whether the current moment is an object contact moment. If not, when the current moment is a model prediction moment, the current end-effector pose of the target robotic arm is determined based on the current joint angle, and the current visual vector is determined based on the current visual image. The current end-effector pose and the current visual vector are input to a trained multimodal motion mapping model corresponding to the target robotic arm type to obtain candidate prediction poses corresponding to candidate prediction moments within a preset prediction time period. Based on the candidate prediction poses, the predicted joint angles at the corresponding candidate prediction moments are determined. The multimodal motion mapping model is trained using a multimodal motion mapping model training method. Based on the predicted joint angles, the operation of the target robotic arm at the corresponding candidate prediction moments is controlled. The above scheme determines whether the current moment is an object contact moment based on the current end-effector contact force and the current visual contact probability. If not, when the current moment is a model prediction moment, the current end-effector pose and the current visual vector are input into the trained multimodal motion mapping model to output candidate predicted poses. Based on the candidate predicted poses, predicted joint angles are determined. Finally, the target robotic arm is controlled to run at the corresponding candidate prediction moment based on the predicted joint angles. This improves the accuracy of the candidate predicted poses determined by the trained multimodal motion mapping model, and improves the accuracy of controlling the target robotic arm based on the predicted joint angles determined by the candidate predicted poses. In other words, it improves the accuracy of controlling the robotic arm based on the joint angles corresponding to the predicted poses determined by the trained multimodal motion mapping model.

[0161] Optionally, the current time determination module 530 includes: The contact force value determination unit is used to determine the current end contact force value based on the current end contact force. The object contact time determination unit is used to determine the current time as the object contact time if the current end contact force value is greater than a preset end contact force threshold and the current visual contact probability is greater than a preset contact probability threshold.

[0162] The robotic arm control device provided in the embodiments of the present invention can execute the robotic arm control method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing each robotic arm control method.

[0163] In the technical solution of this invention, the collection, storage, use, processing, transmission, provision and disclosure of current joint data and current visual images, etc., all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0164] Example 7 Figure 6 This is a schematic diagram of an electronic device for implementing a training method for a multimodal motion mapping model or a robotic arm control method, as provided in Embodiment 7 of the present invention. The electronic device 10 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0165] like Figure 6 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0166] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0167] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as training methods for multimodal motion mapping models or robotic arm control methods.

[0168] In some embodiments, the training method for the multimodal motion mapping model or the robotic arm control method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the training method for the multimodal motion mapping model or the robotic arm control method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the training method for the multimodal motion mapping model or the robotic arm control method by any other suitable means (e.g., by means of firmware).

[0169] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0170] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0171] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0172] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0173] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0174] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0175] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0176] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A training method for a multimodal action mapping model, characterized in that, include: Obtain candidate demonstration data sequences of the target handle corresponding to the target robotic arm under the demonstration training task, and determine the corresponding global duration of the sequence and event detection data based on the candidate demonstration data sequences; wherein, the candidate demonstration data sequences include the demonstration end pose and the demonstration visual vector; Based on the event detection data, determine the semantic events corresponding to the corresponding candidate demonstration data sequences, and based on the candidate demonstration data sequences and the semantic events, determine the sequence score and the start time of the terminal event corresponding to the corresponding candidate demonstration data sequences. The sequence truncation time of the corresponding candidate demonstration data sequence is determined based on the global duration of the sequence, the sequence score, the start time of the terminal event, the preset truncation ratio threshold, and the preset retention duration. Based on the sequence truncation time, the candidate demonstration data sequence is truncated to obtain a reference demonstration data sequence, and the event data under different semantic events in the reference demonstration data sequence are post-processed to obtain the target demonstration data sequence; wherein, the post-processing includes intra-segment resampling and mask padding; The constructed multimodal action mapping model is trained based on the target demonstration data sequence to obtain the trained multimodal action mapping model.

2. The method according to claim 1, characterized in that, The step of determining the sequence score and terminal event start time corresponding to the candidate demonstration data sequence based on the candidate demonstration data sequence and the semantic event includes: Based on the candidate demonstration data sequence and the semantic event, determine the evaluation data and the start time of the terminal event corresponding to the respective candidate demonstration data sequence; Based on the evaluation data, the sequence score corresponding to the candidate demonstration data sequence is determined.

3. The method according to claim 1, characterized in that, The step of determining the sequence truncation time of the corresponding candidate demonstration data sequence based on the global duration of the sequence, the sequence score, the start time of the terminal event, a preset truncation ratio threshold, and a preset retention duration includes: For any candidate demonstration data sequence, the sequence truncation ratio of the candidate demonstration data sequence is determined based on the sequence score of the candidate demonstration data sequence and the truncation ratio threshold. The sequence truncation time of the candidate demonstration data sequence is determined based on the sequence truncation ratio, the global duration of the sequence, the preset retention duration, and the start time of the terminal event.

4. The method according to claim 1, characterized in that, The demonstration end-effector pose is determined based on the following method: For any given data acquisition moment, acquire the visual end-effector pose and inertial end-effector pose of the target handle at that data acquisition moment; Based on the visual end-effector pose, inertial end-effector pose, preset visual weight matrix, and preset inertial weight matrix at the time of data acquisition, the demonstration end-effector pose at the time of data acquisition is determined.

5. The method according to claim 1, characterized in that, The step of training the constructed multimodal action mapping model based on the target demonstration data sequence to obtain the trained multimodal action mapping model includes: The demonstration input pose and demonstration label pose are determined from the target end pose of the target demonstration data sequence, and the demonstration input vector is determined from the target visual vector; The demonstration input pose and the demonstration input vector are input into the constructed multimodal motion mapping model to obtain the demonstration predicted pose; Based on the demonstration predicted pose, the demonstration label pose, the demonstration input pose, the target demonstration data sequence, the intermediate semantic vector of the multimodal action mapping model, and the preset standard semantic vector, the model loss value is determined, and the constructed multimodal action mapping model is trained based on the model loss value to obtain the trained multimodal action mapping model.

6. The method according to claim 5, characterized in that, The step of determining the model loss value based on the demonstrated predicted pose, the demonstrated labeled pose, the demonstrated input pose, the target demonstrated data sequence, the intermediate semantic vector of the multimodal action mapping model, and the preset standard semantic vector includes: The pose loss value is determined based on the predicted pose and the labeled pose. The predicted demonstration trajectory is determined based on the predicted demonstration pose and the predicted demonstration input pose, and the trajectory loss value is determined based on the predicted demonstration trajectory and the target demonstration trajectory corresponding to the target demonstration data sequence. The semantic loss value is determined based on the intermediate semantic vector and the standard semantic vector; wherein the intermediate semantic vector is the latent semantic representation output by the encoder in the multimodal action mapping model; The model loss value is determined based on the pose loss value, the trajectory loss value, and the semantic loss value.

7. A robotic arm control method, characterized in that, include: Acquire the current joint data and current visual image corresponding to each target joint in the target robotic arm at the current moment; wherein, the current joint data includes the current joint angle, the current joint current, and the current joint acceleration; Based on the current joint current and the current joint acceleration, the current end-effector contact force of the target robotic arm is determined, and based on the current visual image, the current visual contact probability is determined. Based on the current end contact force and the current visual contact probability, determine whether the current moment is the moment of object contact; If not, then when the current time is the model prediction time, determine the current end-effector pose of the target robotic arm based on the current joint angle, and determine the current visual vector based on the current visual image; The current end-effector pose and the current visual vector are input into a trained multimodal motion mapping model corresponding to the target robotic arm type to obtain candidate predicted poses corresponding to candidate prediction times within a preset prediction period. Based on the candidate predicted poses, the predicted joint angles at the corresponding candidate prediction times are determined. The multimodal motion mapping model is trained using the training method described in any one of claims 1-6. Based on the predicted joint angle, the operation of the target robotic arm is controlled at the corresponding candidate predicted time.

8. The method according to claim 7, characterized in that, The step of determining whether the current moment is a moment of object contact based on the current end contact force and the current visual contact probability includes: Determine the value of the current end contact force based on the current end contact force; If the current end contact force value is greater than the preset end contact force threshold, and the current visual contact probability is greater than the preset contact probability threshold, then the current moment is determined to be the moment of object contact.

9. A training device for a multimodal action mapping model, characterized in that, include: The demonstration data acquisition module is used to acquire candidate demonstration data sequences of the target handle corresponding to the target robotic arm under the demonstration training task, and determine the corresponding global duration of the sequence and event detection data based on the candidate demonstration data sequences; wherein, the candidate demonstration data sequences include the demonstration end pose and the demonstration visual vector; The sequence scoring determination module is used to determine the semantic event corresponding to the corresponding candidate demonstration data sequence based on the event detection data, and to determine the sequence score and end event start time corresponding to the corresponding candidate demonstration data sequence based on the candidate demonstration data sequence and the semantic event. The sequence truncation time determination module is used to determine the sequence truncation time of the corresponding candidate demonstration data sequence based on the global duration of the sequence, the sequence score, the start time of the terminal event, a preset truncation ratio threshold, and a preset retention time. The target demonstration data sequence determination module is used to truncate the candidate demonstration data sequence according to the sequence truncation time to obtain a reference demonstration data sequence, and to post-process the event data under different semantic events in the reference demonstration data sequence to obtain the target demonstration data sequence; wherein, the post-processing includes intra-segment resampling and mask padding; The multimodal action mapping model training module is used to train the constructed multimodal action mapping model based on the target demonstration data sequence to obtain the trained multimodal action mapping model.

10. A robotic arm control device, characterized in that, include: The current data acquisition module is used to acquire the current joint data and current visual image corresponding to each target joint in the target robotic arm at the current moment; wherein, the current joint data includes the current joint angle, the current joint current and the current joint acceleration; The current visual contact probability determination module is used to determine the current end contact force of the target robotic arm based on the current joint current and the current joint acceleration, and to determine the current visual contact probability based on the current visual image. The current moment determination module is used to determine whether the current moment is the moment of object contact based on the current end contact force and the current visual contact probability. The current visual vector determination module is used to determine the current end-effector pose of the target robotic arm based on the current joint angle, and to determine the current visual vector based on the current visual image, if no, when the current time is the model prediction time. The joint angle prediction module is used to input the current end-effector pose and the current visual vector into a trained multimodal motion mapping model corresponding to the target robotic arm type, to obtain the candidate prediction pose corresponding to the candidate prediction time within a preset prediction period, and to determine the predicted joint angle at the corresponding candidate prediction time based on the candidate prediction pose; wherein, the multimodal motion mapping model is trained using the training method described in any one of claims 1-6. The robotic arm control module is used to control the operation of the target robotic arm at the corresponding candidate prediction time according to the predicted joint angle.