Human production scene labor strategy decoding method and device for humanoid robot and medium

By acquiring electromyographic signals in real time and using serial-parallel encoders and diffusion models to decode hand movement states, the problem of decoding complex and dexterous hand movements and interactive object information was solved, achieving high-precision decoding of labor strategies.

CN121105016APending Publication Date: 2025-12-12TONGJI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511372092.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-24
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

Existing technologies struggle to decode complex and dexterous hand movements and interactive object information in human production scenarios, and traditional recognition models lack robustness.

Method used

By acquiring electromyographic signals from the upper limbs in real time, a semi-structured latent space representation is constructed using a series-parallel encoder module and a diffusion model to decode hand movement states and infer information about interactive objects.

Benefits of technology

It improves the accuracy and robustness of hand motion tracking, effectively decodes complex and dexterous hand movements and interactive object information, and enhances the decoding accuracy of human productive labor strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121105016A_ABST
    Figure CN121105016A_ABST
Patent Text Reader

Abstract

The invention relates to a human production scene labor strategy decoding method and device for a humanoid robot and a medium. The human production scene labor strategy decoding method comprises the steps that electromyographic signals of upper limb parts of a human body are collected in real time; an encoder module based on a series-parallel mixed structure is adopted to encode the collected electromyographic signals to obtain hidden space representation of the electromyographic signals, and meanwhile semi-structured hidden space representation representing the hand motion state is constructed; inputting the implicit space representation of the myoelectricity data into the diffusion model as a condition vector, and analyzing to obtain a semi-structured hand motion state implicit space representation; decoding from the semi-structured hand motion state hidden space representation to obtain real-time angle data of each joint of the hand, and taking the real-time angle data as a finger motion tracking result; a human action strategy is decoded from the finger motion tracking result, and interaction object information is deduced according to the human action strategy; and integrating the finger motion tracking result and the interaction object information to obtain a complete hand labor strategy. Compared with the prior art, the method has the advantages of accuracy, reliability and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robot control, and in particular to a method, device and medium for decoding labor strategies in human production scenarios. Background Technology

[0002] Human labor strategies exhibit high complexity, adaptability, and versatility across various scenarios, making the decoding of these strategies crucial for enhancing the capabilities of humanoid robots. Current skill learning methods based on optical devices are limited in their application in real-world production environments due to the influence of occlusion, lighting conditions, and the material of the work object.

[0003] While electromyography (EMG)-based methods can capture rich muscle movement information, the nonlinearity and non-stationarity of EMG signals, coupled with the complexity of hand movements, pose significant challenges to the accurate decoding of hand movements. Furthermore, there is still no research on how to simultaneously decode interactive object information from hand movement strategies. Summary of the Invention

[0004] The purpose of this invention is to overcome the shortcomings of the prior art by providing a method, device and medium for decoding labor strategies in human production scenarios for humanoid robots.

[0005] The objective of this invention can be achieved through the following technical solutions:

[0006] According to a first aspect of the present invention, a method for decoding labor strategies in human-like production scenarios for humanoid robots is provided, the method comprising:

[0007] Real-time acquisition of electromyographic signals from the upper limbs of the human body;

[0008] The acquired electromyographic signals are encoded in a series-parallel manner to obtain the latent space representation of the electromyographic signals. Based on the latent space representation of the electromyographic signals, a semi-structured latent space representation representing the hand movement state is constructed.

[0009] The latent space representation of electromyography data is used as a conditional vector input into the diffusion model to estimate a semi-structured latent space representation of hand movement state.

[0010] Real-time angle data of each joint of the hand are obtained by decoding from the semi-structured latent space representation of hand motion state, and then the finger motion tracking results are obtained.

[0011] Decode human action strategies from finger motion tracking results, and infer information about interactive objects based on human action strategies;

[0012] By integrating finger motion tracking results with interactive object information, a complete hand labor strategy can be obtained.

[0013] Preferably, an EMG sensor is used to collect electromyographic signals from the human forearm or wrist in real time.

[0014] Preferably, the step of performing serial-parallel encoding on the acquired electromyographic signals to obtain the latent space representation of the electromyographic signals specifically includes:

[0015] An encoder module based on a series-parallel hybrid structure is constructed, including a parallel encoder group, a series encoder, and a series decoder arranged sequentially. Each encoder in the parallel encoder group corresponds one-to-one with each EMG channel of the electromyography signal.

[0016] Preferably, when an EMG sensor fails, the encoding results of the corresponding EMG channel are all set to zero.

[0017] Preferably, the semi-structured latent space representation characterizing the hand movement state includes a pose matrix with physical meaning and a latent vector without physical meaning.

[0018] Preferably, the latent space representation of electromyography data is used as a conditional vector input based on a diffusion model to estimate a semi-structured latent space representation of hand motion state. In estimating the latent space representation of hand motion state, three parallel computations are performed: estimating only the pose matrix with physical meaning, estimating only the latent vector without physical meaning, and estimating the complete latent space representation of hand motion state. The average of the three parallel computation results is then output as the final semi-structured latent space representation of hand motion state.

[0019] Preferably, the step of decoding the real-time angle data of each joint of the hand from the semi-structured hand motion state latent space representation to obtain the finger motion tracking result specifically includes: inputting the semi-structured hand motion state latent space representation into the decoder, decoding to obtain the real-time angle data of each joint, averaging the decoded real-time angle data of each joint with the joint angle represented by the pose matrix as the stable anchor point, and outputting the result to obtain the finger motion tracking information.

[0020] Preferably, the human action strategy is decoded from the finger motion tracking results, and the interaction object information is inferred based on the human action strategy. When inferring the interaction object information, the finger motion tracking information is reconstructed, and the hidden vector part without physical meaning in the latent space representation of the hand motion state is decoded simultaneously with the angles of each finger joint.

[0021] According to a second aspect of the present invention, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the program to implement any of the methods described above.

[0022] According to a third aspect of the invention, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements any of the methods described herein.

[0023] Compared with the prior art, the present invention has the following beneficial effects:

[0024] (1) This invention solves the problem that traditional recognition models are not good at representing complex and dexterous finger movements and cannot estimate the information of interactive objects by using semi-structured hand state and serial hybrid encoding and decoding and labor strategy decoding based on diffusion model.

[0025] (2) Installing an EMG sensor on the forearm or wrist can effectively avoid interference with human labor due to the installation location of the EMG sensor.

[0026] (3) By using the encoder module based on the series-parallel hybrid structure, the characteristic that each encoder in the parallel encoder group corresponds one-to-one with each EMG channel of the electromyography signal is utilized. When an EMG sensor fails, the encoding result of the corresponding EMG channel is set to zero to maintain normal operation, which has higher robustness.

[0027] (4) The electromyographic signal is encoded by an encoder module based on a series-parallel hybrid structure to obtain its latent space representation. At the same time, combined with a denoising module based on transformer, the long-range dependency and sparsity problems in the electromyographic signal are effectively solved, and the accuracy of hand motion tracking is improved.

[0028] (5) When estimating the latent space representation of hand motion state, this invention performs parallel calculations of only the pose matrix with physical meaning, the latent vector without physical meaning, and the complete latent space representation of hand motion state. The average of the three parallel calculation results is output as the final semi-structured latent space representation of hand motion state, which can effectively improve the stability of the estimation results.

[0029] (6) The finger motion tracking information is more accurate when the real-time angle data of each joint obtained by decoding is averaged with the joint angle represented by the pose matrix as the stable anchor point. Attached Figure Description

[0030] Figure 1 This is a flowchart of the method of the present invention.

[0031] Figure 2 This is a schematic diagram of the human labor strategy decoding process in the embodiment. Detailed Implementation

[0032] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0033] Example

[0034] like Figure 1 and Figure 2 As shown, this embodiment provides a method for decoding labor strategies in human production scenarios, the method including:

[0035] S1. Real-time acquisition of electromyographic signals from the upper limbs of the human body;

[0036] S2. The acquired electromyographic signals are encoded in a series-parallel manner to obtain the latent space representation of the electromyographic signals. Based on the latent space representation of the electromyographic signals, a semi-structured latent space representation representing the hand movement state is constructed.

[0037] S3. Input the latent space representation of electromyography data as a conditional vector into the diffusion model to estimate the semi-structured latent space representation of hand movement state;

[0038] S4. Decode the real-time angle data of each joint of the hand from the semi-structured latent space representation of the hand motion state, and then obtain the finger motion tracking results.

[0039] S5. Decode the human action strategy from the finger motion tracking results, and infer the interaction object information based on the human action strategy;

[0040] S6. Integrate finger motion tracking results with interactive object information to obtain a complete hand labor strategy.

[0041] The method of this embodiment will now be described in detail.

[0042] S1. Real-time acquisition of electromyographic signals from the upper limbs of the human body.

[0043] In this embodiment, EMG sensors can be installed only on the forearm or wrist, effectively avoiding interference with human labor due to the installation location of the EMG sensors. Eight sensors are preferred, but six are also acceptable.

[0044] To address the instability of EMG sensors in practical applications and reduce usage costs, the technology can operate even with a limited number of EMG sensor electrodes, and can also eliminate the filtering step in the EMG acquisition process to reduce costs.

[0045] The model training phase employs a large-scale data pre-training method, which can significantly reduce the number of EMG sensor electrodes required and can be used normally even when the number of EMG electrodes varies or when the number of EMG electrodes is reduced due to poor contact.

[0046] S2. The acquired electromyographic signals are encoded in a series-parallel manner to obtain the latent space representation of the electromyographic signals. Based on the latent space representation of the electromyographic signals, a semi-structured latent space representation representing the hand movement state is constructed.

[0047] In this embodiment, an encoder module based on a hybrid serial-parallel structure is constructed, including a parallel encoder group, a serial encoder, and a serial decoder arranged sequentially. Each encoder in the parallel encoder group corresponds one-to-one with each EMG channel of the electromyography (EMG) signal. The serial encoder encodes the encoding results of multiple parallel encoders, and the encoding results of the serial encoder are input to the serial decoder for decoding. The serial decoder further performs channel alignment, dimensionality compression, and enhanced fault tolerance to obtain the latent space representation of the EMG signal. In this embodiment, each individual encoder adopts an encoder based on a VAE (Variational Autoencoder) structure.

[0048] In this embodiment, an encoder is built for each EMG channel, and the encoders of multiple channels are connected in parallel. When an EMG sensor fails, the corresponding encoding results are all set to zero to maintain the normal operation of the entire system.

[0049] In this embodiment, the semi-structured latent space representation of hand movement states includes a pose matrix with physical meaning and a latent space representation of electromyographic signals without physical meaning.

[0050] S3. Input the latent space representation of electromyography data as a conditional vector into the diffusion model to estimate the semi-structured latent space representation of hand movement state.

[0051] S301. In this embodiment, a denoising module with transformer as its backbone is used to process the latent space representation of electromyography data, which solves the long-range dependency and sparsity problems in EMG signals and avoids the impact of EMG data quality on the final estimation effect.

[0052] S302. The latent space representation of the denoised electromyography data is used as a conditional vector input to the diffusion model. Three parallel calculations are performed, and the average value output is a semi-structured latent space representation of hand movement state.

[0053] In this embodiment, when estimating the latent space representation of hand motion state, three parallel computations are performed. The three computation results are obtained by decoding with a decoder with three outputs: estimating only the pose matrix with physical meaning, estimating the latent vector without physical meaning, and estimating the complete latent space representation of hand motion state. The three parallel computation results are averaged and output as the final semi-structured latent space representation of hand motion state, thereby improving the stability of the estimation results.

[0054] S4. Decode the real-time angle data of each joint of the hand from the semi-structured latent space representation of the hand motion state, and then obtain the finger motion tracking result. In this embodiment, the decoder uses a typical MLP structure to improve decoding efficiency.

[0055] In this embodiment, during the decoding process, the pose matrix, which has physical meaning, is not directly used for joint angle decoding, but rather as a stabilizing anchor point. It should be noted that while the pose matrix already has a clear physical meaning, it often has significant errors and therefore does not require further decoding. However, it can be used to provide better initialization for the decoder, thereby improving the decoding effect.

[0056] The latent space representation of the hand movement state is input into the decoder to obtain the real-time angle data of each joint. The real-time angle data of each joint obtained by decoding is averaged with the joint angle represented by the pose matrix used as the stable anchor point and then output to obtain the finger movement tracking information.

[0057] S5. Decode the human action strategy from the finger motion tracking results, and infer the interaction object information based on the human action strategy.

[0058] In this embodiment, a decoder using an MLP structure is used to decode human action strategies from finger motion tracking information and infer interaction object information. When inferring interaction object information, the finger motion tracking information is reconstructed, and the hidden vector part of the hand motion state latent space representation without physical meaning and the angles of each finger joint are decoded simultaneously.

[0059] This invention addresses the shortcomings of traditional recognition models in representing complex and dexterous finger movements and estimating information about interacting objects by employing a semi-structured hand state hybrid encoding and decoding method and a human labor strategy decoding setting based on a diffusion model.

[0060] S6. Integrate finger motion tracking results with interactive object information to obtain a complete hand labor strategy.

[0061] To verify the performance of this invention, experiments were conducted on five adult individuals, and the results were compared with state-of-the-art methods and algorithms reported in international journals. Under the same experimental conditions, the popular VGG16 network and VAE network exhibited angular errors of 9.102 degrees and 7.496 degrees respectively in finger motion tracking. The method of this invention achieved an angular error of 1.258 degrees, significantly improving the accuracy of finger motion tracking, and reported an object recognition accuracy of 95.3%, thereby effectively improving the accuracy of decoding human productive labor strategies.

[0062] The electronic device of this invention includes a central processing unit (CPU), which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) or loaded from a storage unit into random access memory (RAM). The RAM may also store various programs and data required for device operation. The CPU, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.

[0063] Multiple components in the device are connected to the I / O interface, including: input units such as keyboards and mice; output units such as various types of displays and speakers; storage units such as disks and optical discs; and communication units such as network interface cards (NICs), modems, and wireless transceivers. The communication unit allows the device to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0064] The processing unit executes the various methods and processes described above, such as methods S1 to S6. For example, in some embodiments, methods S1 to S6 may be implemented as computer software programs tangibly contained in a machine-readable medium, such as a storage unit. In some embodiments, part or all of the computer program may be loaded and / or installed on the device via ROM and / or a communication unit. When the computer program is loaded into RAM and executed by the CPU, one or more steps of methods S1 to S6 described above may be performed. Alternatively, in other embodiments, the CPU may be configured to execute methods S1 to S6 by any other suitable means (e.g., by means of firmware).

[0065] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload programmable logic devices (CPLDs), and so on.

[0066] The program code used to implement the methods of the present invention can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a standalone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0067] In the context of this invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0068] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for decoding labor strategies in human-like production scenarios for humanoid robots, characterized in that, include: Real-time acquisition of electromyographic signals from the upper limbs of the human body; The acquired electromyographic signals are encoded in a series-parallel manner to obtain the latent space representation of the electromyographic signals. Based on the latent space representation of the electromyographic signals, a semi-structured latent space representation representing the hand movement state is constructed. The latent space representation of electromyography data is used as a conditional vector input into the diffusion model to estimate a semi-structured latent space representation of hand movement state. Real-time angle data of each joint of the hand are obtained by decoding from the semi-structured latent space representation of hand motion state, and then the finger motion tracking results are obtained. Decode human action strategies from finger motion tracking results, and infer information about interactive objects based on human action strategies; By integrating finger motion tracking results with interactive object information, a complete hand labor strategy can be obtained.

2. The method for decoding labor strategies in human-like production scenarios for humanoid robots according to claim 1, characterized in that, EMG sensors are used to collect electromyographic signals from the human forearm or wrist in real time.

3. The method for decoding labor strategies in human-like production scenarios for humanoid robots according to claim 1, characterized in that, The process of performing serial-parallel encoding on the acquired electromyographic signals to obtain the latent space representation of the electromyographic signals specifically includes: An encoder module based on a series-parallel hybrid structure is constructed, including a parallel encoder group, a series encoder, and a series decoder arranged sequentially. Each encoder in the parallel encoder group corresponds one-to-one with each EMG channel of the electromyography signal.

4. The method for decoding labor strategies in human-like production scenarios for humanoid robots according to claim 3, characterized in that, When an EMG sensor fails, the encoding results of the corresponding EMG channel are all set to zero.

5. The method for decoding labor strategies in human-like production scenarios for humanoid robots according to claim 1, characterized in that, The semi-structured latent space representation characterizing the hand movement state includes a pose matrix with physical meaning and latent vectors without physical meaning.

6. The method for decoding labor strategies in human-like production scenarios for humanoid robots according to claim 1, characterized in that, The latent space representation of electromyography data is used as a conditional vector input based on a diffusion model to estimate a semi-structured latent space representation of hand motion state. In estimating the latent space representation of hand motion state, three parallel computations are performed: estimating only the pose matrix with physical meaning, estimating only the latent vector without physical meaning, and estimating the complete latent space representation of hand motion state. The results of the three parallel computations are averaged and output as the final semi-structured latent space representation of hand motion state.

7. The method for decoding labor strategies in human-like production scenarios for humanoid robots according to claim 1, characterized in that, The process of decoding the real-time angle data of each joint of the hand from the semi-structured latent space representation of hand motion state to obtain the finger motion tracking result specifically includes: inputting the semi-structured latent space representation of hand motion state into the decoder, decoding to obtain the real-time angle data of each joint, averaging the real-time angle data of each joint obtained by decoding with the joint angle represented by the pose matrix as the stable anchor point, and outputting the result to obtain the finger motion tracking information.

8. The method for decoding labor strategies in human-like production scenarios for humanoid robots according to claim 1, characterized in that, Human motion strategies are decoded from finger motion tracking results, and interactive object information is inferred based on human motion strategies. When inferring interactive object information, the finger motion tracking information is reconstructed, and the hidden vector part of the hand motion state latent space representation without physical meaning and the angle of each finger joint are decoded simultaneously.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the program, it implements the method as described in any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 8.