Mechanical arm control method and equipment based on human body posture recognition and storage medium

By collecting the operator's human skeleton node coordinate information in real time on the robot and using the action recognition model to dynamically adjust the task execution strategy, the problem that the robotic arm cannot perceive the operator's real-time movement changes is solved, and the efficiency and flexibility of task execution are improved.

CN120395857APending Publication Date: 2025-08-01YOUDI ROBOT (WUXI) CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510660417.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

In the prior art, the robotic arm control mode cannot perceive the operator's real-time movement changes, resulting in manual modification of the program or manual intervention when facing emergencies or task adjustments, reducing task execution efficiency.

Method used

By setting sensors on the robot to collect the operator's human skeleton node coordinate information in real time, use preset action recognition models to identify the action type, and infer the operator's intentions based on scene information and motion trajectory, and dynamically adjust the task execution strategy.

Benefits of technology

It realizes that the robotic arm does not require manual intervention when facing emergencies or task adjustments, which improves the efficiency and flexibility of task execution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120395857A_ABST
    Figure CN120395857A_ABST
Patent Text Reader

Abstract

The invention discloses a mechanical arm control method and device based on human body posture recognition and a storage medium, and relates to the technical field of robots. The mechanical arm control method comprises the steps that human body skeleton joint point coordinate information, collected by a sensor arranged on a robot, of an operator is obtained; on the basis of a preset action recognition model, action type recognition is conducted on the human skeleton joint point coordinate information, and the action type of the operator is obtained; determining the intention of the operator according to the scene information, the action type and the movement track of the operator; and determining a target task associated with the intention of the operator, and controlling the robot to execute the target task. In the invention, the robot can dynamically adjust the task execution strategy according to the real-time action and intention of the operator, so that the task execution efficiency and flexibility are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of robotics, and in particular, to a robotic arm control method, device, and storage medium based on human pose recognition. Background Art

[0002] In current industrial and collaborative scenarios, the control of robotic arms usually relies on an offline pre-programmed instruction set. Before task execution, technicians write a detailed control program using professional programming software based on the working objectives of the robotic arm, setting the motion trajectories, speeds, accelerations of each joint of the robotic arm, and the operation logic of the end effector. However, in this control mode, the robotic arm can only execute tasks according to the preset logic and cannot perceive the real-time action changes of the operator. In the face of emergencies or task adjustments, manual modification of the program or manual intervention must be relied on, which reduces the task execution efficiency.

[0003] The above content is only used to assist in understanding the technical solution of the present application and does not represent an admission that the above content is prior art. Summary of the Invention

[0004] The main purpose of the present application is to provide a robotic arm control method, device, and storage medium based on human pose recognition, aiming to solve the technical problem of how to dynamically adjust the task execution strategy and improve the efficiency and flexibility of task execution.

[0005] To achieve the above objective, the present application proposes a robotic arm control method based on human pose recognition. The robotic arm control method based on human pose recognition includes:

[0006] Obtain the coordinate information of the human skeleton joint points of the operator collected by the sensor provided on the robot;

[0007] Based on a preset action recognition model, perform action type recognition on the coordinate information of the human skeleton joint points to obtain the action type of the operator;

[0008] Determine the operator's intention according to the scene information, the action type, and the operator's motion trajectory;

[0009] Determine the target task associated with the operator's intention and control the robot to execute the target task.

[0010] In an embodiment, the sensor is a camera device. The step of obtaining the coordinate information of the human skeleton joint points of the operator collected by the sensor provided on the robot includes:

[0011] Obtain the image data collected by the camera device, perform human body recognition on the image data according to the target detection algorithm, and obtain the human body region image;

[0012] Perform pose estimation on the human body region image according to the deep learning model to obtain the coordinate information of the human body skeleton joint points.

[0013] In one embodiment, the step of performing pose estimation on the human body region image according to the deep learning model to obtain the coordinate information of the human body skeleton joint points includes:

[0014] Input the preprocessed human body region image into the deep learning model for feature extraction to generate a confidence map and an offset map, where one confidence map corresponds to one human body joint point;

[0015] Obtain the target position with the highest confidence in the confidence map, and confirm the target position as the joint point;

[0016] Determine the connection relationship between the joint points according to the offset map to obtain the human body skeleton structure;

[0017] Obtain the coordinate information of the human body skeleton joint points according to the joint points and the human body skeleton structure.

[0018] In one embodiment, the step of performing action type recognition on the coordinate information of the human body skeleton joint points based on a preset action recognition model to obtain the action type of the operator includes:

[0019] Combine the coordinate information of the joint points of consecutive frames to obtain a joint point sequence;

[0020] Input the joint point sequence into the action recognition model to output the prediction probability of the candidate action;

[0021] Confirm the candidate action with the highest prediction probability as the action type.

[0022] In one embodiment, the step of determining the operator's intention according to the scene information, the action type, and the operator's movement trajectory includes:

[0023] Input the scene information, the action type, and the movement trajectory into the intention prediction model to obtain an intention score;

[0024] Determine the intention with an intention score higher than the preset score threshold as the operator's intention.

[0025] In one embodiment, the step of determining the target task associated with the operator's intention and controlling the robot to execute the target task includes:

[0026] Determine the target task associated with the operator's intention;

[0027] If the target task is a following task, control the robotic arm of the robot to move according to the movement trajectory of the operator;

[0028] If the target task is an avoidance task, control the robotic arm to perform an avoidance action according to a preset safety distance and path planning algorithm.

[0029] In one embodiment, after the step of determining the target task associated with the operator's intention and controlling the robot to execute the target task, it includes:

[0030] According to the task type and task execution stage of the target task, obtain a voice prompt file, a GIF page, and a vibration mode corresponding to the target task;

[0031] Based on the user interface, switch to display the GIF page, and / or call and play the voice file, and / or call the vibration interface to trigger vibration according to the vibration mode.

[0032] In one embodiment, the robotic arm control method based on human body posture recognition further includes:

[0033] Detect the speed value and joint torque value of the robot during the execution of the target task;

[0034] If the speed value is greater than a preset speed threshold, and / or the joint torque value is greater than a preset joint torque threshold, control the robot to perform a deceleration operation or stop moving.

[0035] In addition, to achieve the above object, the present application also proposes a robotic arm control device based on human body posture recognition, the device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, the computer program is configured to implement the steps of the robotic arm control method based on human body posture recognition as described above.

[0036] In addition, to achieve the above object, the present application also proposes a storage medium, the storage medium is a computer-readable storage medium, a computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the robotic arm control method based on human body posture recognition as described above.

[0037] This application proposes a robotic arm control method based on human pose recognition. By means of sensors installed on the robot, the coordinate information of the human skeleton joint points of the operator is collected in real time, and the immediate movement changes of the operator can be sensed. The preset action recognition model is used to process the collected coordinate information of the human skeleton joint points to identify the action types of the operator. Combining the scenario information and the movement trajectory of the operator, the intention of the operator is further inferred, and the task to be executed is determined according to the intention. Since the robot can sense the intention of the operator in real time, when facing sudden situations or task adjustments, there is no need for technicians to repeatedly modify the program or perform manual intervention. The robot can dynamically adjust its task execution strategy according to the real-time actions and intentions of the operator, thereby improving the efficiency and flexibility of task execution. Description of the Drawings

[0038] The drawings here are incorporated into the specification and form a part of this specification, showing the embodiments consistent with this application, and are used together with the specification to explain the principles of this application.

[0039] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0040] Figure 1 The first flow diagram provided for the robotic arm control method based on human pose recognition in this application;

[0041] Figure 2 The second flow diagram provided for the robotic arm control method based on human pose recognition in this application;

[0042] Figure 3 The third flow diagram provided for the robotic arm control method based on human pose recognition in this application;

[0043] Figure 4 The device structure diagram of the hardware operating environment involved in the robotic arm control method based on human pose recognition in the embodiments of this application.

[0044] The implementation, functional features, and advantages of the purpose of this application will be further described with reference to the embodiments and the drawings. Detailed Embodiments

[0045] It should be understood that the specific embodiments described here are only used to explain the technical solutions of this application and are not used to limit this application.

[0046] To better understand the technical solution of this application, the following will be described in detail in conjunction with the accompanying drawings of the specification and specific embodiments.

[0047] The main solution of the embodiment of this application is as follows: Obtain the coordinate information of the human skeleton joint points of the operator collected by the sensor set on the robot; Based on a preset action recognition model, identify the action type of the operator from the coordinate information of the human skeleton joint points, and obtain the action type of the operator; Determine the operator's intention according to the scene information, the action type, and the movement trajectory of the operator; Determine the target task associated with the operator's intention, and control the robot to execute the target task.

[0048] In current industrial and collaborative scenarios, the control of robotic arms usually relies on an offline pre-programmed instruction set. Before task execution, technicians write detailed control programs using professional programming software based on the working objectives of the robotic arm, setting the movement trajectories, speeds, accelerations of each joint of the robotic arm, and the operation logic of the end effector. However, in this control mode, the robotic arm can only execute tasks according to the preset logic and cannot perceive the real-time action changes of the operator. In the face of emergencies or task adjustments, manual modification of the program or manual intervention must be relied on, which reduces the task execution efficiency.

[0049] This application provides a solution. By means of the sensor set on the robot, the coordinate information of the human skeleton joint points of the operator is collected in real time, and the real-time action changes of the operator can be perceived. The preset action recognition model is used to process the collected coordinate information of the human skeleton joint points to identify the action type of the operator. Combining the scene information and the movement trajectory of the operator, the operator's intention is further inferred, and the task to be executed is determined according to the intention. Since the robot can perceive the operator's intention in real time, in the face of emergencies or task adjustments, there is no need for technicians to repeatedly modify the program or perform manual intervention. The robot can dynamically adjust its task execution strategy according to the real-time actions and intentions of the operator, thereby improving the efficiency and flexibility of task execution.

[0050] It should be noted that the execution subject of this embodiment can be a computing service device with network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, a server, a server cluster, etc., or an electronic device, device, etc. that can implement the above functions. Hereinafter, a robotic arm control device based on human pose recognition will be used as an example to illustrate this embodiment and the following embodiments.

[0051] Based on this, the embodiment of this application provides a robotic arm control method based on human pose recognition, referring to Figure 1 , Figure 1 is a schematic flowchart of the first embodiment of the robotic arm control method based on human pose recognition of this application.

[0052] In this embodiment, the robotic arm control method based on human pose recognition includes steps S10 to S40:

[0053] Step S10: Obtain the coordinate information of the human skeleton joint points of the operator collected by the sensors arranged on the robot.

[0054] In this embodiment, visual sensors such as multi-view cameras, RGB, and depth cameras, or wearable sensors such as inertial measurement units (IMUs) are used to collect human key points in real time, such as the spatial positions and movement trajectories of joints and the torso.

[0055] In a feasible implementation manner, an RGB image containing the operator is obtained through an RGB camera, and deep learning models such as OpenPose and HRNet are used to extract the joint point coordinates from the RGB image, and the joint pixel positions are located through a heatmap.

[0056] In a feasible implementation manner, human key point information is collected according to wearable sensors. IMU sensors are arranged at joints such as the hands, wrists, and elbows that the operator needs to track. The IMU sensors collect acceleration and angular velocity information in real time. According to the Kalman filter or complementary filter algorithm, the data of the accelerometer and gyroscope are combined to dynamically estimate the attitude information of the sensor. Specifically, the attitude at the next moment is predicted according to the angular velocity integration of the gyroscope, and the covariance matrix represents the uncertainty of the predicted attitude. The gravity direction of the accelerometer is used as the observation value, and the residual between the observation value and the predicted value is calculated. The optimal weight Kalman gain value is calculated according to the predicted covariance matrix and the unreliability of the accelerometer, and the obtained Kalman gain value is used to mix the predicted value and the observation value to obtain the attitude information of the sensor. Combining the attitude information of the sensor and the preset human skeleton model, the pose information of the operator is calculated joint by joint.

[0057] Step S20: Based on a preset action recognition model, perform action type recognition on the coordinate information of the human skeleton joint points to obtain the action type of the operator.

[0058] In this embodiment, according to the obtained joint point coordinates, the trajectory data of the joint points are determined by the coordinate changes in the time series. The joint point coordinates are converted to a unified anatomical coordinate system (such as taking the waist as the origin), and all joint point coordinates are divided by the operator's height and normalized to the range of [0, 1] to eliminate the influence of individual height and body type differences. The joint point trajectory is segmented according to a fixed time window (such as 0.5 seconds) to form action segments. Exemplarily, a 10-second waving action is divided into 20 time windows, and each window corresponds to an action segment.

[0059] Feature extraction is performed on each action segment to obtain joint point features. The joint point features are input into a preset action recognition model to obtain the action type corresponding to the action segment and its confidence. The action recognition model can be a machine learning model such as a support vector machine (SVM), random forest (RF), K-nearest neighbor (KNN), etc., or a deep learning model such as a recurrent neural network or a graph convolutional network. Optionally, a confidence threshold can be set to filter out low-confidence results and avoid misjudgment.

[0060] In a feasible implementation manner, step S20 may include steps S21 to S23:

[0061] Step S21, combine the joint point coordinate information of consecutive frames to obtain a joint point sequence.

[0062] In this implementation manner, the joint point coordinate information detected in consecutive multiple frames is combined in chronological order to form a joint point sequence. For example, 10 frames of images are collected, and 15 joint points are detected in each frame of image. The obtained joint point sequence is a three-dimensional data structure containing 10 time steps, and each time step has 15 joint point coordinates.

[0063] Optionally, the time steps of the joint point sequence can be evenly distributed by interpolation or frame discarding. The time stamp corresponding to each frame of image is obtained from the sensor, which records the acquisition time of each frame of image. For two consecutive frames of images, calculate the time interval between them. According to the requirements of the action recognition model and the actual application scenario, determine the target time step. For some rapidly changing actions, a shorter target time step is required to capture the details of the action; while for some slowly changing actions, the target time step can be appropriately extended.

[0064] When the overall frame rate is low and there is a large time interval, the interpolation method is used to increase the number of frames. Find adjacent frames with a time interval greater than the target time step. Divide the absolute value of the difference between the time stamps of the adjacent frames by the target time step and subtract 1 to obtain the number of frames to be inserted. Use the linear interpolation formula to calculate the joint point coordinates of the inserted frames. Exemplarily, the time stamp of the first frame is 0 milliseconds, and the joint point coordinates are (10, 20). The time stamp of the second frame is 60 milliseconds, and the joint point coordinates are (40, 50). The target time step is 30 milliseconds. Then 1 frame needs to be inserted, and the time stamp of the inserted frame is 30 milliseconds. The joint point coordinates of the inserted frame are calculated as follows: x = (30 - 0 / 60 - 0) / 60 - 0 × (40 - 10) = 25, y = 20 + (30 - 0 / 60 - 0) × (50 - 20) = 35.

[0065] When the overall frame rate is relatively high and there are some adjacent frames with too small time intervals, the frame discarding method is adopted to reduce the number of frames. Traverse the entire frame sequence and calculate the time intervals between adjacent frames. For adjacent frames with time intervals smaller than the target time step, select to discard one of them. You can choose to discard the frame with a smaller timestamp, or according to other rules, such as according to the movement amplitude of the joint points, discard the frame with a smaller movement amplitude. Repeat the above process until the time intervals between adjacent frames in the entire frame sequence are greater than or equal to the target time step. After interpolating or discarding frames, use a smoothing algorithm to smooth the joint point coordinates to reduce noise and mutations.

[0066] In this embodiment, by adjusting the frame sequence, it is possible to ensure that the time intervals between consecutive frames are consistent, avoid timing misalignment caused by unstable frame rates, and thus improve the accuracy of action recognition.

[0067] Step S22: Input the joint point sequence into the action recognition model and output the prediction probabilities of candidate actions.

[0068] In this embodiment, the spatial features and temporal features of the joint point sequence are extracted, and a weight coefficient is learned for the spatial features and temporal features respectively. These weights are learned through a neural network or a simple linear transformation. For example, use a fully connected layer to map the spatial features and temporal features to a scalar weight w1 and w2 respectively, and then normalize these two weights through the softmax function so that their sum is 1. Multiply the normalized weights by the spatial features and temporal features respectively, and then add the results to obtain the fused features. Input the fused features into the fully connected layer of the deep learning model. The number of neurons in the fully connected layer is equal to the number of candidate actions. For example, if the candidate actions include 4 actions: "standing", "walking", "reaching", and "grasping", then the number of neurons in the fully connected layer is 4. The fully connected layer performs a linear transformation on the input features through the weight matrix W and the bias term b to obtain the scores of each candidate action. Finally, perform Softmax normalization on the scores to convert the scores into a probability distribution.

[0069] Optionally, use a CNN to extract the spatial features of the joint point sequence. After preprocessing each frame image in the joint point sequence, extract the local spatial relationship features between joint points through convolution operations, and reduce the size of the feature map through pooling operations to obtain spatial features.

[0070] Optionally, a GCN is used to extract the spatial features of the joint point sequence. According to the natural connection relationship of human joint points, a human skeleton graph is constructed. For example, joint points such as the head, neck, shoulders, elbows, and wrists are connected by edges in accordance with the actual connection method of the human body to form a graph structure. Each joint point can be regarded as a node in the graph, and the weights of the edges between the nodes can be determined according to factors such as the physical distance and motion correlation between the joint points. Through graph convolution operations, the information of the node and its neighbor nodes is aggregated to update the features of the node. For each node, its new feature value is the weighted sum of its own features and the features of its neighbor nodes. The weights can be obtained through learning or determined according to predefined rules (such as distance-based weights). For example, for the shoulder joint point, its new feature value is the sum of its own feature value and the weighted sum of the feature values of the connected joint points such as the neck and elbows.

[0071] In this embodiment, an LSTM is used to extract the temporal features of the joint point sequence. For each frame in the joint point sequence, the spatial features of this frame are used as the input x at the current moment t , which is combined with the state at the previous moment to generate the hidden state and cell state at the current moment. Among them, the hidden state contains the temporal feature information at the current moment and is used as one of the inputs of the LSTM unit at the next moment. After being processed by multiple layers of LSTM, the hidden state of the last layer is used as the temporal feature representation of the entire sequence.

[0072] Step S23: Confirm the candidate action with the highest prediction probability as the action type.

[0073] In this embodiment, a probability threshold is set, and the selection of the probability threshold can be adjusted according to the requirements of the specific task. Traverse the probability distribution to filter out the action types with probabilities higher than this threshold as candidate actions, and select the candidate action with the highest probability as the final action type. Setting the probability threshold to filter out candidate actions can reduce false alarms.

[0074] Step S30: Determine the operator's intention according to the scene information, the action type, and the operator's movement trajectory.

[0075] In this embodiment, the operator's intention is determined based on the scenario information, action type, and the operator's motion trajectory. The scenario information may include an industrial production line, a home kitchen, a hotel lobby, etc. In different scenarios, the same action may represent different intentions when the specific scenario where the operator is located is different. The operator's actions are classified into basic types, such as grasping, placing, moving, rotating, etc., and the combinations and sequences of actions are analyzed. The combination of multiple actions can form a more complex intention. For example, the action sequence of reaching out first, then grasping, and then moving may represent picking up goods and transporting them to a designated location in a warehouse scenario. Trajectory shape: Analyze the shape of the operator's motion trajectory, such as a straight line, a curve, a broken line, etc. Trajectories of different shapes may represent different intentions. For example, a straight-line trajectory may indicate directly going to the target location, and a curved trajectory may indicate bypassing an obstacle.

[0076] In a feasible implementation manner, a rule library is established according to different scenarios and action types. The rule library contains the intentions corresponding to various combinations of actions and trajectories. The scenario information, action type, and motion trajectory obtained in real time are matched with the rules in the rule library. If the match is successful, the operator's intention is determined.

[0077] In a feasible implementation manner, the scenario information, action type, and motion trajectory are analyzed based on a machine learning algorithm to predict the operator's intention. Step S30 may include steps S31 to S32:

[0078] Step S31, input the scenario information, the action type, and the motion trajectory into an intention prediction model to obtain an intention score.

[0079] Step S32, determine the intention with the intention score higher than a preset score threshold as the operator's intention.

[0080] In this implementation manner, the scenario information may include objects, devices, personnel distribution, etc. in the environment. The scenario information is converted into structured data, for example, using feature vectors or embedding vectors to represent the key elements in the scenario. The action type includes basic actions such as grasping, placing, moving, etc., and the action type can be encoded as a one-hot vector or represented using embedding technology. The motion trajectory includes the motion path, speed, acceleration, etc., and the trajectory data can be discretized into a time series and represented using trajectory features (such as average speed, trajectory length).

[0081] In this embodiment, a deep learning model or a traditional machine learning model can be selected as the intent prediction model. The intent prediction model receives a combined input of scene information, action type, and motion trajectory. The model outputs a score for each possible intent, indicating the likelihood of that intent. The intent scores output by the model are compared with a threshold, and intents with scores higher than the preset threshold are selected as candidate intents. Among the candidate intents, the intent with the highest score is selected as the final operator intent.

[0082] Step S40: Determine the target task associated with the operator intent and control the robot to execute the target task.

[0083] In this embodiment, based on a preset intent-task knowledge base, the operator intent is mapped to a specific target task. For example, in a hotel scenario, after the operator checks into a room in the hotel lobby and is about to reach for their handbag, the robot recognizes the intent as "grasp", and the corresponding task is to assist in retrieving the handbag. This task includes multiple subtasks such as locating the handbag, planning a path, performing the grasp, and delivering the item. Each subtask has a preset task template, and dynamic parameters such as the path and speed in the task template are filled through visual recognition or spatial positioning technology. For example, after determining the current position of the handbag and the desired delivery position of the guest (the direction of hand extension or next to a nearby seat), the robotic arm first picks up the handbag and plans a collision-free path from the handbag position to the guest's hand based on the A* or RRT algorithm and delivers it to the operator's hand.

[0084] In this embodiment, the coordinate information of the human skeleton joint points of the operator is collected in real time through sensors installed on the robot, and the immediate action changes of the operator can be sensed. The preset action recognition model is used to process the collected coordinate information of the human skeleton joint points to identify the action type of the operator. Combining the scene information and the operator's motion trajectory, the operator's intent is further inferred, and the task to be executed is determined according to the intent. Since the robot can sense the operator's intent in real time, when facing unexpected situations or task adjustments, there is no need for technicians to repeatedly modify the program or perform manual intervention. The robot can dynamically adjust its task execution strategy according to the operator's real-time actions and intents, thereby improving the efficiency and flexibility of task execution.

[0085] Based on the first embodiment of the present application, in the second embodiment of the present application, for content that is the same as or similar to the above-mentioned first embodiment, reference can be made to the above introduction and will not be elaborated hereinafter. On this basis, step S10 may include steps A10 to A20:

[0086] Step A10: Obtain the image data collected by the imaging device, perform human body recognition on the image data according to the target detection algorithm, and obtain the human body region image.

[0087] In this embodiment, a video stream is collected in real time by a camera, and the collected image is input into a target detection model. The preprocessed image is input into the target detection model for forward propagation calculation. The model outputs the detected object category, bounding box coordinates, and confidence scores. The higher the confidence score, the more reliable the detection result. A confidence threshold is set according to application requirements, and the detection results with confidence higher than this threshold are retained. For each detected result that passes the screening, its bounding box coordinates are extracted, and according to the bounding box coordinates, the corresponding human body region image is cropped from the original image.

[0088] Step A20: Perform pose estimation on the human body region image according to the deep learning model to obtain the coordinate information of the human body skeleton joint points.

[0089] In this embodiment, a pre-trained pose estimation model, such as OpenPose, HRNet, AlphaPose, etc., is used to detect human body skeleton joint points. The human body region image is input into the pose estimation model, and the model outputs the coordinate information of the human body skeleton joint points (such as shoulders, elbows, wrists, hips, knees, ankles, etc.).

[0090] Specifically, step A20 may include steps A11 to A14:

[0091] Step A11: Input the preprocessed human body region image into the deep learning model for feature extraction to generate a confidence map and an offset map, where one confidence map corresponds to one human body joint point.

[0092] Step A12: Obtain the target position with the highest confidence in the confidence map, and confirm the target position as the joint point.

[0093] In this implementation manner, the preprocessed human body region image is input into the deep learning model for forward propagation calculation, and a confidence map and an offset map are output. Among them, each confidence map corresponds to one human body joint point, indicating the probability that the joint point appears at a certain position in the image. The offset map is used to describe the connection relationship between joint points, indicating the limb direction and strength. Traverse each confidence map, find the pixel position with the highest confidence, and confirm this position as the coordinate of the corresponding joint point.

[0094] Step A13: Determine the connection relationship between the joint points according to the offset map to obtain the human body skeleton structure.

[0095] Step A14: Obtain the coordinate information of the human body skeleton joint points according to the nodes and the human body skeleton structure.

[0096] In this embodiment, the connection relationship between joint points is determined using the vector information in the offset map. The vector information includes the vector direction and the vector intensity. The vector direction represents the direction from one end of the limb (such as the wrist) to the other end (such as the elbow), and the vector intensity represents the possibility of the presence of a limb in that direction. The vector information corresponding to the limb is extracted from the offset map output by the deep learning model. For each limb type (such as the left arm, right arm), the corresponding offset map is extracted. The joint point pairs are traversed. For each pair of potentially connected joint points (such as the wrist and the elbow), all possible paths between these two joint points are traversed. For each point on the path, the vector value of this point in the offset map and the vector integral on the path are calculated. If the vector integral value on the path exceeds a certain vector integral threshold, indicating that the direction and intensity of this path are strong enough, then it is considered that these two joint points belong to the same limb and there is a connection relationship. According to the determined connection relationship, a connection graph is constructed to represent the connection between joint points. A graph matching algorithm, such as using the Hungarian algorithm, is used to connect the joint points into a complete human skeleton. The coordinate information of each joint point is extracted from the skeleton structure to obtain the coordinate information of the joint points of the human skeleton.

[0097] Based on the first embodiment of the present application, in the third embodiment of the present application, the same or similar content as that in the above-mentioned first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, steps S41 to S43 may further be included before step S40:

[0098] Step S41, determining the target task associated with the operator's intention.

[0099] In this embodiment, the target task may be a cooperation mode, a following mode, an avoidance mode, or an assistance mode, and the task mode of the target task is switched according to the intention analysis result. In the following mode, the end of the robotic arm tracks the trajectory of the operator's hand to assist in completing fine operations, such as delivering items and instruments. In the avoidance mode, the safe distance between the human body and the robotic arm is calculated in real time, and active avoidance is achieved through dynamic path planning. In the assistance mode, the task is executed in advance according to the action prediction, such as the robotic arm picking up a handbag to achieve "helping the guest carry the handbag".

[0100] Step S42, if the target task is a following task, controlling the robotic arm of the robot to move according to the movement trajectory of the operator.

[0101] In this embodiment, according to the operator's movement trajectory, the target position of the end effector of the robotic arm is calculated. The target position of the end effector of the robotic arm can be set to a joint point of the operator, such as the position of the wrist. Using the inverse kinematics algorithm, the target position and orientation of the end effector in the Cartesian space are determined. The target position of the end effector is converted into the joint angles of the robotic arm. The calculation is performed according to the DH parameters (Denavit-Hartenberg parameters) or URDF model of the robotic arm. According to the joint angles obtained from the inverse kinematics calculation, a control command for the robotic arm is generated. The control command usually includes joint angle, speed, and acceleration information. The control command is sent to the robotic arm controller to control the robotic arm to move along the planned path.

[0102] Optionally, in the process of determining the target position and orientation of the end effector in the Cartesian space, first, analyze the geometric structure of the robotic arm to identify the relationships between the various links and joints of the robotic arm. According to the target position and orientation, draw the geometric figure of the robotic arm in the Cartesian space. Second, use trigonometric functions (such as sine, cosine, tangent) and geometric relationships to gradually solve the angles of each joint. Starting from the last joint of the end effector, calculate the angles of each joint step by step forward.

[0103] Step S43, if the target task is an avoidance task, according to the preset safety distance and path planning algorithm, control the robotic arm to perform an avoidance action.

[0104] In this embodiment, data on the environment around the robotic arm is collected, and the collected sensor data is processed to identify the position, size, and shape of the obstacles. According to the working environment and task requirements of the robotic arm, a suitable path planning algorithm is selected, such as the A* algorithm, RRT (Rapidly-Exploring Random Tree) algorithm, Dijkstra algorithm, etc. Using the selected path planning algorithm, combined with the position of the obstacles and the safety distance, an avoidance path is generated. The generated avoidance path is converted into a motion instruction for the robotic arm, and a PID controller or other control algorithms are used to control the robotic arm to move along the avoidance path. During the movement of the robotic arm, the position and motion state of the obstacles are monitored in real time. If the obstacles move or new obstacles appear, the avoidance path is adjusted to ensure that the robotic arm and the obstacles always maintain the set safety distance.

[0105] Based on the first embodiment of the present application, in the fourth embodiment of the present application, the same or similar content as in the above-mentioned embodiment one can be referred to the above introduction and will not be elaborated hereinafter. On this basis, please refer to Figure 2 , after step S40, the robotic arm control method based on human pose recognition further includes steps S50 to S60:

[0106] Step S50: Obtain a voice prompt file, a GIF page, and a vibration pattern corresponding to the target task according to the task type and task execution stage of the target task.

[0107] In this embodiment, corresponding voice prompts, GIF pages, and vibration feedback are provided according to the task type and execution stage of the target task. The required feedback type is determined according to the nature of the task (such as navigation, avoidance, grasping, following, etc.). For example, navigation tasks require voice prompts and GIF guidance, while grasping tasks require vibration feedback to enhance the operation feeling. The task is divided into different execution stages, such as start, in progress, completed, abnormal, etc., and each stage has a different feedback form.

[0108] Prepare corresponding voice prompt files for each task type and execution stage. For example, "Start navigation" can be played at the start of a navigation task, and "Navigation completed" can be played when it is completed. Design and prepare GIF pages that match the task type and execution stage. The GIF can display operation steps, status indicators, or progress information. Define different vibration patterns, such as short vibration, long vibration, continuous vibration, etc., to correspond to different task stages or states. For example, short vibration can be used when the task is completed, and continuous vibration can be used in case of an abnormal situation. Create a mapping table or database to associate the task type and execution stage with the corresponding voice prompt file, GIF page, and vibration pattern. For example, when the task type is "avoidance" and the execution stage is "in progress", it is mapped to the voice file "Avoiding.mp3", the GIF page "Avoidance GIF.gif", and the "short vibration" mode.

[0109] Step S60: Based on the user interface, switch to display the GIF page, and / or call and play the voice file, and / or call the vibration interface to trigger vibration according to the vibration pattern.

[0110] In this embodiment, according to the current task type and execution stage, obtain the corresponding voice prompt file, GIF page, and vibration pattern from the mapping table or database. Switch to display the obtained GIF page on the user interface of the device to ensure that the GIF page can clearly display the task status or operation guidance. Call the voice playback interface of the device to play the obtained voice prompt file. Call the vibration interface of the device to trigger vibration according to the obtained vibration pattern. Through visual (UI interface), tactile (vibration reminder), or voice prompts, the status of the robotic arm and the robot is fed back to the operator. For example, switch the UI to the GIF page of "Sorry" and voice broadcast "Sorry, sorry", which enhances the collaboration transparency.

[0111] Based on the first embodiment of the present application, in the fifth embodiment of the present application, the same or similar content as in the above-mentioned first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 3, the method for controlling a robotic arm based on human pose recognition may further include steps B10 to B20:

[0112] Step B10, detecting the speed value and joint torque value of the robot during the execution of the target task.

[0113] Step B20, if the speed value is greater than a preset speed threshold, and / or the joint torque value is greater than a preset joint torque threshold, controlling the robot to perform a deceleration operation or stop moving.

[0114] In this embodiment, safety strategies such as speed limits and joint torque thresholds are introduced to ensure immediate shutdown or deceleration in case of emergencies. Speed sensors such as encoders and tachogenerators, and torque sensors such as strain gauges and torque sensors are installed on each joint and actuator of the robot. The speed value and joint torque value from the sensors are received in real time. According to the design parameters, task requirements, and safety standards of the robot, the speed threshold and joint torque threshold are set. The real-time obtained speed value is compared with the preset speed threshold, and the joint torque value is compared with the preset joint torque threshold. If the speed value exceeds the speed threshold, or the joint torque value exceeds the joint torque threshold, it is determined as an over-limit state. When it is detected that the speed value or joint torque value exceeds the threshold, a deceleration operation is first attempted. By reducing the output power of the actuator or sending a deceleration instruction, the movement speed of the robot is gradually reduced. If the deceleration operation cannot effectively reduce the speed value or joint torque value, or the over-limit state persists, a stop movement instruction is triggered to cut off the power supply of the actuator or send a stop instruction to control the robot to stop moving. By real-time monitoring of the speed value and joint torque value, mechanical failures, collisions, or personnel injuries caused by excessive speed or torque can be detected and prevented in a timely manner. When an over-limit state is detected, the system can respond quickly and perform deceleration or stop movement operations, thereby minimizing potential safety risks.

[0115] It should be noted that the above examples are only for understanding the present application and do not constitute a limitation on the method for controlling a robotic arm based on human pose recognition of the present application. More simple transformations in various forms based on this technical concept are within the protection scope of the present application.

[0116] The present application provides a device for controlling a robotic arm based on human pose recognition. The device for controlling a robotic arm based on human pose recognition includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method for controlling a robotic arm based on human pose recognition in the first embodiment above.

[0117] Next, refer to Figure 4, which shows a schematic structural diagram of a robotic arm control device suitable for implementing the embodiment of the present application based on human pose recognition. The robotic arm control device based on human pose recognition in the embodiment of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, personal digital assistants (PDAs), tablet computers (PADs), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 4 The shown robotic arm control device based on human pose recognition is merely an example and should not impose any limitation on the functions and usage scope of the embodiment of the present application.

[0118] As Figure 4 shown, the robotic arm control device based on human pose recognition may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1002 or the program loaded from the storage device 1003 into the random access memory (RAM) 1004. In the random access memory 1004, various programs and data required for the operation of the robotic arm control device based on human pose recognition are also stored. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. The input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the robotic arm control device based on human pose recognition to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a robotic arm control device with various systems, it should be understood that it is not required to implement or have all the shown systems. More or fewer systems may be alternatively implemented or had.

[0119] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by a processing device 1001, the above-mentioned functions defined in the methods of the embodiments disclosed in the present application are executed.

[0120] The robotic arm control device based on human pose recognition provided by the present application adopts the robotic arm control method based on human pose recognition in the above embodiments, and can solve the technical problem of how to dynamically adjust the task execution strategy and improve the efficiency and flexibility of task execution. Compared with the prior art, the beneficial effects of the robotic arm control device based on human pose recognition provided by the present application are the same as those of the robotic arm control method based on human pose recognition provided in the above embodiments, and other technical features in the robotic arm control device based on human pose recognition are the same as the features disclosed in the method of the previous embodiment, which will not be elaborated here.

[0121] It should be understood that the various parts disclosed in the present application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0122] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0123] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the robotic arm control method based on human pose recognition in the above embodiments.

[0124] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium can be transmitted using any suitable medium, including but not limited to: wires, optical cables, radio frequency (RF), etc., or any suitable combination of the above.

[0125] The above computer-readable storage medium can be included in a robotic arm control device based on human pose recognition; it can also exist independently without being assembled into a robotic arm control device based on human pose recognition.

[0126] The above computer-readable storage medium carries one or more programs. When the one or more programs are executed by a robotic arm control device based on human pose recognition, the robotic arm control device based on human pose recognition is caused to: obtain the human skeleton joint point coordinate information of the operator collected by a sensor disposed on the robot; based on a preset action recognition model, perform action type recognition on the human skeleton joint point coordinate information to obtain the action type of the operator; determine the operator's intention according to the scene information, the action type, and the operator's movement trajectory; determine a target task associated with the operator's intention, and control the robot to execute the target task.

[0127] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by connecting through the Internet using an Internet service provider).

[0128] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0129] The modules involved in the embodiments described in this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation to the unit itself in some cases.

[0130] The readable storage medium provided in this application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for performing the above-mentioned robotic arm control method based on human pose recognition, and can solve the technical problem of how to dynamically adjust the task execution strategy and improve the efficiency and flexibility of task execution. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the robotic arm control method based on human pose recognition provided in the above embodiments, and will not be elaborated here.

[0131] The present application also provides a computer program product, including a computer program which, when executed by a processor, implements the steps of the above-described robotic arm control method based on human pose recognition.

[0132] The computer program product provided by the present application can solve the technical problem of how to dynamically adjust the task execution strategy and improve the efficiency and flexibility of task execution. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as those of the robotic arm control method based on human pose recognition provided in the above embodiments, and will not be elaborated herein.

[0133] The above are only partial embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structural transformation made under the technical concept of the present application by using the content of the specification and drawings of the present application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.

Claims

1. A robotic arm control method based on human pose recognition, characterized in that, The robotic arm control method based on human pose recognition includes: Obtaining the coordinate information of the human skeleton joint points of the operator collected by the sensor set on the robot; Based on a preset action recognition model, performing action type recognition on the coordinate information of the human skeleton joint points to obtain the action type of the operator; Determining the operator's intention according to the scene information, the action type, and the operator's movement trajectory; Determining a target task associated with the operator's intention and controlling the robot to execute the target task.

2. The robotic arm control method based on human body posture recognition according to claim 1, characterized in that, The sensor is a camera device, and the step of obtaining the coordinate information of the human skeleton joint points of the operator collected by the sensor set on the robot includes: Obtaining the image data collected by the camera device, performing human body recognition on the image data according to the target detection algorithm to obtain a human body region image; Performing pose estimation on the human body region image according to a deep learning model to obtain the coordinate information of the human skeleton joint points.

3. The robotic arm control method based on human body posture recognition according to claim 2, characterized in that, The step of performing pose estimation on the human body region image according to a deep learning model to obtain the coordinate information of the human skeleton joint points includes: Inputting the preprocessed human body region image into the deep learning model for feature extraction to generate a confidence map and an offset map, where one confidence map corresponds to one human body joint point; Obtaining the target position with the highest confidence in the confidence map and confirming the target position as a joint point; Determining the connection relationship between the joint points according to the offset map to obtain a human skeleton structure; Obtaining the coordinate information of the human skeleton joint points according to the nodes and the human skeleton structure.

4. The robotic arm control method based on human body posture recognition according to claim 1, wherein The step of performing action type recognition on the coordinate information of the human skeleton joint points based on a preset action recognition model to obtain the action type of the operator includes: Combining the coordinate information of the joint points of consecutive frames to obtain a joint point sequence; Inputting the joint point sequence into the action recognition model and outputting the prediction probability of candidate actions; Confirming the candidate action with the highest prediction probability as the action type.

5. The robotic arm control method based on human body posture recognition according to claim 1, characterized in that, The step of determining the operator's intention according to the scene information, the action type, and the operator's movement trajectory includes: Inputting the scene information, the action type, and the movement trajectory into an intention prediction model to obtain an intention score; Determining the intention with an intention score higher than a preset score threshold as the operator's intention.

6. The robotic arm control method based on human body posture recognition according to claim 1, characterized in that, The step of determining a target task associated with the operator's intention and controlling the robot to execute the target task includes: Determining the target task associated with the operator's intention; If the target task is a following task, controlling the robotic arm of the robot to move according to the operator's movement trajectory; If the target task is an avoidance task, controlling the robotic arm to perform an avoidance action according to a preset safety distance and a path planning algorithm.

7. As described in claim 1 of the robotic arm control method based on human pose recognition, after the step of determining a target task associated with the operator's intention and controlling the robot to execute the target task, it includes: Obtain a voice prompt file, a GIF page, and a vibration mode corresponding to the target task according to the task type and task execution stage of the target task; Based on the user interface, switch to display the GIF page, and / or, call and play the voice file, and / or, call the vibration interface to trigger vibration according to the vibration mode.

8. The robotic arm control method based on human pose recognition according to claim 1, wherein the robotic arm control method based on human pose recognition further comprises: Detect the speed value and joint torque value of the robot during the execution of the target task; If the speed value is greater than a preset speed threshold, and / or, the joint torque value is greater than a preset joint torque threshold, control the robot to perform a deceleration operation or stop moving.

9. A robotic arm control device based on human pose recognition, characterized in that, The device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the robotic arm control method based on human pose recognition according to any one of claims 1 to 8.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by the processor, the steps of the robotic arm control method based on human pose recognition according to any one of claims 1 to 8 are implemented.

Citation Information

Cited By

  • Remote control method and system for health care robot based on artificial intelligence

    CN121492038A

  • A health and rehabilitation robot remote control method and system based on artificial intelligence

    CN121492038B