Method and system for planning disabled-helping robot based on human body posture capture
By using deep learning algorithms with human posture capture and emotion analysis technology in the disabled robot, robot motion planning is optimized, and the existing disabled robot interaction system is not intelligent and inaccurate, realizing personalized rehabilitation plans and more efficient rehabilitation training effects.
Patent Information
- Application Number
- CN202510235084.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-03
AI Technical Summary
The existing disabled-assisted robots have not yet fully developed in intelligent interaction systems, making it difficult to achieve safe, convenient, intelligent and accurate human-computer interaction. The traditional force interaction method has limited perception range, making it difficult to obtain the patient's real-time posture.
The deep learning algorithm based on human posture capture and emotion analysis technology is used to optimize the robot motion planning, and a safe, convenient, intelligent and accurate interaction system is designed. Through high-definition RGB cameras, depth sensors and convolutional neural networks, patients' posture and emotional information are obtained in real time and the robot motion strategy is adjusted.
It has realized personalized rehabilitation plans, which have improved the safety, convenience and effectiveness of rehabilitation training, and is suitable for patients in elderly, mobility-constrained or telemedicine environments.
Smart Images

Figure CN120085757A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of robot control, and particularly relates to a planning method and system for a disabled-assisting robot based on human posture capture. Background Art
[0002] With the intensification of global aging, the demand for rehabilitation medicine and assisted care has increased sharply, and the relevant medical resources are increasingly scarce. The wide application of robotic intelligent rehabilitation platforms is imperative. Most of the existing traditional rehabilitation devices rely on fixed programs and cannot be adjusted in real time according to the individual differences of patients. Traditional methods are difficult to play a role when medical resources are scarce in remote areas, and there are deficiencies in high-intensity and high-risk (such as infectious diseases) in nursing work. With the development of artificial intelligence and machine vision technologies and their extensive application in medicine, the intelligence level of disabled-assisting robots has been continuously improved, which also provides a possibility to solve this problem. Human pose estimation technology can obtain the position information of each joint in the human skeleton in real time through a simple visual sensor. Based on this, it is fully feasible to study the posture and motion information of the human body during movement and use this for information interaction between patients and disabled-assisting robots.
[0003] Currently, due to the increasing demand for domestic rehabilitation medical resources, robotic rehabilitation training platforms have a huge application prospect. In the related technologies of disabled-assisting robots, the structural design and control system design have been relatively mature, while the research on the interaction system, especially the intelligent interaction system, is still in its infancy, and a relatively simple contact force interaction method is still used clinically. There are still the following three technical problems to be solved in the current technical field of disabled-assisting robots:
[0004] (1) Most of the research on disabled-assisting robots focuses on mechanical structures and control systems, and there is less research on intelligent interaction between patients and robots. At present, there has been relatively in-depth research on the mechanical structure designed according to the characteristics of patients themselves and the compliant control of the robotic arm for assisted movement, and numerous prototypes of disabled-assisting robots have emerged. However, due to the complex actual rehabilitation training environment and the possible abnormal states of patients, the clinical application prospect of disabled-assisting robots is still not optimistic in the absence of a safe, convenient, intelligent and accurate human-machine interaction technology method.
[0005] (2) Common force interaction methods play a significant role in the safe and smooth interaction between the actuators of assistive robots and patients. However, due to their limited perception range, it is difficult to obtain the real-time posture of patients during training, and it is difficult to meet the needs of actual training environments. Existing studies use various types of sensors to collect patient motion parameters and physiological signal data to detect the patient's status to achieve more intelligent interaction. However, considering the high cost of physiological signal acquisition equipment and the complex and cumbersome use process, its widespread clinical application has certain limitations. How to develop a simple and convenient non-contact posture detection system is critical to the scientificity and effectiveness of robotic rehabilitation training.
[0006] (3) At present, the research on the interaction technology of assistive robots based on visual information and virtual reality is still in its initial stage, and has not fully utilized the technical advantages brought by mature machine vision algorithms. Existing research mainly focuses on using machine vision algorithms to obtain human posture information and realize feedback-based information interaction. This method can help assistive robots understand abnormal information during patient training and adjust control strategies to improve the efficiency of rehabilitation. However, relying entirely on feedback-based interaction ignores the patient's own training willingness and fails to fully utilize the patient's own characteristics and machine vision technology to assist them in completing autonomous intention training. Summary of the invention
[0007] In response to the above technical problems, the present invention proposes a planning method and system for a disabled-assistive robot based on human posture capture. The idea is to use posture capture and emotion analysis technology, optimize robot motion planning through deep learning algorithms, and design a safe, convenient, intelligent and precise interactive system for clinical applications. It can implement personalized rehabilitation plans for patients who are elderly, have limited mobility or are in remote medical environments, thereby improving the effect of rehabilitation training.
[0008] In a first aspect, the present invention proposes a method for planning a disabled assistance robot based on capturing human posture, comprising the following steps:
[0009] S1: Fix the patient to the system body through a flexible fixing device to ensure posture capture accuracy and prevent accidental slipping;
[0010] S2: Use a high-definition RGB camera and depth sensor to obtain the patient's real-time posture joint data, including the position of the head, limbs and torso, based on the human posture estimation algorithm;
[0011] S3: Real-time solution of the captured data, referring to the posture database, to generate the optimal solution for the initial posture of rehabilitation;
[0012] S4: Use cameras, microphones and other sensors to collect visual, voice and physiological signals, and use convolutional neural networks for feature extraction and fusion to improve recognition accuracy;
[0013] S5: Through emotion recognition technology, combined with voice emotion analysis, determine the patient's satisfaction with the current posture and adjust the system according to the patient's satisfaction.
[0014] More specifically, step S1 includes:
[0015] S11: Waiting for the patient to be in place, the system enters a standby state and is ready to receive instructions from the user;
[0016] S12: The patient or user directly interacts with the system through voice module commands and inputs commands to the assistive robot through spoken language, that is, the voice signal is transmitted to the processor;
[0017] S13: The pre-processing module processes the input speech features into acoustic feature vectors that can be recognized by the encoder, and converts the real text labels corresponding to the speech signals into text embedding sequences. The sequences obtained by the pre-processing module are added to their respective position codes and then enter the multi-level encoder and decoder. The specific calculation formula of the position code is shown in (1-1), which is used to convert the speech signal into a feature vector that can be recognized by the encoder.
[0018]
[0019] More specifically, step S2 includes:
[0020] S21: Entering the data collection phase, the laser radar sensor is used to scan the surrounding environment within a set frequency range to obtain detailed three-dimensional point cloud data;
[0021] S22: Entering the data preprocessing stage, the three-dimensional point cloud data collected in step S21 is subjected to real-time smooth filtering to effectively improve the data quality; considering that the abnormal value is usually a sudden change value that is extremely different from the actual data, a data smoothing operation is performed;
[0022] Definition of kp i represents the coordinate value of a key point at time i, e is the error threshold between two adjacent frames of data, the initial value is set to the difference between the first two frames of data, A is the amplitude gain constant; when continuously collecting data, if the difference between two adjacent frames of data is kp i -kp i-1 >Ae, the current data is considered to be an abnormal value, and kp is updated using formula (2-1) i And the value of e:
[0023]
[0024] Among them, α 1 and β 1 is a weight constant and satisfies α1 +β 1 = 1; To remove the interference of outliers, generally let α 1 be much greater than β 1 .
[0025] If kp i -kp i-1 < Ae, it is considered that the current data is normal, and only the value of e is updated, as shown in Equation (2-2):
[0026] e = α 2 e + β 2 (kp i -kp i-1 ) (2-2)
[0027] Among them, α 2 and β 2 are weight constants, and also satisfy α 2 + β 2 = 1; The difference between the two can represent the degree of emphasis on the time series characteristics of the data, and to a certain extent, eliminate the influence of historical data on the current data.
[0028] In the environment faced by the present invention, according to the results of actual data processing, A takes the value of 5, α 1 takes the value of 0.95, β 1 takes the value of 0.05, α 2 takes the value of 0.6, β 2 takes the value of 0.4, and a good filtering effect can be obtained.
[0029] The statistical outlier removal method is adopted for the three-dimensional point cloud data, and the calculation formula is as shown in Equation (2-3):
[0030]
[0031] Among them, p i is the reference point, p j is the jth nearest neighbor point of, d(p i , p j ) is the Euclidean distance between two points, μ is the average value of the average distances of all points of, n is the total number of points in the point cloud, and α is the user-defined standard deviation multiple threshold;
[0032] S23: Explore and predict human motion parameters based on the human pose estimation algorithm, including joint angles, joint movement speeds, and Euler angles, and construct a digital model of the patient during the training process from the key point detection data;
[0033] S24: Accurately project the preprocessed three-dimensional point cloud data to generate a corresponding two-dimensional point cloud map;
[0034] S25: Detect the area where the human body is located from the two-dimensional point cloud map obtained in S24. Further, locate the positions of a certain number of joint points from the human body area and connect them according to the human body structure to form a skeleton model that can reflect the motion state;
[0035] S26: Use a deep neural network through the network model DeepPose based on a convolutional neural network to predict the position information of human joint points in a single image; Detect the human body from the two-dimensional point cloud map obtained in S24 and locate the two-dimensional pixel coordinates of the joint points;
[0036] S27: Obtain higher prediction accuracy through a cascaded pose regression model with multi-stage prediction. In the first stage, extract the rough pose contour of the human body, and in subsequent stages, continuously update and optimize the information of unknown points based on the positions of known key points; The 3D point P(X, Y, Z) exists in the camera coordinate system with O as the origin, and its projection coordinates on the imaging plane with o(μ, v) as the origin are p(x, y); Define the focal length as f, the pixel width as p x , and the height as p y , f x = f / p x , f y = f / p y ; In the X-Z plane, according to the principle of similar triangles, the formula (2-4) is derived:
[0037]
[0038] In the Y-Z plane, the formula (2-5) is obtained as follows:
[0039]
[0040] According to the rotation and translation of the coordinate system, the mapping relationship from the two-dimensional pixel coordinate system to an arbitrary three-dimensional coordinate system is shown in Equation (2-6):
[0041]
[0042] Among them, is the internal parameter matrix of the camera, is the external parameter matrix; Each parameter of the matrix can be calculated through multiple groups of 2D-3D point pairs with known coordinates, thereby establishing the mapping relationship from 2D key points to 3D space;
[0043] S28: For the evaluation indicators of human pose estimation, include PCK (Percentage of Correct Keypoint) and mAP (mean Average Percision);
[0044] Taking the longest distance of the human body as the normalization standard, the threshold is set to σ, and the PCK value of the i-th human body key point is defined as PCK i , and the calculation formula is shown in Equation (2-7):
[0045]
[0046] where the total number of samples is N, y represents the predicted value, Y represents the true value, represents the longest distance from the top to the bottom of the human body; PCK can intuitively reflect the detection accuracy of key points in a single human body;
[0047] mAP is a commonly used indicator in the field of object detection. When calculating, a certain standard is first required to measure the similarity between the predicted value and the true value; in the field of human pose estimation, OKS (Object Keypoint Similarity) is usually used to calculate the similarity; mAP represents the mean value of AP values under different thresholds. If the threshold is s, the corresponding AP value is denoted as AP@s, as shown in Equation (2-8):
[0048]
[0049] where p represents the p-th human body, δ is the Kronecker function, indicating that only the number of key points with similarity greater than the threshold is counted; OKS p represents the OKS value of the p-th human body, and the calculation formula is Equation (2-9):
[0050]
[0051] where i is the number value, measured by the Euclidean distance, represents the scale factor related to the distance, is the normalization factor, and the Kronecker function δ(v pi = 1) is used to measure the visibility of key points.
[0052] More specifically, step S3 includes:
[0053] S31: Identifying the type of patient's limb movement using the information obtained by the human pose estimation algorithm;
[0054] S32: The input of the action recognition method based on the bone model is a one-dimensional vector composed of the coordinates or motion information of multiple joint points of the human bone;
[0055] S33: Using a Graph Convolutional Neural Networks (GCN) to fit the input data in the form of a graph, where the input data includes limb length, limb angle, joint linear velocity, and limb angular velocity;
[0056] The limb length is an estimate of the actual length of each limb in the upper body of the human body. The Euclidean distance between two adjacent joints is used to measure the length of a certain limb, denoted as l ab ; a and b are the key point numbers and it is stipulated that a < b; for example, l 23 refers to the Euclidean distance between key points 2 and 3, representing the length of the right upper arm.
[0057] The angle between limbs is used to measure the positional relationship between two limbs at a certain moment, and is represented by the spatial angle formed by the two, denoted as θ xyz , where x, y, and z are all key point numbers; θ 189 represents the angle between limb 1-8 and limb 8-9, and can be used to describe the inclination angle of the torso. If the spatial vector of limb 1-8 is (x 1 , y 1 , z 1 ), and the spatial vector of limb 8-9 is (x 2 , y 2 , z 2 ), the calculation formula (3-1) is obtained from the cosine theorem:
[0058]
[0059] The joint linear velocity refers to the instantaneous linear velocity of a certain key point. The mathematical form is the differential of the key point displacement with respect to time, denoted as v i , i is the key point number; it is estimated by the quotient of the displacement of the same key point in two consecutive frames of data and the time difference, and the calculation formula is as shown in Equation (3-2).
[0060]
[0061] Among them, represents the x coordinate of key point i at time t, and Δt is the time difference between two frames of data;
[0062] The limb angular velocity refers to the rotation speed of the limb during the rehabilitation training process and is an important parameter of the human limb during rotational motion; when the human body is simplified into a fulcrum (joint), connecting rod (limb) model, the angular velocity of the limb can be calculated by the connecting rod motion rule, and can be divided into two cases: circular motion around a fixed point and circular motion around a moving point; in the case of knowing the elbow joint linear velocity v 3 and the wrist joint linear velocity v 4 , the angular velocity of the limb rotating around a fixed point can be calculated by formula (3-3):
[0063] ω 23 = v 3 / l 23 (3-3)
[0064] The angular velocity of the limb around the moving point can be calculated by formula (3-4):
[0065] ω 34 =(v 4 -v 3 ) / l 34 (3-4)
[0066] It can be seen from this that after obtaining the accurate spatial positions of the key points, various basic motion parameters of the human body during static and dynamic processes can be obtained through the mathematical relationships between limb coordinates; thus, a mathematical description of the entire human body can be achieved, providing the possibility for the robot to understand the patient's behavior.
[0067] Euler angles are a combination of angles that describe the rotational relationship between two coordinate systems in three-dimensional space. Depending on the order of rotation of the axes, Euler angles have different rotation conventions, namely XYZ, XZY, YXZ, YZX, ZXY, ZYX rotating around three axes, and XYX, YXY, XZX, ZXZ, YZY, ZYZ rotating around two axes. In this paper, ZXY is used to represent Euler angles.
[0068] The original coordinate system is the XYZ coordinate system, and after rotation, it is the xyz coordinate system. The intersection line N is the intersection line of the XY plane and the xy plane; define the angle between the X axis and the N axis as α, the angle between the Z axis and the z axis as β, and the angle between the N axis and the x axis as γ. Then, (α, β, γ) is a set of Euler angles that describe the current rotation transformation;
[0069] Method for estimating the Euler angles of limb rotation through the three-dimensional coordinates of key points: According to the single-axis rotation transformation matrix in three-dimensional space, define the Euler angles corresponding to a certain rotation transformation as (α, β, γ), and obtain the rotation matrix R(α, β, γ) by multiplying the single-axis transformation matrices, as shown in formula (3-5).
[0070]
[0071] Let sinα = s 1 , sinβ = s 2 , sinγ = s 3 , cosα = c 1 , cosβ = c 2 , cosγ = c 3 , and formula (3-6) is obtained by matrix multiplication calculation:
[0072]
[0073] According to the obtained rotation matrix R(α, β, γ), calculate the Euler angles through formula (3-7):
[0074]
[0075] The rotation matrix derived from the coordinate information of the limb joints must be converted with the help of axis angles and quaternions. Consider the limb as a vector v in three-dimensional space, the rotation axis when lifting is n, the angle is θ, and the upper arm vector after rotation is v * , the axis angle is calculated by formula (3-8):
[0076]
[0077] Define the vector obtained after normalization as (a, b, c), and construct the quaternion q=w+xi+yj+zk corresponding to the rotation transformation from the axis angle θ. The mathematical relationship between the quaternion parameters and the axis angle is shown in formula (3-9):
[0078]
[0079] Finally, the rotation matrix is constructed using the quaternion rotation operation, as shown in formula (3-10):
[0080]
[0081] S34: Reversely calculate the size of the Euler angle from the rotation matrix, mathematically describe the rotation movement in three-dimensional space, and restore and reproduce it in a virtual environment, providing a way for patients to interact with the virtual environment; enable the assistive robot to understand the patient's autonomous training intention and assist the affected limb in motor training.
[0082] More specifically, step S4 includes:
[0083] S41: Use the recurrent neural network (RNN) to calculate the output value. The input data at time t is the superposition of the original data and the output data at time t-1. The calculation formula is shown in formula (4-1):
[0084] h t =tanh(W ih x t +b ih +W hh h t-1 +b hh ) (4-1)
[0085] Among them, W ih and b ih are the input-oriented weights and biases, W hh and b hh are the weights and biases for the previous state;
[0086] S42: Apply the attention mechanism in the encoder and decoder to guide the encoding and decoding process to process the important components in the original data;
[0087] The Multilayer Perceptron (MLP) is a machine learning prediction model that simulates the information processing principle of the nervous system. The MLP uses multiple layers of basic units interconnected to form a fully connected network model to achieve efficient parallel processing of input information. It has a powerful learning and generalization ability and can be used for tasks such as classification and regression. The specific form of the basic unit of the MLP is called a "neuron".
[0088] Using the multiple layers of basic units of the multilayer perceptron interconnected to form a fully connected network model, each basic unit of the multilayer perceptron obtains multiple signals x i (i = 1, 2, …, n) passed from the units of the previous layer, and sets a weight component w i (i = 1, 2, …, n) for each input. All x i multiplied by their respective weights w i and accumulated is the input value of the neuron; subsequently, the neuron adds the bias b and obtains the output value y after being processed by the activation function. The internal calculation process is represented by formula (4-2):
[0089]
[0090] where f(.) is selected from the sigmoid or tanh function, as shown in formulas (4-3) and (4-4):
[0091] sigmoid(a) = 1 / (1 + e -a ) (4-3)
[0092] tanh(a) = (e a - e -a ) / (e a + e -a ) (4-4)
[0093] S43: Use the CNN encoder to gradually obtain feature information from the input information through multiple convolutional operations, and at the same time change the convolutional kernel size and the number of receptive fields sufficient to change the convolutional operation to obtain features in different regions;
[0094] S44: After the input information is processed by two layers of 1D CNN, multi-channel spatial and short-time sequence features are extracted; subsequently, these features are input into the channel attention module to encode the attention information of each channel to reduce the resources of the model to process irrelevant channel information. The general calculation formulas of the 1D CNN convolutional layer are shown in (4-5) and (4-6),
[0095]
[0096] Define x i with a length of lx , the convolution kernel w j with a length of l w , and a stride of s, the output vector y can be obtained k with a length of l y ; when the input is a multi-channel vector composed of time series signals, the convolution operation based on the sliding window fuses the data in different regions to obtain the time series features;
[0097] S45: Using the decoder, according to the encoded features, the compressed features are restored to the length before compression through the mapping relationship; the decoded feature vector can be the same as the input or a better-quality feature with the same length as the input. After processing the compressed features, longer time series features with more classification value will be extracted, and this process is realized by the LSTM network.
[0098] S46: First, the input information is processed by the self-attention module to obtain the attention distribution of each part of the data. Subsequently, the data with attention information is decoded by two layers of LSTM networks to extract the time series information; through a sigmoid function to process the combination information of h t-1 and x t to obtain the forgetting component f t , and its product with the memory information flow determines the amount of information retained to the current moment, and this process can be shown by formula (4-7):
[0099]
[0100] where the symbols W and b refer to the weights and biases of the corresponding network layers respectively;
[0101] The input gate multiplies the input component obtained by processing the input information through the sigmoid and tanh functions respectively and then adds the output of the forgetting gate to obtain the memory amount output at the current moment, and this process is shown in formula (4-8):
[0102]
[0103] i t is the output of the sigmoid function of the input layer, c t is the output of the tanh function of the input layer, and the output value of the input layer is both the input of the output layer and the memory information retained to the next moment. The input information is multiplied by the memory information C t to obtain the output of the current state, and formula (4-9) is the operation process of the output gate:
[0104]
[0105] More specifically, the steps of S5 include:
[0106] S51: Relying on the force sensor installed on the end effector or exoskeleton to obtain the force information between the patient's arm and the actuator, it is used for the compliant interactive control of the robot-assisted training;
[0107] S52: Use wearable sensors that are independent of the robot platform, such as inertial sensors and surface electromyography sensors, to obtain the patient's motion or physiological parameters during training, so as to obtain the patient's intentions and achieve more intelligent interaction.
[0108] S53: Use the brain-computer interface system to detect the patient's EEG signals, identify the patient's movement intentions, and then control the robot to assist in completing the movement training movements of reaching, grasping and releasing. The interaction method based on physiological signals can meet the needs of patients in the flaccid paralysis period and those with insufficient muscle strength to interact with the assistive robot for autonomous intentions.
[0109] The self-attention mechanism is used to determine the attention resources that the model should allocate to different parts of the data. The input data is I. The query matrix Q, key matrix K and value matrix V are obtained through the mapping relationship. The calculation formula is shown in formula (5-1):
[0110]
[0111] Among them, W Q , W K and W V Represents the mapping relationship, which can be a function or an MLP network. The final attention value is obtained through these three matrices, as shown in formula (5-2):
[0112]
[0113] Q and K have the same dimension d k , represents the normalized attention matrix;
[0114] S54: If it is judged to be comfortable, the current posture is maintained to avoid unnecessary adjustments; if it is judged to be uncomfortable, the system actively controls the rehabilitation equipment to perform actions such as lifting, rotating or extending to adjust the patient's posture to ensure comfort and safety.
[0115] In a second aspect, the present invention proposes a planning system for a disabled assistance robot based on human posture capture, comprising:
[0116] The data acquisition and preprocessing module uses the lidar sensor to scan the surrounding environment, obtain detailed 3D point cloud data, and improve the data quality through filtering and denoising. The microphone module collects information and transmits it to the central processor for model algorithm processing;
[0117] The data analysis and path planning module is used to build a 2D map and use convex segmentation technology to divide the environment into several convex polygon areas to reduce the difficulty of path planning; then the RRT algorithm is used to search for paths in the convex polygon areas, and the optimal path is found through continuous iteration and optimization;
[0118] The state estimation and dynamic obstacle avoidance control law module is used to perform local path planning to avoid dynamic obstacles encountered through the DWA algorithm, and use the extended Kalman filter (EKF) to estimate the current state of the assistive robot to ensure that the robot executes according to the path during the path planning process.
[0119] In a third aspect, the present invention proposes a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the present invention's method for planning a disabled-assistive robot based on human posture capture.
[0120] In a fourth aspect, the present invention proposes a computing device comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the planning method of an assistive robot based on human posture capture of the present invention is implemented.
[0121] In a fifth aspect, the present invention proposes a computer program product, including a computer program, which, when executed by a processor, implements the planning method of an assistive robot based on human posture capture of the present invention.
[0122] The beneficial effects of the present invention are:
[0123] (1) Optimize the human posture estimation algorithm and build a visual motion capture platform. Aiming at the actual environment of rehabilitation training, a fast, high-precision, and stable human posture estimation algorithm is used to design a machine vision-based motion capture platform to quickly obtain the motion parameters of patients' rehabilitation training.
[0124] (2) Design a posture recognition model and an interactive method based on posture information. Based on the patient motion parameters collected by the visual motion capture platform, a method that can accurately identify the patient's exercise training posture is studied to improve the effectiveness and efficiency of rehabilitation training.
[0125] (3) Design an action recognition model and an interactive method based on action information. Based on the patient's limb action information, the action type is identified and used as the input signal of the robot arm during passive training. The affected limb is pulled to perform mirror movement training, and the patient's active training intention interaction with the assistive robot is realized, thereby increasing the patient's participation and enthusiasm. BRIEF DESCRIPTION OF THE DRAWINGS
[0126] Figure 1 The present invention is a flowchart of a method for planning a disabled-assisting robot based on human posture capture.
[0127] Figure 2 This is a schematic flow diagram of the present invention for obtaining real-time pose joint point data in step S2.
[0128] Figure 3 This is a schematic flow diagram of the present invention for real-time calculation of captured data in step S3.
[0129] Figure 4 This is a schematic flow diagram of the present invention for feature signal acquisition and recognition in step S4.
[0130] Figure 5 This is a schematic flow diagram of the present invention for voice command conversion and input in step S5.
[0131] Figure 6 This is a schematic diagram of the computing device of the present invention. Detailed implementation manners
[0132] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0133] Embodiment 1
[0134] As Figure 1 shown, this embodiment relates to a method for planning a disability assistance robot based on human pose capture, including the following steps:
[0135] S1: The patient is fixed to the system main body through a flexible fixing device to ensure the accuracy of pose capture and prevent accidental slipping;
[0136] S2: Using a high-definition RGB camera and a depth sensor, based on the human pose estimation algorithm, obtain the real-time pose joint point data of the patient, including the positions of the head, limbs, and torso;
[0137] S3: Perform real-time calculation on the captured data, refer to the pose database, and generate the most suitable initial rehabilitation pose;
[0138] S4: Use a camera, a microphone, and other sensors to collect visual, voice, and physiological signals, and adopt a convolutional neural network for feature extraction and fusion to improve the recognition accuracy;
[0139] S5: Through emotion recognition technology, combined with voice emotion analysis, judge the patient's satisfaction with the current pose, and adjust the system according to the patient's satisfaction.
[0140] In some embodiments, step S1 includes:
[0141] S11: Wait for the patient to be in position, and the system enters the standby state, ready to receive the user's instructions;
[0142] S12: The patient or user directly interacts with the system through the voice module instructions, inputs instructions to the disability assistance robot through spoken language, that is, transmits the voice signal to the processor;
[0143] S13: The preprocessing module processes the input voice features into acoustic feature vectors that the encoder can recognize, and converts the real text labels corresponding to the voice signals into text embedding sequences. After the sequences obtained by the preprocessing module are added to their respective position encodings, they enter the multi-level encoder and decoder; the specific calculation formula of the position encoding is shown in (1-1), which is used to convert the voice signal into a feature vector recognizable by the encoder.
[0144]
[0145] In some embodiments, as Figure 2 shown, step S2 includes:
[0146] S21: Enter the data acquisition stage, and use the lidar sensor to scan the surrounding environment at a set frequency range to obtain detailed three-dimensional point cloud data;
[0147] S22: Enter the data preprocessing stage, perform real-time smoothing filtering on the three-dimensional point cloud data collected in step S21 to effectively improve the data quality; considering that outliers are usually mutation values with a large difference from the actual data, data smoothing operations are performed;
[0148] Define kp i to represent the coordinate value of a certain key point at time i, e is the error threshold between two adjacent frames of data, and the initial value is set to the difference between the first two frames of data. A is the amplitude gain constant; when continuously collecting data, if the difference between two adjacent frames of data kp i -kp i-1 > Ae, it is considered that the current data is an outlier, and the values of kp i and e are updated with formula (2-1):
[0149]
[0150] where, α 1 and β 1 are weight constants and satisfy α 1 + β 1 = 1; in order to remove the interference of outliers, generally make α 1 much larger than β 1 .
[0151] If \(k_p\) i \(-k_p\) i-1 \(< A_e\), it is considered that the current data is normal, and only the value of \(e\) is updated, as shown in Equation (2-2):
[0152] \(e=\alpha\) 2 \(e + \beta\) 2 \((k_p\) i \(-k_p\) i-1 ) (2-2)
[0153] Where \(\alpha\) 2 and \(\beta\) 2 are weight constants, and also satisfy \(\alpha\) 2 + \(\beta\) 2 = 1; the difference between the two can represent the degree of emphasis on the data time series characteristics, and to a certain extent, eliminate the influence of historical data on the current data.
[0154] In the environment faced by the present invention, according to the actual data processing results, \(A\) takes the value of 5, \(\alpha\) 1 takes the value of 0.95, \(\beta\) 1 takes the value of 0.05, \(\alpha\) 2 takes the value of 0.6, \(\beta\) 2 takes the value of 0.4, and a good filtering effect can be obtained.
[0155] The statistical outlier removal method is adopted for the three-dimensional point cloud data, and the calculation formula is as shown in Equation (2-3):
[0156]
[0157] Where \(p\) i is the reference point, \(p\) j is the \(j\)th nearest neighbor point of \(p\), \(d(p\) i , \(p\) j ) is the Euclidean distance between the two points, \(\mu\) is the average value of the average distances of all points of all points, \(n\) is the total number of points in the point cloud, and \(\alpha\) is the user-defined standard deviation multiple threshold;
[0158] S23: Discuss and predict human motion parameters based on the human pose estimation algorithm, including joint angles, joint movement speeds, and Euler angles, and construct a digital model of the patient during the training process from the key point detection data;
[0159] S24: Accurately project the preprocessed three-dimensional point cloud data to generate a corresponding two-dimensional point cloud map;
[0160] S25: Detect the human body area from the two-dimensional point cloud map obtained in S24, and further locate the positions of a certain number of joint points from the human body area and connect them according to the human body structure to form a skeleton model that can reflect the motion state;
[0161] S26: Use a deep neural network through the network model DeepPose based on a convolutional neural network to predict the position information of human joint points in a single image; Detect the human body from the two-dimensional point cloud map obtained in S24 and locate the two-dimensional pixel coordinates of the joint points.
[0162] S27: Obtain higher prediction accuracy through a cascaded pose regression model with multi-stage prediction. In the first stage, extract the rough pose contour of the human body, and in the subsequent stages, continuously update and optimize the information of unknown points based on the positions of known key points; The 3D point P(X, Y, Z) exists in the camera coordinate system with O as the origin, and its projection coordinates on the imaging plane with o(μ, v) as the origin are p(x, y); Define the focal length as f, the pixel width as p x , and the height as p y , f x = f / p x , f y = f / p y ; In the X-Z plane, the formula (2-4) is derived from the principle of similar triangles:
[0163]
[0164] In the Y-Z plane, the formula (2-5) is obtained as follows:
[0165]
[0166] According to the rotation and translation of the coordinate system, the mapping relationship from the two-dimensional pixel coordinate system to an arbitrary three-dimensional coordinate system is shown in Equation (2-6):
[0167]
[0168] Among them, is the internal parameter matrix of the camera, is the external parameter matrix; Each parameter of the matrix can be calculated through multiple groups of 2D-3D point pairs with known coordinates, thereby establishing the mapping relationship from 2D key points to 3D space;
[0169] S28: For the evaluation indicators of human pose estimation, it includes PCK (Percentage of Correct Keypoint) and mAP (mean Average Percision);
[0170] Taking the longest distance of the human body as the normalization standard, setting the threshold as σ, the PCK value of the i-th human key point is defined as PCK i , and the calculation formula is shown in Equation (2-7):
[0171]
[0172] Among them, the total number of samples is N, y represents the predicted value, and Y represents the true value. represents the longest distance from the top to the bottom of the human body; PCK can intuitively reflect the detection accuracy of key points in a single human body.
[0173] mAP is a commonly used metric in the field of object detection. When calculating, a certain standard is first needed to measure the similarity between the predicted value and the true value; in the field of human pose estimation, OKS (Object Keypoint Similarity) is usually used to calculate the similarity; mAP represents the mean of AP values under different thresholds. If the threshold is s, the corresponding AP value is denoted as AP@s, as shown in Equation (2-8):
[0174]
[0175] Among them, p represents the p-th human body, and δ is the Kronecker function, indicating that only the number of key points with similarity greater than the threshold is counted; OKS p represents the OKS value of the p-th human body, and the calculation formula is Equation (2-9):
[0176]
[0177] Among them, i is the numbered value. Measured by the Euclidean distance. represents the scale factor related to the distance. is the normalization factor, and the Kronecker function δ(v pi = 1) is used to measure the visibility of key points.
[0178] In some embodiments, as Figure 3 shown, step S3 includes:
[0179] S31: Identify the types of patient limb movements using the information obtained by the human pose estimation algorithm.
[0180] S32: The action recognition method based on the bone model inputs a one-dimensional vector composed of the coordinates or motion information of multiple joint points of the human bone.
[0181] S33: Use a Graph Convolutional Neural Networks (GCN) to fit the input data in the form of a graph, and the input data includes limb length, limb angle, joint linear velocity, and limb angular velocity.
[0182] The limb length is an estimate of the actual length of each limb of the upper body of the human body. The Euclidean distance between two adjacent joints is used to measure the length of a certain limb, denoted as l ab; a and b are key point numbers and it is stipulated that a < b; for example, l 23 refers to the Euclidean distance between key points 2 and 3, representing the length of the upper right arm.
[0183] The angle between limbs is used to measure the positional relationship between two segments of limbs at a certain moment, represented by the spatial angle formed by the two, denoted as θ xyz , where x, y, and z are all key point numbers; θ 189 represents the angle between limb 1-8 and limb 8-9, which can be used to describe the inclination angle of the torso. If the spatial vector of limb 1-8 is (x 1 , y 1 , z 1 ), and the spatial vector of limb 8-9 is (x 2 , y 2 , z 2 ), the calculation formula (3-1) is obtained from the cosine theorem:
[0184]
[0185] The joint linear velocity refers to the instantaneous linear velocity of a certain key point, and its mathematical form is the differential of the key point displacement with respect to time, denoted as v i , i is the key point number; it is estimated by the quotient of the displacement of the same key point in two consecutive frames of data and the time difference, and the calculation formula is shown in Equation (3-2).
[0186]
[0187] Among them, represents the x coordinate of key point i at time t, and Δt is the time difference between two frames of data;
[0188] The limb angular velocity refers to the rotation speed of the limb during the rehabilitation training process and is an important parameter of the human limb during rotational motion; when the human body is simplified into a fulcrum (joint), connecting rod (limb) model, the angular velocity of the limb can be calculated through the connecting rod motion rule, which can be divided into two cases: circular motion around a fixed point and circular motion around a moving point; in the case of knowing the elbow joint linear velocity v 3 and the wrist joint linear velocity v 4 , the angular velocity of the limb rotating around a fixed point can be calculated by formula (3-3):
[0189] ω 23 = v 3 / l 23 (3-3)
[0190] The angular velocity of the limb rotating around a moving point can be calculated by formula (3-4):
[0191] ω 34 = (v 4 - v3 ) / l 34 (3-4)
[0192] It can be seen from this that after obtaining the accurate spatial positions of the key points, various basic motion parameters of the human body during static and dynamic processes can be obtained through the mathematical relationships between limb coordinates; furthermore, a mathematical description of the entire human body can be achieved, providing the possibility for the robot to understand the patient's behavior.
[0193] Euler angles are a combination of angles that describe the rotational relationship between two coordinate systems in three-dimensional space. According to different orders of axis rotation, Euler angles have different rotation sequences, namely XYZ, XZY, YXZ, YZX, ZXY, ZYX that rotate around three axes, and XYX, YXY, XZX, ZXZ, YZY, ZYZ that rotate around two axes. In this paper, ZXY is used to represent Euler angles.
[0194] The original coordinate system is the XYZ coordinate system, and after rotation, it is the xyz coordinate system. The intersection line N is the intersection line of the XY plane and the xy plane; define the angle between the X axis and the N axis as α, the angle between the Z axis and the z axis as β, and the angle between the N axis and the x axis as γ. Then, (α, β, γ) is a set of Euler angles that describe the current rotation transformation;
[0195] A method for estimating the Euler angles of limb rotation from the three-dimensional coordinates of key points: According to the single-axis rotation transformation matrix in three-dimensional space, define the Euler angles corresponding to a certain rotation transformation as (α, β, γ), and obtain the rotation matrix R(α, β, γ) by multiplying the single-axis transformation matrices, as shown in formula (3-5).
[0196]
[0197] Let sinα = s 1 , sinβ = s 2 , sinγ = s 3 , cosα = c 1 , cosβ = c 2 , cosγ = c 3 , and by matrix multiplication, formula (3-6) is obtained:
[0198]
[0199] According to the obtained rotation matrix R(α, β, γ), calculate the Euler angles through formula (3-7):
[0200]
[0201] To obtain the rotation matrix from the coordinate information of limb joints, the conversion must be completed with the help of axis angles and quaternions. Consider the limb as a vector v in three-dimensional space. When it is lifted, the rotation axis is n, the rotation angle is θ, and the vector of the upper arm after rotation is v* , the axis angle is calculated by formula (3-8):
[0202]
[0203] Define the vector obtained after normalization as (a, b, c), and construct the quaternion q=w+xi+yj+zk corresponding to the rotation transformation from the axis angle θ. The mathematical relationship between the quaternion parameters and the axis angle is shown in formula (3-9):
[0204]
[0205] Finally, the rotation matrix is constructed using the quaternion rotation operation, as shown in formula (3-10):
[0206]
[0207] S34: Reversely calculate the size of the Euler angle from the rotation matrix, mathematically describe the rotation movement in three-dimensional space, and restore and reproduce it in a virtual environment, providing a way for patients to interact with the virtual environment; enable the assistive robot to understand the patient's autonomous training intention and assist the affected limb in motor training.
[0208] In some embodiments, Figure 4 As shown, step S4 includes:
[0209] S41: Use the recurrent neural network (RNN) to calculate the output value. The input data at time t is the superposition of the original data and the output data at time t-1. The calculation formula is shown in formula (4-1):
[0210] h t =tanh(W ih x t +b ih +W hh h t-1 +b hh ) (4-1)
[0211] Among them, W ih and b ih are the input-oriented weights and biases, W hh and b hh are the weights and biases for the previous state;
[0212] S42: Apply the attention mechanism in the encoder and decoder to guide the encoding and decoding process to process the important components in the original data;
[0213] The Multilayer Perceptron (MLP) is a machine learning prediction model that simulates the information processing principle of the nervous system. The MLP uses multiple layers of basic units interconnected to form a fully connected network model to achieve efficient parallel processing of input information. It has a powerful learning and generalization ability and can be used for tasks such as classification and regression. The specific form of the MLP basic unit is called a "neuron".
[0214] Using the multiple layers of basic units of the multilayer perceptron interconnected to form a fully connected network model, each basic unit of the multilayer perceptron obtains multiple signals x i (i = 1, 2, …, n) passed from the units of the previous layer, and sets a weight component w i (i = 1, 2, …, n) for each input. All the x i and their respective weights w i are multiplied and accumulated, and the resulting value is the input value of the neuron. Subsequently, the neuron adds a bias b and, after being processed by an activation function, obtains an output value y. The internal calculation process is represented by formula (4-2):
[0215]
[0216] where f(.) is selected from the sigmoid or tanh functions, as shown in formulas (4-3) and (4-4):
[0217] sigmoid(a) = 1 / (1 + e -a ) (4-3)
[0218] tanh(a) = (e a - e -a ) / (e a + e -a ) (4-4)
[0219] S43: Using the CNN encoder to gradually obtain feature information from the input information through multiple convolution operations, while changing the convolution kernel size and the number of receptive fields sufficient to change the convolution operation to obtain features in different regions;
[0220] S44: After the input information is processed by two layers of 1D CNN, multi-channel spatial and short-time sequence features are extracted. Subsequently, these features are input into the channel attention module to encode the attention information of each channel to reduce the resources of the model for processing irrelevant channel information. The general formulas for the 1D CNN convolutional layer are shown in (4-5) and (4-6),
[0221]
[0222]
[0223] Define x i with length l x , and convolution kernel w j with length l w . With a stride of s, the output vector y can be obtained k with length l y ; when the input is a multi-channel vector composed of time series signals, the convolution operation based on the sliding window fuses the data in different regions to obtain the time series features accordingly;
[0224] S45: Use the decoder to restore the compressed features to the length before compression according to the encoded features through the mapping relationship; the decoded feature vector can be the same as the input or a better-quality feature with the same length as the input. After processing the compressed features, longer time series features with more classification value can be extracted, and this process is achieved by the LSTM network.
[0225] S46: First process the input information through the self-attention module to obtain the attention distribution of each part of the data. Subsequently, the data with attention information is decoded by two layers of the LSTM network to extract the time series information therein; process h t-1 and x t combined information to obtain the forgetting component f t , and the product of it and the memory information flow determines the amount of information retained to the current moment, and this process can be shown by formula (4-7):
[0226]
[0227] where the symbols W and b refer to the weights and biases of the corresponding network layers respectively;
[0228] The input gate adds the product of the input component obtained by processing the input information through the sigmoid and tanh functions respectively to the output of the forgetting gate to obtain the memory amount output at the current moment, and this process is as shown in formula (4-8):
[0229]
[0230] i t is the output of the sigmoid function of the input layer, c t is the output of the tanh function of the input layer, and the output value of the input layer is both the input of the output layer and the memory information retained to the next moment. The input information is multiplied by the memory information C t to obtain the output of the current state, and formula (4-9) is the operation process of the output gate:
[0231]
[0232] In some embodiments, such asFigure 5 As shown, step S5 includes:
[0233] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0234] S51: Relying on the force sensor installed on the end effector or exoskeleton to obtain the force information between the patient's arm and the actuator, it is used for the compliant interactive control of the robot-assisted training;
[0235] S52: Use wearable sensors that are independent of the robot platform, such as inertial sensors and surface electromyography sensors, to obtain the patient's motion or physiological parameters during training, so as to obtain the patient's intentions and achieve more intelligent interaction.
[0236] S53: Use the brain-computer interface system to detect the patient's EEG signals, identify the patient's movement intentions, and then control the robot to assist in completing the movement training movements of reaching, grasping and releasing. The interaction method based on physiological signals can meet the needs of patients in the flaccid paralysis period and those with insufficient muscle strength to interact with the assistive robot for autonomous intentions.
[0237] The self-attention mechanism is used to determine the attention resources that the model should allocate to different parts of the data. The input data is I. The query matrix Q, key matrix K and value matrix V are obtained through the mapping relationship. The calculation formula is shown in formula (5-1):
[0238]
[0239] Among them, W Q , W K and W V Represents the mapping relationship, which can be a function or an MLP network. The final attention value is obtained through these three matrices, as shown in formula (5-2):
[0240]
[0241] Q and K have the same dimension d k , represents the normalized attention matrix;
[0242] S54: If it is judged to be comfortable, maintain the current posture to avoid unnecessary adjustments; if it is judged to be uncomfortable, the system actively controls the rehabilitation device to perform actions such as lifting, rotating, or telescoping to adjust the patient's posture and ensure their comfort and safety.
[0243] Embodiment 2
[0244] This embodiment relates to a human posture capture-based disabled person assistance robot planning system for implementing the human posture capture-based disabled person assistance robot planning method described in Embodiment 1, including:
[0245] A data acquisition and preprocessing module uses a lidar sensor to scan the surrounding environment to obtain detailed three-dimensional point cloud data, and improves the data quality through steps such as filtering and denoising. The microphone module collects information and transmits it to the central processor for model algorithm processing;
[0246] A data analysis and path planning module is used to establish a 2D map and use the convex segmentation technique to divide the environment into several convex polygon regions to reduce the difficulty of path planning; subsequently, the RRT algorithm is used to search for paths in the convex polygon regions, and through continuous iteration and optimization, the optimal path is found;
[0247] A state estimation and dynamic obstacle avoidance control law module is used to avoid local path planning for dynamic obstacles encountered through the DWA algorithm, and uses an extended Kalman filter (EKF) to estimate the current state of the disabled person assistance robot to ensure that the robot executes the path during the path planning process.
[0248] Embodiment 3
[0249] This embodiment relates to a computer-readable storage medium with a program stored thereon. When the program is executed by a processor, it implements the human posture capture-based disabled person assistance robot planning method of Embodiment 1.
[0250] Embodiment 4
[0251] Referring to Figure 6 , this embodiment relates to a computing device including a memory and a processor. Among them, an executable code is stored in the memory, and when the processor executes the executable code, it implements the above-mentioned human posture capture-based disabled person assistance robot planning method.
[0252] Embodiment 5
[0253] This embodiment relates to a computer program product including a computer program. When the computer program is executed by a processor, it implements the human posture capture-based disabled person assistance robot planning method of Embodiment 1.
[0254] The above are only the preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the concept of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, several improvements and refinements made without departing from the principle of the present invention should also be regarded as the present invention.
Claims
1. A method for planning a disabled assistance robot based on human posture capture, comprising the following steps: S1: Fix the patient to the system body through a flexible fixing device; S2: Use a high-definition RGB camera and depth sensor to obtain the patient's real-time posture joint data, including the position of the head, limbs and torso, based on the human posture estimation algorithm; S3: Real-time solution of the captured data, referring to the posture database, to generate the optimal solution for the initial posture of rehabilitation; S4: Use cameras, microphones and other sensors to collect visual, voice and physiological signals, and use convolutional neural networks for feature extraction and fusion to improve recognition accuracy; S5: Through emotion recognition technology, combined with voice emotion analysis, determine the patient's satisfaction with the current posture and adjust the system according to the patient's satisfaction.
2. According to claim 1, a method for planning a disabled-assisting robot based on human posture capture is characterized in that: The S1 step specifically includes: S11: Waiting for the patient to be in place, the system enters a standby state and is ready to receive instructions from the user; S12: The patient or user directly interacts with the system through voice module commands and inputs commands to the assistive robot through spoken language, that is, the voice signal is transmitted to the processor; S13: The pre-processing module processes the input speech features into acoustic feature vectors that can be recognized by the encoder, and converts the real text labels corresponding to the speech signals into text embedding sequences. The sequences obtained by the pre-processing module are added to their respective position codes and then enter the multi-level encoder and decoder. The specific calculation formula of the position code is shown in (1-1), which is used to convert the speech signal into a feature vector that can be recognized by the encoder.
3. The method for planning a disabled-assisting robot based on human posture capture according to claim 1, characterized in that: The S2 step specifically includes: S21: Entering the data collection phase, the laser radar sensor is used to scan the surrounding environment within a set frequency range to obtain detailed three-dimensional point cloud data; S22: Entering the data preprocessing stage, performing real-time smoothing filtering processing on the three-dimensional point cloud data collected in step S21; Definition of kp i represents the coordinate value of a key point at time i, e is the error threshold between two adjacent frames of data, the initial value is set to the difference between the first two frames of data, A is the amplitude gain constant; when continuously collecting data, if the difference between two adjacent frames of data is kp i -kp i-1 >Ae, the current data is considered to be an abnormal value, and kp is updated using formula (2-1) i And the value of e: Among them, α1 and β1 are weight constants, and satisfy α1+β1=1; If kp i -kp i-1 < Ae, it is considered that the current data is normal, and only the value of e is updated, as shown in Equation (2-2): e=α2e+β2(kp i -kp i-1 ) (2-2) Among them, α2 and β2 are weight constants, which also satisfy α2+β2=1; The statistical outlier removal method is used for the three-dimensional point cloud data, and the calculation formula is shown in formula (2-3): Among them, p i is the reference point, p j is the jth nearest neighbor, d(p i ,p j ) is the Euclidean distance between two points, μ is the average distance of all points The average value of , n is the total number of points in the point cloud, and α is the user-defined standard deviation multiple threshold; S23: Based on the human posture estimation algorithm, we explore the prediction of human motion parameters, including joint angles, joint motion speeds, and Euler angles, and build a digital model of the patient during training from key point detection data; S24: accurately projecting the preprocessed three-dimensional point cloud data to generate a corresponding two-dimensional point cloud image; S25: detecting the region where the human body is located from the two-dimensional point cloud image obtained in S24, further locating a number of joint points from the human body region, and connecting them according to the human body structure to form a skeleton model that can reflect the motion state; S26: Use a deep neural network to predict the position information of human joints in a single image through the network model DeepPose based on convolutional neural network; detect the human body from the two-dimensional point cloud obtained in S24 and locate the two-dimensional pixel coordinates of the joints; S27: A cascaded posture regression model with multi-stage prediction is used to obtain higher prediction accuracy. The first stage extracts the rough posture contour of the human body, and in the subsequent stages, the information of unknown points is continuously updated and optimized based on the positions of known key points. The 3D point P (X, Y, Z) exists in the camera coordinate system with O as the origin, and its projection coordinates on the imaging plane with o (μ, v) as the origin are p (x, y); the focal length is defined as f, and the pixel width is p x , height p y , f x =f / p x , f y =f / p y ; In the XZ plane, the formula (2-4) is derived from the principle of similar triangles: In the YZ plane, formula (2-5) is obtained as follows: According to the rotation and translation of the coordinate system, the mapping relationship from the two-dimensional pixel coordinate system to any three-dimensional coordinate system is shown in formula (2-6): in, is the camera’s intrinsic parameter matrix, It is an external parameter matrix; each parameter of the matrix can be calculated through multiple sets of 2D-3D point pairs with known coordinates, thereby establishing a mapping relationship from 2D key points to 3D space; S28: Evaluation indicators for human posture estimation, including PCK and mAP; Taking the longest distance of the human body as the normalization standard, setting the threshold as σ, the PCK value of the i-th human key point is defined as PCK i , the calculation formula is shown in formula (2-7): Among them, the total number of samples is N, y represents the predicted value, and Y represents the true value. Represents the longest distance from the top to the bottom of the human body; mAP is a commonly used indicator in the field of target detection. When calculating, a certain standard is first required to measure the similarity between the predicted value and the true value. In the field of human posture estimation, OKS (Object Keypoint Similarity) is usually used to calculate the similarity. mAP represents the mean of AP values under different thresholds. If the threshold is s, the corresponding AP value is recorded as AP@s, as shown in formula (2-8): Among them, p represents the pth person, δ is the Kronecker function, which means that only the number of key points with similarity greater than the threshold is counted; OKS p Represents the OKS value of the pth person, and the calculation formula is formula (2-9): Among them, i is the number value, Measured by Euclidean distance, represents the distance-related scale factor, is the normalization factor.
4. The method for planning a disabled-assisting robot based on human posture capture according to claim 1, characterized in that: The S3 step specifically includes: S31: using the information obtained by the human posture estimation algorithm to identify the patient's limb movement type; S32: The action recognition method based on the skeleton model inputs a one-dimensional vector composed of coordinates or motion information of multiple joints of the human skeleton; S33: fitting input data in a graph form using a graph convolutional neural network, wherein the input data includes limb length, angle between limbs, joint linear velocity, and limb angular velocity; Consider the limb as a vector v in three-dimensional space. When lifting, the rotation axis is n, the rotation angle is θ, and the upper arm vector after rotation is v * , the axis angle is calculated by formula (3-8): Define the vector obtained after normalization as (a, b, c), and construct the quaternion q=w+xi+yj+zk corresponding to the rotation transformation from the axis angle θ. The mathematical relationship between the quaternion parameters and the axis angle is shown in formula (3-9): Finally, the rotation matrix is constructed using the quaternion rotation operation, as shown in formula (3-10): S34: Reversely calculate the size of the Euler angle from the rotation matrix, mathematically describe the rotation movement in three-dimensional space, and restore and reproduce it in a virtual environment, providing a way for patients to interact with the virtual environment; enable the assistive robot to understand the patient's autonomous training intention and assist the affected limb in motor training.
5. The method for planning a disabled-assisting robot based on human posture capture according to claim 1 is characterized in that: The S4 step specifically includes: S41: Calculate the output value using a recurrent neural network. The input data at time t is the superposition of the original data and the output data at time t-1. The calculation formula is shown in formula (4-1): h t =tanh(W ih x t +b ih +W hh h t-1 +b hh ) (4-1) Among them, W ih and b ih are the input-oriented weights and biases, W hh and b hh are the weights and biases for the previous state; S42: Apply the attention mechanism in the encoder and decoder to guide the encoding and decoding process to process the important components in the original data; The multi-layer basic units of the multi-layer perceptron are interconnected to form a fully connected network model. Each basic unit of the multi-layer perceptron obtains multiple signals x transmitted by the previous layer unit. i (i=1,2,…,n), and set a weight component w for each input i (i=1,2,…,n), all x i With their respective weights w i The multiplied and accumulated value is the input value of the neuron; then, the neuron superimposes the bias b and obtains the output value y after being processed by the activation function; the internal calculation process is expressed by formula (4-2): Where f(.) is selected from the sigmoid or tanh function, as shown in formulas (4-3) and (4-4): sigmoid(a)=1 / (1+e -a ) (4-3) tanh(a)=(and a -And -a ) / (And a +e -a ) (4-4) S43: Using the CNN encoder to gradually obtain feature information from the input information through multiple convolution operations, while changing the convolution kernel size and the number of receptive fields sufficient to change the convolution operation to obtain features of different regions; S44: After the input information is processed by two layers of 1DCNN, the spatial and short-term features of multiple channels are extracted; these features are then input into the channel attention module to encode the attention information of each channel; the generalized calculation formula of the 1DCNN convolutional layer is shown in (4-5) and (4-6), Define x i The length is l x , convolution kernel w j The length is l w , the step length is s, then the output vector y can be obtained k The length l y ; When the input is a multi-channel vector composed of time series signals, the convolution operation based on the sliding window fuses the data of different regions to obtain the time series features; S45: using a decoder to restore the compressed features to the length before compression through a mapping relationship according to the encoded features; S46: The input information is first processed by the self-attention module to obtain the attention distribution of each part of the data. Then the data with attention information is decoded by a two-layer LSTM network to extract the timing information; h is processed by a sigmoid function t-1 and x t Combining information to obtain the forgetting component f t , and its product with the memory information flow determines the amount of information retained to the current moment. The process can be shown by formula (4-7): Among them, the symbols W and b refer to the weight and bias of the corresponding network layer respectively; The input gate processes the input information through the sigmoid and tanh functions respectively, and then multiplies the input components obtained by adding them to the output of the forget gate to obtain the memory output at the current moment. The process is shown in formula (4-8): i t is the output of the sigmoid function of the input layer, c t is the output of the tanh function of the input layer. The input information is processed by the sigmoid function and combined with the memory information C t The multiplication is the output of the current state. Formula (4-9) is the output gate operation process:
6. The method for planning a disabled-assisting robot based on human posture capture according to claim 1, characterized in that: Step S5 specifically includes: S51: Obtaining the force information between the patient's arm and the actuator by means of a force sensor installed on the end actuator or the exoskeleton; S52: Use wearable sensors that are independent of the robot platform to obtain the patient's movement or physiological parameters while participating in training, so as to obtain the patient's intention. S53: Use the brain-computer interface system to detect the patient's EEG signals, identify the patient's movement intention, and then control the robot to assist in completing the movement training movements of reaching, grasping and releasing; The self-attention mechanism is used to determine the attention resources that the model should allocate to different parts of the data. The input data is I. The query matrix Q, key matrix K and value matrix V are obtained through the mapping relationship. The calculation formula is shown in formula (5-1): Among them, W Q , W K and W V Represents the mapping relationship, and the final attention value is obtained through these three matrices, as shown in formula (5-2): Q and K have the same dimension d k , represents the normalized attention matrix; S54: If it is judged to be comfortable, the current posture is maintained to avoid unnecessary adjustments; if it is judged to be uncomfortable, the system actively controls the rehabilitation equipment to perform actions such as lifting, rotating or extending to adjust the patient's posture to ensure comfort and safety.
7. A planning system for a disabled-assisting robot based on human posture capture, characterized in that: include: The data acquisition and preprocessing module uses the lidar sensor to scan the surrounding environment, obtain detailed 3D point cloud data, and improve the data quality through filtering and denoising. The microphone module collects information and transmits it to the central processor for model algorithm processing; The data analysis and path planning module is used to build a 2D map and divide the environment into several convex polygonal areas using convex segmentation technology; then the RRT algorithm is used to search for paths in the convex polygonal areas, and the optimal path is found through continuous iteration and optimization; The state estimation and dynamic obstacle avoidance control law module is used to perform local path planning to avoid dynamic obstacles encountered through the DWA algorithm, and use the extended Kalman filter (EKF) to estimate the current state of the assistive robot to ensure that the robot executes according to the path during the path planning process.
8. A computer-readable storage medium, characterized in that: A program is stored thereon, and when the program is executed by a processor, the method for planning a disabled-assisting robot based on human posture capture as described in any one of claims 1-6 is implemented.
9. A computing device comprising a memory and a processor, wherein: The memory stores executable codes, and when the processor executes the executable codes, the method for planning a disabled-assisting robot based on human posture capture described in any one of claims 1 to 6 is implemented.
10. A computer program product, characterized in that It includes a computer program, which, when executed by a processor, implements the method for planning a disabled-assistive robot based on human posture capture as described in any one of claims 1 to 6.
Citation Information
Cited By
Fall detection method, processing device and storage medium
CN115457659A
Rehabilitation evaluation system and method based on data feedback
CN120938361A
Human body posture recognition method and system based on double-attention structured position coding
CN121600558A
Human pose recognition method and system based on double-attention structured position encoding
CN121600558B