Vision-based intelligent prosthetic system for natural grasping

Through the head-mounted RGB-D camera and intelligent prosthetic hand system, combined with 2D object detection and 3D mapping, a gesture model library is built, which solves the problem that prosthetic hand is difficult to naturalize in a multi-object environment, and realizes efficient grasping intention recognition and anthropomorphic control.

CN116236328BActive Publication Date: 2025-08-12SHENYANG INST OF AUTOMATION - CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310149331.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-22
Publication Date
2025-08-12
Estimated Expiration
2043-02-22

AI Technical Summary

Technical Problem

The prior art is difficult to achieve natural grasping of prosthetic hands, especially in a multi-object environment, which is difficult to accurately identify the grasping intention and perform anthropomorphic control.

Method used

A head-mounted RGB-D camera is used to collect visual images, combine the grasping gesture learning function module and the five-finger dexterity hand control module to build a gesture model library through 2D object detection, 3D mapping and gesture estimation to realize the natural capture of prosthetic hands.

Benefits of technology

The crawling success rate in a single object environment reached 95.43%, the crawling time was close to manpower, the intent estimate accuracy rate in a multi-object environment was 94.34%, and the crawling success rate was 88.75%, achieving fast, natural and accurate crawling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116236328B_ABST
    Figure CN116236328B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of intelligent prosthetic control technology, and in particular to a vision-based intelligent prosthetic system that can achieve natural grasping. The present invention proposes a powered prosthetic hand system with vision, which includes an RGB-D camera worn on the head and a powered prosthetic hand. The system includes two sub-functional modules: a grasping gesture learning functional module and a prosthetic hand control functional module. The grasping gesture learning functional module is used to learn gesture data of the human hand grasping an object and generate a gesture model library, which provides a motion trajectory reference for the control of the prosthetic hand. The prosthetic hand control functional module is used to control the prosthetic hand to achieve autonomous, natural and effective grasping.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent prosthetic control, and in particular to a vision-based intelligent prosthetic system capable of achieving natural grasping. Background Art

[0002] Healthy people rely on their hands to manipulate objects in daily life. For hand amputees, the loss of their hands makes it difficult to effectively manipulate objects, which in turn impacts their daily lives. Having a five-finger powered prosthetic hand that can naturally grasp a variety of everyday objects in an anthropomorphic manner is a long-cherished dream of hand amputees.

[0003] The motion control of a prosthetic hand depends on the amputee's intention to move. Currently, common intention signals used to control a prosthetic hand include electroencephalogram (EEG) and electromyogram (EMG). The human hand's movement intention is generated in the brain. Amputees use a brain-computer interface (BCI) to extract EEG from the cerebral cortex to identify movement intentions and thus control their prosthetic hand. The signal quality of BCI directly affects the accuracy of intention recognition, and it includes two basic forms: embedded EEG interface and wearable EEG interface. Although the signal quality of embedded EEG interface is better than that of wearable EEG interface, it still has problems such as poor signal-to-noise ratio, susceptibility to environmental influences, and low intention resolution. Therefore, the current EEG-based method is difficult to achieve effective grasping of a prosthetic hand.

[0004] Using electromyographic signals from the residual limb to control a prosthetic hand is a common method. The earliest related research used electromyographic thresholds to control simple opening and closing movements of a prosthetic hand. Subsequently, researchers have proposed pattern recognition methods to achieve refined motion control using electromyographic signals, using machine learning algorithms or deep neural networks to identify gestures. However, these methods merely classify finger movements, rather than controlling movements during grasping. Furthermore, these methods, when used for prosthetic hand control, face challenges such as large individual electromyographic signal variability, muscle atrophy in amputees, signal interference caused by sweat, and electrode position shifts. Therefore, achieving accurate grasping of diverse objects with an EMG-based prosthetic hand presents significant challenges.

[0005] When people grasp an object, they adjust their grasping gesture based on the object's shape, size, and other information. In order for a prosthetic hand to adapt to grasping different objects, its control also needs to rely on the characteristic information of the grasped object. To obtain object features, researchers have proposed installing visual sensors on prosthetic hands. However, the field of view that this visual system can capture is very limited, and it can only capture local features. In order to obtain global environmental information, the method of wearing visual sensors on the head or eyes has been proposed. However, these research methods only focus on the results of grasping, directly controlling the movement of the prosthetic hand to the pre-grasping gesture, and ignoring the naturalness of the grasping process. Summary of the Invention

[0006] In view of the shortcomings of the existing technology, the present invention provides a vision-based intelligent prosthetic system that can achieve natural grasping, including:

[0007] RGB-D camera, used to capture visual images: multiple frames of images of a human hand grasping an object during training and learning, and multiple frames of images of a five-fingered dexterous hand grasping an object during real-time control;

[0008] The controller includes a processor and a memory. The processor stores program modules: a grasping gesture learning function module and a five-finger dexterous hand control function module. The processor loads and executes the program modules. The grasping gesture learning function module is used to learn the spatial sequence of gestures used by a human hand to grasp different objects and generate a gesture model library for different objects. The five-finger dexterous hand control function module is used to control the movement of the five-finger dexterous hand using the gesture model library data as a tracking target and combining real-time object and five-finger dexterous hand state information.

[0009] The five-finger dexterous hand is a humanoid five-finger robotic hand that is used to achieve natural grasping according to the instructions of the controller.

[0010] The RGB-D camera is worn on the human head, so that the camera can obtain global environmental information from a first-person perspective.

[0011] The five-finger dexterous hand has six driving degrees of freedom: the thumb has two driving degrees of freedom, which are used to drive the bending and rotation of the thumb respectively, and the other four fingers each have one driving degree of freedom, which controls the bending movement of each finger respectively.

[0012] The grabbing gesture learning function module includes the following program function modules:

[0013] 1) 2D Object Detection Module: (i) Object Detection: Identify the object category in the image space and locate the object's 2D coordinate position based on the object's bounding box; (ii) Hand Detection: Locate the bounding box of the hand region in the image space and estimate the hand's 2D coordinate position;

[0014] 2) 3D Mapping and Gesture Estimation Module: (i) 3D spatial mapping: maps 2D positions to 3D space, constructs a 3D model of the object, and calculates the 3D spatial position relationship between the hand and the target object and the 3D size of the object; (ii) Gesture Estimation: estimates the posture changes during hand movement;

[0015] 3) Gesture post-processing module: processes the changes in hand posture when fitting the human hand grasping different target objects and generates gesture functions;

[0016] 4) Constructing a gesture model library module: Constructing a gesture model library based on the process of human hand grasping objects:

[0017]

[0018] Among them, C obj is the category of the object, is the size parameter of the object, and F(D) is the grasping gesture sequence function obtained through regression and fitting.

[0019] The object detection in the 2D target detection module adopts the YOLOv5 lightweight target detection network model, takes the 2D RGB image as the input of the model, and obtains the object category C obj and bounding box B obj , and use the center point of the bounding box as the center 2D position of the object

[0020] The hand detection in the 2D object detection module uses the SRHandNet real-time network model, which regresses the bounding box B of the hand area in the RGB image. h , and estimate the 2D position of the wrist keypoints

[0021] In the 3D mapping module, the depth image and the RGB image are first aligned, and then the and Calculate the position of the object and wrist key points in the world coordinate system and The formula is as follows:

[0022]

[0023] Among them, x, y, z are or The coordinate value in the world coordinate system, h is the depth value of the corresponding point in the depth image, (u, v) is the coordinate of the corresponding point in the pixel coordinate system, (p x ,p y ) is the coordinate of the principal point, f x and f y are the focal lengths of the camera on the x-axis and y-axis respectively;

[0024] The gesture estimation in the 3D mapping and gesture estimation module is used to estimate the object bounding box B obj In the region, the target object is mapped to the three-dimensional space by equation (1), and the 3D model of the object is obtained after removing the background and outliers, and then its size parameters are obtained. It is also used to determine the hand bounding box B h Locate the hand area and use the IntagHand model to obtain the hand mesh vertex V h , build a mano hand model.

[0025] The gesture post-processing is used to smooth the estimated gesture motion and control the prosthetic hand, and includes performing the following three steps on the data of the grasping process:

[0026] 1) Key point regression of the hand: Using the joint regression method of the mano hand model, the vertex V h The regression results in 21 node coordinates;

[0027] 2) Posture calculation: Since the prosthetic hand is controlled based on the angle between the metacarpophalangeal joint-fingertip line and the reference plane or reference axis, the gesture is mapped to an angle vector α = [α p ,α r ,α m ,α i ,α tb ,α tr ];

[0028] where α p ,α r ,α m ,α i ,α tb are the bending angles of the little finger, ring finger, middle finger, index finger and thumb, α tr is the thumb rotation angle; α p ,α r ,α m ,α i ,α tr The reference surface for calculation is the palm plane, which is fitted by key points 0, 5, 9, 13 and 17; α tb The reference axis is the thumb rotation axis, which passes through key points 1 and 5;

[0029] 3) Gesture sequence fitting: To measure the spatial position relationship between the object and the hand, the Euclidean distance sequence between the hand and the object during the grasping process is defined as D

[0030]

[0031] Define the gesture sequence F(D) during the grasping process = [F p ,F r ,F m ,F i ,F tb ,F tr ]; among them, F p ,F r ,F m ,F i ,F tb ,F tr is the polynomial fitting function of the angle vector α with respect to the distance D;

[0032] The most ideal gesture transformation curve can be obtained by using fourth-order polynomial fitting. The formula is as follows:

[0033] F j (D) = a 4,j D 4 +a 3,j D 3 +a 2,j D 2 +a 1,j D 1 +a 0,j (3) where a 0,j ,a 1,j ,a 2,j ,a 3,j ,a 4,j are polynomial coefficients, j∈{p,r,m,i,tb,tr} are the same;

[0034] The gesture function F(D) is obtained through regression and fitting, which is used to express the continuous and smooth gesture transformation process.

[0035] The five-finger dexterous hand control function module includes the following program modules:

[0036] 1) 2D object detection and 3D spatial mapping module: used to calculate the category C of the object during real-time grasping tgt ,Location and size parameters and a 3D spatial model of the prosthetic hand;

[0037] 2) Grasping intention estimation module: In this stage, the joints of the prosthetic hand are stationary, and the system determines the subject's grasping intention towards an object in the current field of view and determines the target object to be pre-grasped;

[0038] 3) Real-time grasping gesture control module: After determining the target object, as the subject moves the prosthetic hand towards the object, the system performs anthropomorphic gesture control on the prosthetic hand until it grasps the object.

[0039] The intention to grasp the object is estimated by collecting the position of each prosthetic wrist. Where n ≥ 3; a spatial straight line is regressed based on these wrist positions to predict the movement direction of the prosthetic hand and thus estimate the grasping intention;

[0040] The wrist of the prosthesis is labeled with white markers. In three-dimensional Euclidean space, since a spatial straight line is obtained by the intersection of two planes, two plane equations are constructed to represent the regression line.

[0041]

[0042] For plane 1, the residual squared loss Loss is defined as

[0043]

[0044] Among them, the weight of plane 1 w=[w1,w2,w0] T , A i =[y i ,z i ,1] T ,A=[A1,A2,…,A n ] T ,x=[x1,x2,…,x n ] T , let the partial derivative of the loss function Loss with respect to w be equal to zero

[0045]

[0046] When w satisfies equation (7), the loss function Loss reaches its minimum value

[0047] w=(A T A) -1 A T x (8)

[0048] A T A is a full rank matrix; w is the parameter of plane 1, which represents the plane closest to the sample point in the x-axis direction; similarly, the parameter w of plane 2 can be obtained ′ =[w1 ′ ,w2 ′ ,w0 ′ ] T , Plane 2 represents the plane closest to the sample point in the y-axis direction;

[0049] Combine the two plane regression models to obtain the regression line;

[0050] The prosthetic hand is defined as a right-handed hand. When grasping an object, the subject usually moves to the right of the target object; that is, the object to be grasped is located on the left side of the regression line, and the distance between the target object and the regression line is the shortest.

[0051] The regression line and the vector in the positive direction of the oy axis form a space dividing surface, whose plane parameter is w s =[w x ,w y ,w z ,w d ] T (w x <0); then when the object is in the left space, the following equation is satisfied

[0052]

[0053] where p = [x p ,y p ,z p ,1] T is the homogeneous position vector of the object; when there are multiple objects on the left side of the spatial segmentation plane, the object closest to the regression line in the left space is taken as the target object, and then the object that the subject intends to grasp is determined; the 2D target detection and 3D space mapping modules are called to obtain the category C of the target object tgt ,Location and size parameters

[0054] The grasping gesture control module calls the corresponding gesture function F(D) in the gesture model library for control according to the spatial distance D between the object and the prosthetic hand; the prosthetic hand is equipped with a metal strain force sensor on the push rod of each degree of freedom to detect the grasping force.

[0055] The present invention has the following beneficial effects and advantages:

[0056] 1. This invention proposes a prosthetic hand system with head-mounted vision. The system's visual sensors are located on the head, close to the human eye, enabling global positioning and visual servoing feedback capabilities similar to those of the human eye. This vision-based prosthetic hand can obtain information about the position and shape of objects to be grasped in advance, enabling appropriate, real-time motion control.

[0057] 2. This paper proposes a novel vision-based natural gesture learning method. The grasping gesture learning module constructed using this method can learn the natural gesture transitions of healthy individuals. It is highly scalable and adaptable to a wide variety of grasping gestures, overcoming the limitations of previous single, limited grasping models.

[0058] 3. This invention proposes an anthropomorphic gesture control method for a prosthetic hand. This method derives grasping gestures from the healthy hand's grasping process and applies them to the motion control of the prosthetic hand. This method enables the prosthetic hand to quickly achieve anthropomorphic grasping of novel objects, enabling a wide range of grasping gestures tailored to the specific object being grasped.

[0059] 4. The present invention proposes a vision-based grasping intention estimation method, through which the system can identify the human body's grasping intention in a multi-object environment, determine the target object to be pre-grasped, and then call the corresponding gesture data from the gesture model library to complete the grasping.

[0060] 5. The system constructed by the present invention can simultaneously take into account the speed, naturalness, accuracy of intention estimation and grasping success rate of grasping. In the single-object environment, the grasping success rate is 95.43%, and the grasping time is 3.07±0.41s, which is close to the time it takes for a human hand to grasp an object normally. The similarity between the grasping action of the prosthetic hand and the human hand (determination coefficient R2 ) is 0.911. In a multi-object environment, the accuracy of intention estimation is 94.34%, and the grasping success rate is 88.75%. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 is a schematic diagram of the intelligent prosthetic hand system of the present invention;

[0062] Figure 2 It is a schematic diagram of the intention estimation of the present invention. DETAILED DESCRIPTION

[0063] To make the above-mentioned objects, features, and advantages of the present invention more readily apparent, the specific implementation methods of the present invention are described in detail below with reference to the accompanying drawings. The following description sets forth many specific details to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways than those described herein, and those skilled in the art may make similar modifications without departing from the scope of the invention. Therefore, the present invention is not limited to the specific implementation methods disclosed below.

[0064] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art of the art to which the present invention pertains. The terms used in the specification of the invention herein are for the purpose of describing specific embodiments only and are not intended to limit the present invention.

[0065] The present invention proposes a vision-based grasping intention estimation, which enables the system to identify the human body's grasping intention in a multi-object environment, determine the target object to be pre-grasped, and then call the corresponding gesture data from the gesture model library to complete the grasping.

[0066] The hardware of the powered prosthetic hand system with vision consists of three parts (such as Figure 1 As shown in the figure, the upper part is the workflow of the grasping gesture learning function module, and the lower part is the workflow of the prosthetic hand control function module): an RGB-D camera worn on the human head, a controller, and a five-finger dexterous hand. The RGB-D camera is worn on the human head, allowing the camera to obtain global environmental information from a first-person perspective: including one or more objects to be grasped, the human hand during training and learning, and the five-finger dexterous hand during real-time control. The size of the dexterous hand is comparable to that of a human hand, and it has 6 actuated degrees of freedom (DOF). The thumb has 2 actuated degrees of freedom, which are used to drive the bending and rotational movements of the thumb respectively. The other four fingers each have 1 actuated degree of freedom, which controls the bending movement of each finger.

[0067] The system consists of two sub-functional modules, such as Figure 1The following figure shows the grasping gesture learning module and the prosthetic hand control module. The grasping gesture learning module learns the spatial sequence of human hand gestures used to grasp different objects, ultimately generating a library of gesture models for different objects, providing motion trajectory references for prosthetic hand control. The prosthetic hand control module tracks the gesture model library data and combines real-time object and prosthetic hand state information to control the prosthetic hand's motion and achieve natural grasping.

[0068] 1. Grasping gesture learning function module

[0069] The present invention takes the right prosthesis as an example for research. The data of the grasping gesture learning module comes from the right hand of a healthy person grasping various objects. The processing flow of the grasping gesture learning module is as follows: Figure 1 As shown in Figure 2, the camera coordinate system is used as the world coordinate system during data processing.

[0070] 1.1 2D Object Detection Module

[0071] Object detection uses the YOLOv5 lightweight target detection network model. Using a 2D RGB image as the input of the model, the object category C can be obtained. obj and bounding box B obj The center point of the bounding box is then used as the center 2D position of the object.

[0072] Hand detection uses the SRHandNet real-time network model. This model regresses the bounding box B of the hand area in the RGB image. h , and estimate the 2D position of the wrist keypoints

[0073] 1.2 3D Mapping and Gesture Estimation Module

[0074] First, align the depth image with the RGB image, and then and Calculate the position of the object and wrist key points in the world coordinate system and The formula is as follows:

[0075]

[0076] Among them, x, y, z are or The coordinate value in the world coordinate system, h is the depth value of the corresponding point in the depth image, (u, v) is the coordinate of the corresponding point in the pixel coordinate system, (p x ,p y ) is the principal point coordinate (the intersection of the camera optical axis and the imaging plane), f x and f yare the focal lengths of the camera on the x-axis and y-axis respectively.

[0077] In the object bounding box B obj In the region, the target object is mapped to three-dimensional space by equation 1, and the 3D model of the object is obtained after removing the background and outliers, and then its size parameters are obtained. According to the hand bounding box B h The hand area is located, and 778 hand mesh vertices V are obtained using the IntagHand model. h , and then a mano hand model can be constructed. The hand model rendered when a human hand grasps an object is as follows Figure 1 shown.

[0078] 1.3 Gesture post-processing module

[0079] In order to make the estimated gesture smooth and use it for prosthetic hand control, it is necessary to perform key point regression, posture calculation and gesture sequence fitting on the data of the grasping process.

[0080] 1.3.1 Key point regression of the hand: Using the joint regression method of the mano hand model, the vertex V h The regression results in 21 node coordinates. The node positions are as follows: Figure 1 As shown in the figure, the red marked points on the hand model are the positions of the corresponding nodes.

[0081] 1.3.2 Posture calculation: Since the prosthetic hand is controlled based on the angle between the metacarpophalangeal joint-fingertip line and the reference plane or reference axis, the gesture is mapped to an angle vector α = [α p ,α r ,α m ,α i ,α tb ,α tr ], where α p ,α r ,α m ,α i ,α tb are the bending angles of the little finger, ring finger, middle finger, index finger and thumb, α tr α is the rotation angle of the thumb. p ,α r ,α m ,α i ,α tr The reference surface used for calculation is the palm plane, which is fitted by key points 0, 5, 9, 13 and 17. tb The reference axis is the thumb rotation axis, which passes through key points 1 and 5.

[0082] 1.3.3 Gesture sequence fitting: In order to measure the spatial position relationship between the object and the hand, the Euclidean distance sequence between the hand and the object during the grasping process is defined as D

[0083]

[0084] The gesture sequence in the grasping process is defined as F(D) = [F p ,F r ,F m ,F i ,F tb ,F tr ]. Among them, F p ,F r ,F m ,F i ,F tb ,F tr is the polynomial fitting function of the angle vector α with respect to the distance D. The fourth-order polynomial fitting can obtain the most ideal gesture transformation curve, and the formula is as follows:

[0085] F j (D) = a 4,j D 4 +a 3,j D 3 +a 2,j D 2 +a 1,j D 1 +a 0,j (3)

[0086] where a 0,j ,a 1,j ,a 2,j ,a 3,j ,a 4,j are polynomial coefficients, j∈{p,r,m,i,tb,tr}. The gesture function F(D) obtained through regression and fitting can express the continuous and smooth gesture transformation process.

[0087] 1.4 Gesture model library construction

[0088] A gesture model library is built based on the process of human hand grasping objects. The data representation of the model library is as follows:

[0089]

[0090] The model library is scalable and can construct different gesture data according to different grasped objects and size parameters.

[0091] 2. Prosthetic hand control function module

[0092] Prosthetic hand control process Figure 1As shown in Figure 1, the control process is divided into two stages: 1) intention estimation stage: in this stage, the joints of the prosthetic hand are stationary, and the system determines the subject's intention to grasp an object, especially in a multi-object environment, and determines the target object to be pre-grasped; 2) grasping stage: after determining the target object, as the prosthetic hand approaches the object, it is controlled by anthropomorphic gestures until the object is grasped.

[0093] 2.1 Estimation of Grasping Intention

[0094] When the human body grasps an object, the motion curve of the wrist in space can usually be approximated as a straight line. Based on this principle, we will collect the position of each prosthetic wrist Where n ≥ 3. Based on these wrist positions, a spatial straight line is regressed to predict the movement direction of the prosthetic hand, thereby realizing the estimation of the grasping intention, such as Figure 2 As shown. In order to facilitate the system to locate the position of the prosthetic wrist, we put white markers on its wrist. In three-dimensional Euclidean space, since a spatial straight line is obtained by the intersection of two planes, two plane equations are constructed to represent the regression line.

[0095]

[0096] For plane 1, the residual squared loss Loss is defined as

[0097]

[0098] Among them, the weight of plane 1 w=[w1,w2,w0] T , A i =[y i ,z i ,1] T ,A=[A1,A2,…,A n ] T ,x=[x1,x2,…,x n ] T , let the partial derivative of the loss function Loss with respect to w be equal to zero

[0099]

[0100] When w satisfies Equation 7, the loss function Loss reaches its minimum value

[0101] w=(A T A) -1 A T x (8)

[0102] A T A is a full rank matrix. w is the parameter of plane 1, which represents the plane closest to the sample point in the x-axis direction. Similarly, the parameter w of plane 2 can be obtained′ =[w1 1 ,w2 ′ ,w0 ′ ] T , plane 2 represents the plane closest to the sample point in the y-axis direction. By combining the two plane regression models, we can get the regression line.

[0103] The present invention uses a right prosthetic hand, and the subject usually moves to the right of the target object when grasping an object. That is, the object to be grasped is located on the left side of the regression line, and the distance between the target object and the regression line is the shortest.

[0104] The regression line and the vector in the positive direction of the oy axis form a space dividing surface, whose plane parameter is w s =[w x ,w y ,w z ,w d ] T (w x <0). Then when the object is in the left space, the following equation is satisfied:

[0105]

[0106] where p = [x p ,y p ,z p ,1] T is the homogeneous position vector of the object. When there are multiple objects on the left side of the spatial segmentation plane, the object closest to the regression line in the left space is taken as the target object, and the category of the target object during real-time grasping is obtained. Location and size parameters Then compare it with the gesture model library to determine the type, size and gesture function of the object.

[0107] 2.2 Gesture Control of Grasping

[0108] For the prosthetic hand system, when the type of object C is determined obj and size Then, the corresponding gesture function F(D) in the gesture model library is called for control based on the spatial distance D between the object and the hand. The prosthetic hand integrates metal strain sensors on the push rods of each degree of freedom to detect grip force. A grip force threshold is set to ensure effective grip. At the end of the prosthetic hand's grasping process, two situations may occur: 1) The angle preset by the gesture function is not reached, but the current threshold is reached. In this case, the angle of that degree of freedom is locked; 2) The final angle preset by the gesture function is reached, but the preset current threshold is not reached. In this case, the angle of the corresponding degree of freedom is further contracted until the current threshold or angle contraction threshold is reached.

[0109] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should be regarded as within the scope of protection of the present invention.

Claims

1. A vision-based intelligent prosthetic system that enables natural grasping, characterized by: include: RGB-D camera, used to capture visual images: multiple frames of images of a human hand grasping an object during training and learning, and multiple frames of images of a five-fingered dexterous hand grasping an object during real-time control; The controller includes a processor and a memory. The processor stores program modules: a grasping gesture learning function module and a five-finger dexterous hand control function module. The processor loads the program modules and executes them. The grasping gesture learning function module is used to learn the spatial sequence of gestures used by a human hand to grasp different objects and generate a gesture model library for different objects. The five-finger dexterous hand control function module is used to control the movement of the five-finger dexterous hand using the gesture model library data as a tracking target and combining real-time object and five-finger dexterous hand state information. The grasping gesture learning function module includes the following program function modules: a 2D target detection module, a 3D mapping and gesture estimation module, a gesture post-processing module, and a gesture model library construction module. The gesture post-processing module is used to smooth the estimated gesture movement and control the prosthetic hand, including performing the following three steps on the data of the grasping process: 1) Key point regression of the hand: Using the joint regression method of the mano hand model, the vertex V h The regression results in 21 node coordinates; 2) Posture calculation: Since the prosthetic hand is controlled based on the angle between the metacarpophalangeal joint-fingertip line and the reference plane or reference axis, the gesture is mapped to an angle vector α = [α p ,α r ,α m ,α i ,α tb ,α tr ]; where α p ,α r ,α m ,α i ,α tb are the bending angles of the little finger, ring finger, middle finger, index finger and thumb, α tr is the thumb rotation angle; α p ,α r ,α m ,α i ,α tr The reference surface for calculation is the palm plane, which is fitted by key points 0, 5, 9, 13 and 17; α tb The reference axis is the thumb rotation axis, which passes through key points 1 and 5; 3) Gesture sequence fitting: To measure the spatial position relationship between the object and the hand, the Euclidean distance sequence between the hand and the object during the grasping process is defined as D: Define the gesture sequence F(D) during the grasping process = [F p ,F r ,F m ,F i ,F tb ,F tr ]; among them, F p ,F r ,F m ,F i ,F tb ,F tr is the polynomial fitting function of the angle vector α with respect to the distance D; The most ideal gesture transformation curve can be obtained by using fourth-order polynomial fitting. The formula is as follows: F j (D)=a 4,j D 4 +a 3,j D 3 +a 2,j D 2 +a 1,j D 1 +a 0,j (3) where a 0,j ,a 1,j ,a 2,j ,a 3,j ,a 4,j are polynomial coefficients, j∈{p,r,m,i,tb,tr} are the same; The gesture function F(D) is obtained through regression and fitting, which is used to express the continuous and smooth gesture transformation process; the five-finger dexterous hand is a humanoid five-finger robotic hand that is used to achieve natural grasping according to the instructions of the controller.

2. The vision-based intelligent prosthetic system capable of natural grasping according to claim 1, characterized in that: The RGB-D camera is worn on the human head, so that the camera can obtain global environmental information from a first-person perspective.

3. The vision-based intelligent prosthetic system capable of natural grasping according to claim 1, characterized in that: The five-finger dexterous hand has six driving degrees of freedom: the thumb has two driving degrees of freedom, which are used to drive the bending and rotation of the thumb respectively, and the other four fingers each have one driving degree of freedom, which controls the bending movement of each finger respectively.

4. The vision-based intelligent prosthetic system capable of natural grasping according to claim 1, characterized in that: The grabbing gesture learning function module includes the following program function modules: 1) 2D Object Detection Module: (i) Object Detection: Identify the object category in the image space and locate the object's 2D coordinate position based on the object's bounding box; (ii) Hand Detection: Locate the bounding box of the hand region in the image space and estimate the hand's 2D coordinate position; 2) 3D Mapping and Gesture Estimation Module: (i) 3D spatial mapping: maps 2D positions to 3D space, constructs a 3D model of the object, and calculates the 3D spatial position relationship between the hand and the target object and the 3D size of the object; (ii) Gesture Estimation: estimates the posture changes during hand movement; 3) Gesture post-processing module: processes the changes in hand posture when fitting the human hand grasping different target objects and generates gesture functions; 4) Constructing a gesture model library module: Constructing a gesture model library based on the process of human hand grasping objects: Among them, C obj is the category of the object, is the size parameter of the object, and F(D) is the grasping gesture sequence function obtained through regression and fitting.

5. The vision-based intelligent prosthetic system capable of natural grasping according to claim 4, characterized in that: The object detection in the 2D target detection module adopts the YOLOv5 lightweight target detection network model, takes the 2D RGB image as the input of the model, and obtains the object category C obj and bounding box B obj , and use the center point of the bounding box as the center 2D position of the object The hand detection in the 2D object detection module uses the SRHandNet real-time network model, which regresses the bounding box B of the hand area in the RGB image. h , and estimate the 2D position of the wrist keypoints 6. The vision-based intelligent prosthetic system capable of natural grasping according to claim 4, characterized in that: In the 3D mapping module, the depth image and the RGB image are first aligned, and then the and Calculate the position of the object and wrist key points in the world coordinate system and The formula is as follows: Among them, x, y, z are or The coordinate value in the world coordinate system, h is the depth value of the corresponding point in the depth image, (u, v) is the coordinate of the corresponding point in the pixel coordinate system, (p x ,p y ) is the coordinate of the principal point, f x and f y are the focal lengths of the camera on the x-axis and y-axis respectively; The gesture estimation in the 3D mapping and gesture estimation module is used to estimate the object bounding box B obj In the region, the target object is mapped to the three-dimensional space by equation (1), and the 3D model of the object is obtained after removing the background and outliers, and then its size parameters are obtained. It is also used to determine the hand bounding box B h Locate the hand area and use the IntagHand model to obtain the hand mesh vertex V h , build a mano hand model.

7. The vision-based intelligent prosthetic system capable of natural grasping according to claim 1, characterized in that: The five-finger dexterous hand control function module includes the following program modules: 1) 2D object detection and 3D spatial mapping module: used to calculate the category C of the object during real-time grasping tgt ,Location and size parameters and a 3D spatial model of the prosthetic hand; 2) Grasping intention estimation module: When all joints of the prosthetic hand are stationary, the system determines the subject's grasping intention towards an object in the current field of view and determines the target object to be pre-grasped; 3) Real-time grasping gesture control module: After determining the target object, as the subject moves the prosthetic hand towards the object, the system performs anthropomorphic gesture control on the prosthetic hand until it grasps the object.

8. The vision-based intelligent prosthetic system capable of natural grasping according to claim 7, characterized in that: The intention to grasp the object is estimated by collecting the position of each prosthetic wrist. Where n ≥ 3; a spatial straight line is regressed based on these wrist positions to predict the movement direction of the prosthetic hand and thus estimate the grasping intention; The wrist of the prosthetic limb is affixed with white marking points. In three-dimensional Euclidean space, since a spatial straight line is obtained by the intersection of two planes, two plane equations are constructed to represent the regression line. For plane 1, the residual squared loss Loss is defined as Among them, the weight of plane 1 w=[w1,w2,w0] T , A i =[y i ,z i ,1] T ,A=[A1,A2,…,A n ] T ,x=[x1,x2,…,x n ] T , let the partial derivative of the loss function Loss with respect to w be zero; When w satisfies equation (7), the loss function Loss reaches its minimum value; w=(A T A) -1 A T x (8) A T A is a full rank matrix; w is the parameter of plane 1, which represents the plane closest to the sample point in the x-axis direction; similarly, the parameter of plane 2 is w′=[w′1,w′2,w′0] T , Plane 2 represents the plane closest to the sample point in the y-axis direction; Combine the two plane regression models to obtain the regression line; The prosthetic hand is defined as a right-handed hand. When grasping an object, the subject usually moves to the right of the target object; that is, the object to be grasped is located on the left side of the regression line, and the distance between the target object and the regression line is the shortest. The regression line and the vector in the positive direction of the oy axis form a space dividing surface, whose plane parameter is w s =[w x ,w y ,w z ,w d ] T (w x <0); then when the object is in the left space, the following equation is satisfied: where p = [x p ,y p ,z p ,1] T is the homogeneous position vector of the object; when there are multiple objects on the left side of the spatial segmentation plane, the object closest to the regression line in the left space is taken as the target object, and then the object that the subject intends to grasp is determined; the 2D target detection and 3D space mapping modules are called to obtain the category C of the target object tgt ,Location and size parameters 9. The vision-based intelligent prosthetic system capable of natural grasping according to claim 7, characterized in that: The grasping gesture control module calls the corresponding gesture function F(D) in the gesture model library for control according to the spatial distance D between the object and the prosthetic hand; the prosthetic hand is equipped with a metal strain force sensor on the push rod of each degree of freedom to detect the grasping force.

Citation Information

Patent Citations

  • Method for positioning and grabbing irregular workpiece based on single-frame RGB-D image deep learning

    CN111553949A

  • Methods and apparatus for human centric "hyper UI for devices"architecture that could serve as an integration point with multiple target / endpoints (devices) and related methods / system with dynamic context aware gesture input towards a "modular" universal controller platform and input device virtualization

    WO2016189372A2