Full-automatic mouth following meal assisting method and system based on machine vision
By improving the direct linear transformation algorithm and lightweight YOLOv5n optimization model, the problem that the meal aid robot cannot accurately estimate the mouth position is solved, real-time dynamic follow-up and delivery of the robot arm, improving the accuracy and user experience of feeding.
Patent Information
- Application Number
- CN202510822808.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-19
AI Technical Summary
The existing meal aid robots cannot accurately estimate the mouth position, resulting in low accuracy, poor flexibility, and complex operation, which is inconvenient for users with limited mobility.
By improving the direct linear transformation algorithm, the mouth position of the key points of the two-dimensional image of the face is obtained, combining the orthogonal constraints between the rotation vectors and the singular value decomposition strategy, a robotic arm base coordinate system is established, and the mouth movements are monitored in real time, and the food presence is detected using the lightweight YOLOv5n optimization model, and the food delivery trajectory is dynamically planned to realize the standardized food collection and food delivery operations of the robotic arm.
It improves the accuracy and robustness of mouth positioning, adapts to different foods and containers, avoids feeding with empty spoons, improves the accuracy and safety of food collection, realizes contactless natural interaction, and simplifies the operation process.
Smart Images

Figure CN120326641A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of human-computer interaction of feeding assistance robotic arms, and particularly relates to a fully automatic mouth-following meal assistance method and system based on machine vision. Background Art
[0002] With the aggravation of population aging and the continuous increase in the number of people with upper limb dysfunction caused by reasons such as stroke, spinal cord injury, and neurodegenerative diseases, how to help such people achieve independent daily eating has become a key issue of social concern. To reduce the burden on caregivers and improve user autonomy, meal assistance robots, as a new type of auxiliary device, have gradually entered the research and application stage.
[0003] Researchers have tried to introduce machine vision technology into meal assistance robots to obtain the user's facial information through a camera and locate the position of the mouth. Currently, some research uses 2D images to detect the user's mouth, but it is impossible to accurately estimate the spatial position due to the lack of depth information; some systems obtain 3D poses through a depth camera combined with a facial key point detection algorithm, but the processing speed is slow and the accuracy is greatly affected by light and occlusion. In addition, there are already some commercial meal assistance robot products, such as Handy 1 (UK), MYSPOON (Japan), Bestic (Sweden), etc. These systems generally adopt a preset fixed feeding trajectory, but the fixed feeding trajectory has low accuracy and poor flexibility. Summary of the Invention
[0004] The purpose of the embodiments of the present invention is to provide a fully automatic mouth-following meal assistance method and system based on machine vision, which solves the problem of being unable to accurately estimate the mouth pose and dynamically generate a meal delivery trajectory according to the real-time mouth pose.
[0005] To achieve the above purpose, the embodiments of the present invention provide a fully automatic mouth-following meal assistance method based on machine vision, and the meal assistance method includes: Establish a robotic arm base coordinate system according to the positional relationship between the robotic arm and the depth camera; Obtain a two-dimensional image of a human face; Obtain the mouth pose of the key points of the two-dimensional human face image based on an improved direct linear transformation algorithm; Judge whether the feeding condition is satisfied according to the mouth pose; In the case where it is judged that the mouth pose satisfies the feeding condition, the robotic arm obtains a feeding trajectory according to teach learning to perform a feeding operation; Judge whether the robotic arm has completed the feeding operation; In the case where it is judged that the robotic arm has completed the feeding operation, update the mouth pose and obtain a meal delivery trajectory according to the mouth pose; The robotic arm performs a meal delivery operation according to the meal delivery trajectory.
[0006] Optionally, establish the base coordinate system of the robotic arm according to the positional relationship between the robotic arm and the depth camera, including: Install an ArUco marker at the end of the robotic arm; Obtain the first spatial position of the ArUco marker through the depth camera; Obtain the second spatial position of the end of the robotic arm in its own coordinate system through the robotic arm; Calculate the rigid body transformation matrix according to the first spatial position, the second spatial position and the ICP algorithm to transform the camera coordinate system into the base coordinate system of the robotic arm.
[0007] Optionally, obtain the mouth pose of the key points of the two-dimensional face image based on the improved direct linear transformation algorithm, including: Obtain the key points of the two-dimensional face image and the three-dimensional key points of the corresponding calibrated three-dimensional face model according to formulas (1) to (2), , (1) , (2) wherein, is the key point of the two-dimensional face image, is the three-dimensional key point, is the abscissa of the key point of the two-dimensional image, is the ordinate of the key point of the two-dimensional image, is the coordinate of the three-dimensional key point in the X-axis direction, is the coordinate of the three-dimensional key point in the Y-axis direction, is the coordinate of the three-dimensional key point in the Z-axis direction, is the th key point, is the integer number; Normalize the key points of the two-dimensional face image; Construct a cross product constraint algebraic model according to formula (3), , (3) wherein, is the anti-symmetric cross product matrix and , is the rotation matrix, is the translation vector; Obtain the direction projection residual according to formula (4), , (4) wherein, is the direction projection residual, is the abscissa of the normalized key point of the two-dimensional image, is the ordinate of the normalized key point of the two-dimensional image, is the first component of the translation vector, is the second component of the translation vector, is the x - direction of the translation vector, is the y - direction of the translation vector, is the first component of the rotation matrix, is the second component of the rotation matrix; Construct the objective function according to formulas (5) to (6), , (5) , (6) wherein, is the residual quadratic form matrix, is the first residual coefficient matrix, and , is the second residual coefficient matrix, and , is the objective function, is the rotation matrix and , is the number of key points; Minimize the said objective function to obtain the said rotation matrix and translation vector.
[0008] Optionally, minimizing the said objective function to obtain the said rotation matrix and translation vector includes: Minimize the said objective function to obtain the first component, second component of the said rotation matrix and the first component and second component of the translation vector; Obtain the third component of the rotation matrix according to geometric relations; Perform singular value decomposition on the said rotation matrix to obtain the optimal rotation matrix; Obtain the third component of the translation vector according to the depth coordinates of the said 3D key points; Obtain the pitch angle, yaw angle and 3D coordinates of the mouth pose according to the said optimal rotation matrix and translation vector.
[0009] Optionally, judging whether the feeding condition is satisfied according to the said mouth pose includes: Obtain the first feeding time; Judge whether the first feeding time is greater than or equal to the first feeding period; In the case of judging that the first feeding time is greater than or equal to the first feeding period, obtain the pitch angle of the mouth pose; Judge whether the pitch angle is greater than the first set threshold; In the case of judging that the pitch angle is greater than the first set threshold, determine that the feeding condition is satisfied; When it is determined that the pitch angle is not greater than the first set threshold, it is determined that the feeding condition is not satisfied; Obtain the second feeding time; Judge whether the second feeding time is greater than or equal to the second feeding cycle; When it is determined that the second feeding time is greater than or equal to the second feeding cycle, obtain the swing angle of the mouth position; Judge whether the swing angle is greater than the second set threshold; When it is determined that the swing angle is greater than the second set threshold, it is determined that the feeding condition is not satisfied; When it is determined that the swing angle is not greater than the second set threshold, it is determined that the feeding condition is satisfied.
[0010] Optionally, judging whether the robotic arm has completed the food-taking operation includes: Obtain the image information of the spoon at the end of the robotic arm according to the RGB camera; Preprocess the image information; Construct a lightweight YOLOv5n optimized model; Input the image information into the lightweight YOLOv5n optimized model to obtain the binary classification result of whether there is food on the spoon; When the binary classification result that there is food on the spoon is obtained, it is determined that the robotic arm has completed the food-taking operation; When the binary classification result that there is no food on the spoon is obtained, it is determined that the robotic arm has not completed the food-taking operation, and the robotic arm executes the food-taking trajectory obtained according to the teaching learning to perform the food-taking operation.
[0011] Optionally, constructing a lightweight YOLOv5n optimized model includes: Obtain the YOLOv5n backbone network as the backbone network of the lightweight YOLOv5n optimized model; Delete the Conv module, C3 module and SPPF module of the YOLOv5n backbone network; Remove the P5 detection head and the Anchor configuration corresponding to the P5 detection head; Configure lightweight parameters.
[0012] Optionally, updating the mouth position and obtaining the food delivery trajectory according to the mouth position includes: Obtain the three-dimensional coordinates of the mouth position based on the improved direct linear transformation algorithm; Determine the food delivery target point according to the three-dimensional coordinates; Plan the food delivery trajectory according to the food delivery target point and the position of the spoon at the end of the robotic arm.
[0013] Optionally, updating the mouth position and posture and obtaining a food delivery trajectory according to the mouth position and posture includes: Obtaining an end point of the food delivery trajectory according to the food delivery trajectory; Judging whether the distance between the food delivery target point and the end point of the food delivery trajectory is greater than a third set threshold; In the case where it is judged that the short distance between the food delivery target point and the end point of the food delivery trajectory is greater than the third set threshold, re-planning the food delivery trajectory; In the case where it is judged that the food delivery target point and the end point of the food delivery trajectory are not greater than the third set threshold, performing a food delivery operation; Obtaining a third feeding time; Judging whether the third feeding time is greater than a third feeding period; In the case where it is judged that the third feeding time is greater than the third feeding period, the robotic arm performs a food taking operation according to the food taking trajectory obtained by teaching learning.
[0014] On the other hand, the present invention provides a full-automatic mouth following assisted meal system based on machine vision, and the system includes: A robotic arm, arranged on a tabletop, and a spoon is arranged at the end of the robotic arm; A vision module, the vision module includes a depth camera and an RGB camera, the depth camera is used to obtain a two-dimensional face image, and the RGB camera is used to obtain spoon image information; A control module, connected to the vision module and the robotic arm, for executing any one of the above-mentioned assisted meal methods.
[0015] Through the above technical solutions, the present invention provides an assisted meal method and system with full-automatic mouth following based on machine vision. By improving the direct linear transformation algorithm, the mouth position and posture of the key points of the two-dimensional face image are obtained. On the basis of the classical DLT structure, combined with the orthogonal constraint between rotation vectors and the singular value decomposition strategy, the PnP problem is transformed into a linear algebra structure for solution. On the premise of ensuring accuracy, the calculation efficiency is greatly improved, and the error amplification and degradation problems caused by the solution of mixed variables in the traditional method are avoided. It is suitable for real-time pose estimation using only sparse structure points, improving the detection accuracy and robustness. The food taking trajectory obtained by teaching learning enables the robotic arm to perform standardized food taking, adapting to different food types and container shapes. By constructing a lightweight YOLOv5n optimization model, the success rate of scooping food is improved, empty spoon feeding is avoided, and the user experience is enhanced. And it can monitor and update the mouth position and posture in real time, adapt to the real-time action changes of the user, and generate a smooth and collision-free trajectory according to the mouth position and posture, improving the accuracy and safety of food taking.
[0016] Other features and advantages of the embodiments of the present invention will be described in detail in the subsequent specific implementation section. Description of the Drawings
[0017] The accompanying drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the following specific implementation manners, they are used to explain the embodiments of the present invention, but do not constitute a limitation to the embodiments of the present invention. In the accompanying drawings: Figure 1 is a flowchart of a meal assistance method according to an embodiment of the present invention; Figure 2 is a flowchart of space calibration according to an embodiment of the present invention; Figure 3 is a flowchart of obtaining the mouth pose according to an embodiment of the present invention; Figure 4 is a flowchart of minimizing the objective function according to an embodiment of the present invention; Figure 5 is a flowchart of determining whether the mouth pose meets the food-taking condition according to an embodiment of the present invention; Figure 6 is a flowchart of determining whether the mouth pose meets the food-taking condition according to an embodiment of the present invention; Figure 7 is a flowchart of determining whether the robotic arm has completed the food-taking operation according to an embodiment of the present invention; Figure 8 is a flowchart of constructing a lightweight YOLOv5n optimization model according to an embodiment of the present invention; Figure 9 is a flowchart of updating the mouth pose and obtaining the food delivery trajectory according to the mouth pose according to an embodiment of the present invention; Figure 10 is a flowchart of updating the mouth pose and obtaining the food delivery trajectory according to the mouth pose according to an embodiment of the present invention; Figure 11 is a physical diagram of a meal assistance system according to an embodiment of the present invention; Figure 12 is a detection result diagram of space calibration according to an embodiment of the present invention; Figure 13 is a detection result diagram of obtaining the mouth pose according to an embodiment of the present invention; Figure 14 is a detection result diagram of the binary classification result according to an embodiment of the present invention. Specific implementation manners
[0018] The following details the specific implementation manners of the embodiments of the present invention with reference to the accompanying drawings. It should be understood that the specific implementation manners described herein are only used to illustrate and explain the embodiments of the present invention, and do not limit the embodiments of the present invention.
[0019] It should be noted that in the technical solution of this application, the acquisition, transmission, storage, use, processing, etc. of data all comply with the relevant regulations of laws and regulations. In the embodiments of this application, some existing solutions in the industry such as certain software, components, models, etc. may be mentioned. They should be regarded as exemplary. The purpose is only to illustrate the feasibility in the implementation of the technical solution of this application, but it does not mean that the applicant has already or necessarily used this solution.
[0020] Figure 1 is a flowchart of a meal assistance method according to an embodiment of the present invention. In this figure, the meal assistance method includes: In step S1, a robotic arm base coordinate system is established according to the positional relationship between the robotic arm and the depth camera. In the present invention, before the robotic arm performs the food-taking and food-delivering actions, coordinate system one is a prerequisite for ensuring the accuracy of the task. Since the depth camera and the robotic arm are in different coordinate systems, in order to accurately transform the pose of the mouth detected by the camera into the working coordinate system of the robotic arm, it is necessary to complete the spatial calibration between the coordinate systems before the system runs. As Figure 12 shown, since the present invention adopts a "vision configuration with eyes outside the hand", that is, a fixed depth camera is used to observe the working area and the user's face, and the depth camera and the robotic arm are in different coordinate systems, so it is necessary to complete the spatial calibration between the coordinate systems before the system runs.
[0021] In step S2, a two-dimensional face image is acquired.
[0022] In step S3, based on the improved direct linear transformation algorithm, the mouth pose of the key points of the two-dimensional face image is obtained.
[0023] In step S4, it is judged whether the food-taking condition is satisfied according to the mouth pose. When it is judged that the mouth pose satisfies the food-taking condition, step S5 is executed; otherwise, step S2 is executed. The start and stop control of the robotic arm is realized through the mouth pose, and the user does not need to operate with hands, and the interaction is intuitive and contactless.
[0024] In step S5, the robotic arm obtains a food-taking trajectory according to the teaching learning to perform the food-taking operation. The robotic arm obtains a food-taking trajectory according to the teaching learning, and automatically executes the actions of "scoop - scrape - lift" according to the food-taking trajectory, specifically including: the spoon goes deep into the dinner plate to scoop food; the spoon scrapes the edge structure of the dinner plate to scrape off the excess food; the spoon is lifted to the detection position. Since the food-taking trajectory is obtained through the teaching learning, it can adapt to different food types and container shapes.
[0025] In step S6, it is judged whether the robotic arm has completed the food-taking operation. When it is judged that the robotic arm has completed the food-taking operation, step S7 is executed; otherwise, step S5 is executed.
[0026] In step S7, update the mouth pose and obtain the food delivery trajectory according to the mouth pose.
[0027] In step S8, the robotic arm performs the food delivery operation according to the food delivery trajectory. Among them, steps S5 to S8 are repeatedly executed to ensure the completion of feeding. When executing steps S5 to S8, if it is determined that the mouth pose does not meet the food-taking condition, return to execute step S2. That is, when repeatedly executing steps S5 to S8, if it is detected that the food-taking condition is not met, the robotic arm immediately stops working.
[0028] In steps S1 to S8, the present invention provides a full-automatic mouth-following meal assistance method based on machine vision, which obtains the mouth pose of the key points of the human face two-dimensional image based on the improved direct linear transformation algorithm, avoiding the error amplification and degradation problems caused by the solution of mixed variables in the traditional method, being applicable to real-time pose estimation using only sparse structure points, and improving the accuracy of mouth positioning. The food-taking trajectory obtained through teach learning enables the robotic arm to perform standardized food-taking, adapting to different food types and container shapes. And it can monitor and update the mouth pose in real time, adapt to the real-time movement changes of the user, generate a smooth and collision-free trajectory according to the mouth pose, and improve the accuracy and safety of food-taking.
[0029] In step S1, in order to ensure the accuracy of the coordinate mapping from the camera coordinate system to the robotic arm working space, it is necessary to establish the robotic arm base coordinate system according to the positional relationship between the robotic arm and the depth camera. There are various ways to perform the spatial calibration between coordinate systems known to those skilled in the art. In one example of the present invention, specifically, the method for performing the spatial calibration between coordinate systems may include, for example, Figure 2 the method shown in Figure 2 In this In step S11, install an ArUco marker at the end of the robotic arm.
[0030] In step S12, obtain the first spatial position of the ArUco marker through the depth camera.
[0031] In step S13, obtain the second spatial position of the end in its own coordinate system through the robotic arm.
[0032] In step S14, calculate the rigid body transformation matrix according to the first spatial position, the second spatial position and the ICP algorithm to transform the camera coordinate system into the robotic arm base coordinate system. Among them, the rigid body transformation matrix is composed of a rotation matrix and a translation vector, representing the fixed spatial relationship between the camera of the depth camera and the robotic arm.
[0033] In steps S11 to S14, the rigid body transformation matrix of the robotic arm base coordinate system is used to convert the points in the camera coordinate system into the robotic arm base coordinate system, store the calibration results, and all subsequent visual detection results are converted into the robotic arm base coordinate system through this matrix. Calibration ensures the accuracy of the coordinate mapping from the camera coordinate system to the robotic arm working space and is the basis for subsequent mouth positioning and meal delivery trajectory planning.
[0034] In step S2, a two-dimensional image of the user's face is obtained through a fixedly installed depth camera. Based on the key points detected in the two-dimensional face image and the corresponding three-dimensional key points of the calibrated three-dimensional face model, the mouth pose is further obtained. The depth camera detects the position and pose of the mouth, representing the real-time pose of the mouth relative to the camera of the depth camera. There are various ways to obtain the mouth pose known to those skilled in the art. In a preferred example of the present invention, specifically, the way to obtain the mouth pose may include, for example Figure 3 the method shown in Figure 3 . In this In step S31, according to formulas (1) to (2), the key points of the two-dimensional face image and the three-dimensional key points of the corresponding calibrated three-dimensional face model are obtained. , (1) , (2) where is the key point of the two-dimensional face image, is the three-dimensional key point, is the abscissa of the key point of the two-dimensional image, is the ordinate of the key point of the two-dimensional image, is the coordinate of the three-dimensional key point in the X-axis direction, is the coordinate of the three-dimensional key point in the Y-axis direction, is the coordinate of the three-dimensional key point in the Z-axis direction, is the th key point, is an integer number used to index a set of key points.
[0035] In step S32, the key points of the two-dimensional face image are normalized. Among them, in order to unify the coordinate scale and eliminate the influence of the camera internal parameters, the key points of the two-dimensional image can be projected into the normalized line-of-sight direction according to formula (7). , (7) where is the normalized key point of the two-dimensional image, in is the th key point, is the camera intrinsic matrix and , is the horizontal focal length of the camera, is the vertical focal length of the camera, is the abscissa of the optical center in the image, is the ordinate of the optical center in the image, is the horizontal direction (x - direction) of the image, is the vertical direction (y - direction) of the image, is the abscissa of the key point of the normalized two - dimensional image, is the ordinate of the key point of the normalized two - dimensional image, and , .
[0036] Theoretically, the normalized light direction is consistent with the direction of the world point after pose transformation, , (8) According to the collinearity of vectors, there is , (9) where, is the coordinate of the th three - dimensional key point in the camera coordinate system, is the rotation matrix, is the translation vector.
[0037] Furthermore, define as the skew - symmetric cross - product matrix of , and substitute into formula (9). In step S33, construct the cross - product constraint algebraic model according to formula (3), , (3) where, is the skew - symmetric cross - product matrix and , is the rotation matrix, is the translation vector.
[0038] Take the third row of the skew - symmetric cross - product matrix to construct the error term. In step S34, obtain the direction projection residual according to formula (4), , (4) where, is the direction projection residual, is the abscissa of the key point of the normalized two - dimensional image, is the ordinate of the key point of the normalized two - dimensional image, is the first component of the translation vector, is the second component of the translation vector, is the x - direction of the translation vector, is the y - direction of the translation vector, is the first component of the rotation matrix, is the second component of the rotation matrix.
[0039] Set the rotation matrix according to formula (10), , (10) where, is the rotation matrix, is the first component of the rotation matrix, is the second component of the rotation matrix. Due to the inherent structure of the rotation matrix: any two rows are unit vectors, different two rows are orthogonal, and the third row can be obtained by the cross - product of the first two rows. Therefore, we only need to find out , to form the base plane, and then directly construct using orthogonality, which can ensure that there is no deviation in the rotation structure. While the traditional DLT does not consider this structure, the three " " rows obtained may not form an orthogonal system and must be corrected by post - processing, which often leads to error accumulation.
[0040] Determine the residual coefficient row vector of each key point according to formulas (11) to (12), , (11) , (12) where, is the linear coefficient of the direction projection residual with respect to the rotation variable , that is, the linear coefficient of the direction projection residual of the th key point with respect to the rotation variable , is the linear coefficient of the direction projection residual with respect to the translation variable , that is, the linear coefficient of the direction projection residual of the th key point with respect to the translation variable . Stacking all the is . Stacking all the is .
[0041] After eliminating the translation terms , in formula (4), construct the residual quadratic form matrix. In step S35, construct the objective function according to formulas (5) to (6), , (5) , (6) where, is the residual quadratic form matrix, is the first residual coefficient matrix, and , is the abscissa of the th key point of the normalized two-dimensional image, is the ordinate of the th key point of the normalized two-dimensional image, is the th three-dimensional key point, is the second residual coefficient matrix, and , is the objective function, is the rotation matrix and , is the first component of the rotation matrix, is the second component of the rotation matrix, is the number of key points.
[0042] In step S36, the objective function is minimized to obtain the rotation matrix and the translation vector.
[0043] In steps S31 to S36, in order to achieve real-time and accurate perception of the spatial position and orientation (6-DoF) of the user's mouth by the robot, i.e., the robotic arm, the present invention proposes a fast PnP solution method based on orthogonal constraint minimization, that is, by improving the direct linear transformation to solve the PnP problem to estimate the pose of the mouth from the face key points. Based on the classical DLT structure, this method combines the orthogonal constraint between rotation vectors to transform the PnP problem into a linear algebra structure for solution, which greatly improves the calculation efficiency on the premise of ensuring accuracy. It is applicable to the low-power and real-time embedded assisted dining robot platform, and the inference speed is about 33FPS on a laptop cpu.
[0044] In this embodiment, there are various ways to minimize the objective function known to those skilled in the art. In a preferred example of the present invention, specifically, the way to minimize the objective function may include, for example, the method shown in Figure 4 . In this Figure 4 , the method to minimize the objective function may include: In step S361, the objective function is minimized to obtain the first component and the second component of the rotation matrix and the first component and the second component of the translation vector. The first two rows of the rotation matrix and the first two dimensions of the translation vector are solved by the algebraic residual quadratic form matrix. Since has a rank of 4 and the dimension of its homogeneous solution space is 2, therefore, the general solution of the rotation matrix can be set according to formula (12), , (12) where, is the general solution, is the first parameter, is the first basis vector, is the second basis vector. Extract , representing the rotation matrix take the first three elements, corresponding to the first row of the rotation matrix, , representing the rotation matrix the 4th to 6th elements, corresponding to the second row of the rotation matrix, and apply the orthogonal constraint matrix , to obtain , (13) wherein, is the first coefficient, is the second coefficient, is the third coefficient, and , , , is corresponding to the part is the first row subspace component of the rotation matrix , is corresponding to the part is the second row subspace component of the rotation matrix , is corresponding to the part is the first row's another orthogonal basis direction component of the rotation matrix , is corresponding to the part is the second row's another orthogonal basis direction component of the rotation matrix . Solve and substitute it into formula (12) to obtain the unique legal solution rotation matrix .
[0045] Furthermore, according to formula (14), restore the first component and the second component of the translation vector, , (14) In step S362, obtain the third component of the rotation matrix according to the geometric relationship. Specifically, obtain the rotation matrix according to formula (15), , (15) wherein, is the third component of the rotation matrix. Furthermore, the rotation matrix can be obtained.
[0046] In step S363, the rotation matrix is subjected to singular value decomposition to obtain the optimal rotation matrix. According to formulas (16) to (17), the rotation matrix is subjected to singular value decomposition to obtain the optimal rotation matrix. , (16) , (17) where is the left singular vector matrix, is the right singular vector matrix, is the singular value diagonal matrix, is the rotation matrix, is the optimal rotation matrix.
[0047] In step S364, the third component of the translation vector is obtained according to the depth coordinates of the three-dimensional key points. The third component of the translation vector is obtained according to formula (18). , (18) where is the third component of the translation vector, is the th key point's depth coordinate in the camera coordinate system, that is, the position of the transformed point in the z-axis direction of the camera coordinate system. Finally, the translation vector can be obtained.
[0048] In step S365, the pitch angle, yaw angle, and three-dimensional coordinates of the mouth pose are obtained according to the optimal rotation matrix and the translation vector.
[0049] In steps S361 to S365, the present invention divides the estimation of the mouth pose into two stages. In the first stage, the first two rows of the rotation matrix , and the first two-dimensional components of the translation vector , are solved through the algebraic residual matrix. In the second stage, the third row of the rotation matrix is analytically solved based on the geometric structure of the rotation matrix, and the third dimension component of the depth translation vector is inversely solved in combination with the line-of-sight direction to ensure the orthogonality of the rotation matrix. In the first stage, algebraic least squares is used to construct to control the solution structure, decouple rotation and translation, reduce variables, and improve robustness. In the second stage, based on geometric priors, orthogonality is restored, and then is calculated, By inverse solution based on geometric relationships, the stability is stronger, the physical constraints are clarified, and deviations caused by post-correction are avoided. Compared with traditional DLT that requires post-processing to correct the orthogonality of the rotation matrix, this method ensures that the solution process itself meets the geometric consistency conditions through homogeneous solution space analysis and analytical orthogonality constraints, improving numerical stability and physical consistency. Due to the characteristics of the mouth key points such as a small number of points (n = 6 key points), partial approximate coplanarity, and large confidence changes, this method can still stably estimate the three-dimensional pose of the mouth under the condition of only 6 key points, solving the problems of easy degradation and instability of traditional DLT when the number of points is small and the distribution is dense. By real-time estimating the mouth pose, deviations caused by fixed trajectories are avoided, realizing dynamic positioning and following, which is applicable to scenarios such as intelligent feeding and visual interaction. Further, the key point confidence weighted residual row can be extended and introduced to increase the anti-interference ability. Among them, the obtained mouth pose can be as Figure 13 shown.
[0050] In step S4, it is judged whether the feeding condition is met according to the mouth pose. For the method of judging whether the feeding condition is met, there are various methods known to those skilled in the art. In an example of the present invention, specifically, the method of judging whether the feeding condition is met may include the methods shown in Figure 5 and Figure 6 . In this Figure 5 and Figure 6 , the feeding assistance method may include the following steps: In step S41, the first feeding time is obtained.
[0051] In step S42, it is judged whether the first feeding time is greater than or equal to the first feeding cycle. If it is judged that the first feeding time is greater than or equal to the first feeding cycle, step S43 is executed; otherwise, step S41 is executed.
[0052] In step S43, the pitch angle of the mouth pose is obtained.
[0053] In step S44, it is judged whether the pitch angle is greater than the first set threshold. If it is judged that the pitch angle is greater than the first set threshold, step S45 is executed; otherwise, step S46 is executed. Among them, the first set threshold can be 10°.
[0054] In step S45, it is determined that the feeding condition is met.
[0055] In step S46, it is determined that the feeding condition is not met.
[0056] In step S47, the second feeding time is obtained.
[0057] In step S48, it is judged whether the second feeding time is greater than or equal to the second feeding cycle. In the case where it is judged that the second feeding time is greater than or equal to the second feeding cycle, step S49 is executed; otherwise, step S47 is executed.
[0058] In step S49, the swing angle of the mouth pose is obtained.
[0059] In step S410, it is judged whether the swing angle is greater than the second set threshold. In the case where it is judged that the swing angle is greater than the second set threshold, step S46 is executed; otherwise, step S45 is executed. Among them, the second set threshold can be 15°. Further, during the process of looping through steps S5 to S9, the swing angle of the mouth pose is obtained in real time and it is judged whether the swing angle is greater than the second set threshold. In the case where it is judged that the swing angle is greater than the second set threshold, it is determined that the feeding condition is not met, and the robotic arm immediately stops working.
[0060] In steps S41 to S410, by monitoring the pitch angle and the swing angle, the feeding requirements of the user can be obtained. When the pitch angle is greater than the first set threshold, it indicates that the user nods, and when the feeding condition is met, the robotic arm starts to pick up food; when the swing angle is greater than the second set threshold, it indicates that the user shakes their head, and when the feeding condition is not met, the robotic arm does not pick up food or stops picking up food. Existing meal assistance robots mostly use buttons or touch screens for control, which is complex and not intuitive, and some users are difficult to operate due to mobility problems. In the present invention, the user does not need to operate with their hands, realizing contactless natural interaction and significantly improving the operation convenience.
[0061] In step S6, it is judged whether the robotic arm has completed the food picking operation. For the method of judging whether the robotic arm has completed the food picking operation, there are various methods known to those skilled in the art. In one example of the present invention. Specifically, judging whether the robotic arm has completed the food picking operation may include as Figure 7 shown in. In this Figure 7 the meal assistance method may include the following steps: In step S61, image information of the spoon at the end of the robotic arm is obtained according to the RGB camera.
[0062] In step S62, the image information is preprocessed. The image is scaled to 640x640 pixels and normalized.
[0063] In step S63, a lightweight YOLOv5n optimization model is constructed.
[0064] In step S64, the image information is input into the lightweight YOLOv5n optimization model to obtain a binary classification result of whether there is food on the spoon. Among them, the detection result can be as Figure 14 shown.
[0065] In step S65, when the binary classification result that there is food on the spoon is obtained, it is determined that the robotic arm has completed the food-taking operation, and step S7 is executed.
[0066] In step S66, when the binary classification result that there is no food on the spoon is obtained, it is determined that the robotic arm has not completed the food-taking operation, and step S5 is executed.
[0067] In steps S61 to S66, if it is determined that the robotic arm has not completed the food-taking operation, then return to execute step S5. If it is continuously determined three times that the robotic arm has not completed the food-taking operation, the buzzer on the robotic arm control board emits an alarm signal, and at the same time the robotic arm stops working. By cooperating the RGB camera with the robotic arm to form a food-taking closed loop, the success rate of scooping food is improved, empty-spoon feeding is avoided, and the user experience is enhanced.
[0068] In this embodiment, there are various ways to construct the lightweight YOLOv5n optimization model, which are known to those skilled in the art. In a preferred example of the present invention, specifically, the way to construct the lightweight YOLOv5n optimization model may include as Figure 8 shown in the method. In this Figure 8 method, the method to construct the lightweight YOLOv5n optimization model may include the following steps: In step S631, obtain the YOLOv5n backbone network as the backbone network of the lightweight YOLOv5n optimization model.
[0069] In step S632, delete the Conv module, C3 module and SPPF module of the YOLOv5n backbone network. Among them, the number of channels of the feature maps output by the Conv module and the C3 module is 1024, which can also be expressed as Conv(1024), C3(1024). The deep P5 structure in the original YOLOv5n backbone network is mainly used to detect large targets, which does not match the "food detection in the spoon" scenario in the present invention, and its computational complexity is relatively high. Therefore, delete the Conv module, C3 module and SPPF module included in the 7th to 9th layers, and only retain up to P4 (P4 / 16) as the deepest feature extraction, which significantly reduces the model volume and improves the response speed of small targets.
[0070] In step S633, the P5 detection head and the corresponding Anchor configuration are removed. In the detection head part, the original YOLOv5n model performs multi-scale detection on the feature maps of the P3, P4, and P5 layers. In the present invention, the P5 detection head is removed, and only two detection heads of P3 (suitable for extremely small targets) and P4 (suitable for medium and small targets) are retained. The expression ability of the shallow and middle feature maps for tiny food blocks is fully utilized to improve the detection sensitivity and reduce redundant inference calculations. Along with the pruning of the P5 layer, its corresponding Anchor configuration is also deleted synchronously, and only two sets of Anchor sizes corresponding to P3 / 8 and P4 / 16 are retained. This synchronous modification ensures the rationality of the detection box regression mechanism and avoids the problem of mismatch between the anchor and the feature layer.
[0071] In step S634, lightweight parameters are configured. The original channel width scaling and depth scaling strategies of YOLOv5n are retained in the network structure, and width_multiple = 0.25 and depth_multiple = 0.33 are set to maintain extremely low model parameter quantities and computational amounts, making it convenient to be deployed on resource-constrained platforms such as Jetson Nano, Raspberry Pi, and edge boxes.
[0072] In steps S631 to S634, improvements are made on the basis of the original YOLOv5n v6.0 framework. While retaining its structural advantages, the model scale and running speed are optimized. The final model size is 1.7MB, and the inference speed on a notebook mx250 graphics card is 23FPS, making it more suitable for deployment on edge computing devices.
[0073] In order to dynamically generate the food delivery target point to improve the accuracy of food delivery, in this embodiment, the mouth pose needs to be updated. There are various ways to update the mouth pose, which are known to those skilled in the art. In an example of the present invention, specifically, the ways to update the mouth pose may include, for example Figure 9 shown in the method. In this Figure 9 , the food assistance method may include the following steps: In step S71, the three-dimensional coordinates of the mouth pose are obtained based on the improved direct linear transformation algorithm.
[0074] In step S72, the food delivery target point is determined according to the three-dimensional coordinates.
[0075] In step S73, the food delivery trajectory is planned according to the food delivery target point and the position of the spoon at the end of the robotic arm. Specifically, it includes: Generating a food delivery trajectory function according to formula (19),
[0076] where is the food delivery trajectory function, is the starting position of the trajectory, is the initial time, is the first acceleration coefficient, is the second acceleration coefficient, is the third acceleration coefficient, is the fourth acceleration coefficient, is the fifth acceleration coefficient, is the current time, is the initial time. And the robotic arm controls the spoon to stop at a position about 2 cm in front of the mouth, effectively preventing collisions and ensuring the smoothness and safety of meal delivery.
[0077] In steps S71 to S73, the improved direct linear transformation algorithm obtains the three-dimensional coordinates of the mouth to accurately judge the position of the mouth. Secondly, the meal delivery target point is determined according to the three-dimensional coordinates to ensure that the robot can accurately deliver the food to the user's mouth. Finally, by planning the movement trajectory of the spoon at the end of the robotic arm, the meal delivery process is ensured to be smooth and safe, avoiding food splashing or errors. By real-time capturing the position of the user's mouth and dynamically adjusting the meal delivery trajectory according to the current spoon position and the mouth position, the accuracy of meal delivery is ensured and the meal delivery efficiency is improved.
[0078] To prevent the spoon from colliding with the mouth and ensure the smoothness and safety of meal delivery. In this embodiment, specifically, the method of updating the mouth pose and obtaining the meal delivery trajectory according to the mouth pose may include the method as shown in Figure 10 . In this Figure 10 , the meal assistance method may include the following steps: In step S74, obtain the end point of the meal delivery trajectory according to the meal delivery trajectory.
[0079] In step S75, judge whether the distance between the meal delivery target point and the end point of the meal delivery trajectory is greater than the third set threshold. Among them, the third set threshold is 4 cm. Here, the third set threshold may also include a threshold range, and the third set threshold range may include ±4 cm. In the case where the distance between the meal delivery target point and the end point of the meal delivery trajectory is greater than the third set threshold, execute step S76, otherwise, execute step S77.
[0080] In step S76, re-plan the meal delivery trajectory.
[0081] In step S77, perform the meal delivery operation.
[0082] In step S78, obtain the third feeding time.
[0083] In step S79, it is judged whether the third feeding time is greater than the third feeding period. Wherein, the third feeding period can be 15s. In the case of judging that the third feeding time is greater than the third feeding period, step S5 is executed to perform the next food intake. Otherwise, S78 is executed. Wait for 15s to ensure that the user has finished eating, and then continue to perform the next food intake.
[0084] In steps S74 to S79, by real-time monitoring the mouth position, the food delivery trajectory is dynamically adjusted to ensure that when the mouth deviates from the target by more than ±4 cm, the trajectory can be re-planned in time to avoid food delivery deviation.
[0085] On the other hand, the present invention provides a full-automatic mouth-following meal assistance system based on machine vision. The system includes a robotic arm, a vision module, and a control module. Specifically, it can be as Figure 11 shown. Among them, the robotic arm adopts a 6-degree-of-freedom light robotic arm, and each joint movement is controlled by an ESP32 main control board. It is installed on the desktop, and its base is fixedly connected to the desktop, and a spoon is provided at the end of the robotic arm. The vision module includes a depth camera and an RGB camera. The RGB-D depth camera is of the Intel RealSense D415 model, which is fixedly installed in front of the table, with a height of 60 cm and a distance of about 70 cm from the user, to obtain user face and workbench information, and provide a two-dimensional face image and depth data. The RGB camera is installed at the end of the robotic arm, facing the spoon, and is used to take an image inside the spoon after scooping food, and is connected to the control module through a USB interface. The control module is connected to the vision module and the robotic arm, and is used to execute any of the above food delivery methods. For the type of the control module, there can be various types known to those skilled in the art. In an example of the present invention, the control module may include a computer. The computer runs the Ubuntu 16.04 operating system, is equipped with ROS, YOLOv5n detection module and improved direct linear transformation algorithm, connects the depth camera and the end RGB camera through USB, and communicates with the ESP32 through the MQTT protocol at the same time. In addition, the meal assistance system may further include a dinner plate. The dinner plate is placed within the working space of the robotic arm, and a scraping structure is integrated on the edge of the dinner plate to facilitate scraping off excess food after scooping.
[0086] Through the above technical solution, the present invention provides a full-automatic mouth-following meal assistance method and system based on machine vision. By improving the direct linear transformation algorithm, the mouth pose of the key points of the human face two-dimensional image is obtained. On the basis of the classical DLT structure, combined with the orthogonal constraint between rotation vectors and the singular value decomposition strategy, the PnP problem is transformed into a linear algebra structure for solution, which greatly improves the calculation efficiency on the premise of ensuring accuracy, and avoids the error amplification and degradation problems caused by the solution of mixed variables in the traditional method. It is applicable to real-time pose estimation using only sparse structure points, improving the detection accuracy and robustness. The feeding trajectory obtained through teaching learning enables the robotic arm to perform standardized feeding and adapt to different food types and container shapes. By constructing a lightweight YOLOv5n optimization model, the success rate of scooping food is improved, empty spoon feeding is avoided, and the user experience is enhanced. Moreover, it can monitor and update the mouth pose in real time, adapt to the real-time movement changes of the user, generate a smooth and collision-free trajectory according to the mouth pose, and improve the accuracy and safety of feeding.
[0087] It should also be noted that the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, commodity or device including the element.
[0088] The above are only embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A fully automatic mouth-following meal assistance method based on machine vision, characterized in that, The described meal assistance method includes: Establishing a robotic arm base coordinate system according to the positional relationship between the robotic arm and the depth camera; Obtaining a two-dimensional facial image; Obtaining the mouth pose of the key points of the two-dimensional facial image based on an improved direct linear transformation algorithm; Judging whether the feeding condition is satisfied according to the mouth pose; When it is judged that the mouth pose satisfies the feeding condition, the robotic arm obtains a feeding trajectory according to teaching learning to perform a feeding operation; Judging whether the robotic arm has completed the feeding operation; When it is judged that the robotic arm has completed the feeding operation, updating the mouth pose and obtaining a food delivery trajectory according to the mouth pose; The robotic arm performs a food delivery operation according to the food delivery trajectory.
2. The meal assistance method according to claim 1, wherein Establishing a robotic arm base coordinate system according to the positional relationship between the robotic arm and the depth camera includes: Installing an ArUco marker at the end of the robotic arm; Obtaining the first spatial position of the ArUco marker through the depth camera; Obtaining the second spatial position of the end in its own coordinate system through the robotic arm; Calculating a rigid body transformation matrix according to the first spatial position, the second spatial position and the ICP algorithm to transform the camera coordinate system into the robotic arm base coordinate system.
3. The meal assistance method according to claim 1, wherein Obtaining the mouth pose of the key points of the two-dimensional facial image based on an improved direct linear transformation algorithm includes: Obtaining the key points of the two-dimensional facial image and the three-dimensional key points of the corresponding calibrated three-dimensional facial model according to formulas (1) to (2); ,(1) ,(2) Among them, is the key point of the two-dimensional face image, is the three-dimensional key point, is the abscissa of the key point of the two-dimensional image, is the ordinate of the key point of the two-dimensional image, is the coordinate of the three-dimensional key point in the X-axis direction, is the coordinate of the three-dimensional key point in the Y-axis direction, is the coordinate of the three-dimensional key point in the Z-axis direction, is the th key point, is an integer number; Performing normalization processing on the key points of the two-dimensional facial image; Constructing a cross product constraint algebraic model according to formula (3); ,(3) Among them, is an anti-symmetric cross-product matrix and , is a rotation matrix, is a translation vector; Obtaining a direction projection residual according to formula (4); ,(4) wherein, is the directional projection residual, is the abscissa of the key point of the normalized two-dimensional image, is the ordinate of the key point of the normalized two-dimensional image, is the first component of the translation vector, is the second component of the translation vector, is the x direction of the translation vector, is the y direction of the translation vector, is the first component of the rotation matrix, is the second component of the rotation matrix; Constructing an objective function according to formulas (5) to (6); ,(5) ,(6) Among them, is the residual quadratic form matrix, is the first residual coefficient matrix, and , is the second residual coefficient matrix, and , is the objective function, is the rotation matrix and , is the number of key points; Minimizing the objective function to obtain the rotation matrix and the translation vector.
4. The meal assistance method according to claim 3, wherein Minimizing the objective function to obtain the rotation matrix and the translation vector includes: Minimizing the objective function to obtain the first component, the second component of the rotation matrix and the first component and the second component of the translation vector; Obtaining the third component of the rotation matrix according to geometric relationships; Performing singular value decomposition on the rotation matrix to obtain an optimal rotation matrix; Obtaining the third component of the translation vector according to the depth coordinates of the three-dimensional key points; Obtaining the pitch angle, the swing angle and the three-dimensional coordinates of the mouth pose according to the optimal rotation matrix and the translation vector.
5. The meal assistance method according to claim 1, wherein Judging whether the feeding condition is satisfied according to the mouth pose includes: Obtaining a first feeding time; Judging whether the first feeding time is greater than or equal to a first feeding period; When it is judged that the first feeding time is greater than or equal to the first feeding period, obtaining the pitch angle of the mouth pose; Judging whether the pitch angle is greater than a first set threshold; When it is judged that the pitch angle is greater than the first set threshold, determining that the feeding condition is satisfied; When it is judged that the pitch angle is not greater than the first set threshold, determining that the feeding condition is not satisfied; Obtaining a second feeding time; Judging whether the second feeding time is greater than or equal to a second feeding period; When it is judged that the second feeding time is greater than or equal to the second feeding period, obtaining the swing angle of the mouth pose; Judging whether the swing angle is greater than a second set threshold; When it is determined that the swing angle is greater than the second set threshold, it is determined that the feeding condition is not satisfied; When it is determined that the swing angle is not greater than the second set threshold, it is determined that the feeding condition is satisfied.
6. The meal assistance method according to claim 1, characterized in that, Judging whether the robotic arm has completed the food-taking operation includes: Obtaining the image information of the spoon at the end of the robotic arm according to the RGB camera; Preprocessing the image information; Constructing a lightweight YOLOv5n optimized model; Inputting the image information into the lightweight YOLOv5n optimized model to obtain a binary classification result of whether there is food on the spoon; When obtaining the binary classification result that there is food on the spoon, it is determined that the robotic arm has completed the food-taking operation; When obtaining the binary classification result that there is no food on the spoon, it is determined that the robotic arm has not completed the food-taking operation, and the robotic arm executes the food-taking trajectory obtained by teaching learning to perform the food-taking operation.
7. The meal assistance method according to claim 6, wherein Constructing a lightweight YOLOv5n optimized model includes: Obtaining the YOLOv5n backbone network as the backbone network of the lightweight YOLOv5n optimized model; Deleting the Conv module, C3 module and SPPF module of the YOLOv5n backbone network; Removing the P5 detection head and the corresponding Anchor configuration for the P5 detection head; Configuring lightweight parameters.
8. The meal assistance method according to claim 1, characterized in that Updating the pose of the mouth and obtaining the food delivery trajectory according to the pose of the mouth includes: Obtaining the three-dimensional coordinates of the mouth pose based on the improved direct linear transformation algorithm; Determining the food delivery target point according to the three-dimensional coordinates; Planning the food delivery trajectory according to the food delivery target point and the position of the spoon at the end of the robotic arm.
9. The meal assistance method according to claim 8, wherein Updating the pose of the mouth and obtaining the food delivery trajectory according to the pose of the mouth includes: Obtaining the end point of the food delivery trajectory according to the food delivery trajectory; Judging whether the distance between the food delivery target point and the end point of the food delivery trajectory is greater than the third set threshold; When it is judged that the short distance between the food delivery target point and the end point of the food delivery trajectory is greater than the third set threshold, re-planning the food delivery trajectory; When it is judged that the food delivery target point and the end point of the food delivery trajectory are not greater than the third set threshold, execute the food delivery operation; Obtaining the third feeding time; Judging whether the third feeding time is greater than the third feeding period; When it is judged that the third feeding time is greater than the third feeding period, the robotic arm executes the food-taking trajectory obtained by teaching learning to perform the food-taking operation.
10. A fully automatic mouth-following meal assistance system based on machine vision, characterized in that, The system includes: A robotic arm, which is arranged on the table, and a spoon is arranged at the end of the robotic arm; A vision module, the vision module includes a depth camera and an RGB camera, the depth camera is used to obtain a two-dimensional face image, and the RGB camera is used to obtain spoon image information; A control module, which is connected to the vision module and the robotic arm, and is used to execute the meal assistance method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Accurate 6D pose measuring and grabbing method for large sparse feature pallet
CN112109072A
Autonomous meal assisting robot based on active flexibility of joints
CN116494263A
User face pose detection method and device based on improved PnP, and medium
CN117315018A
Operation method and system of humanoid agricultural robot based on path planning
CN119924090A
Method employing reinforcement learning to optimize trajectory of spray painting robot
WO2020134254A1
Cited By
Control method of feeding robot based on multi-modal large model
CN121179437A
A control method of a multi-modal large model feeding robot
CN121179437B