Manual monitoring device and manual monitoring method

The manual monitoring device uses an RGB-D sensor and deep learning to recognize detailed handwashing actions, providing feedback to improve hand hygiene accuracy and compliance, addressing the limitations of existing systems in recognizing finger movements.

JP7854380B2Active Publication Date: 2026-05-01CYBERDYNE INC
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
CYBERDYNE INC
Filing Date
2022-10-25
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing handwashing monitoring systems struggle to accurately recognize detailed finger movements during handwashing, such as 'rubbing fingertips' or 'rubbing between fingers', making it difficult to determine if hands are washed correctly, especially for elderly or care-requiring individuals.

Method used

A manual monitoring device that uses an RGB-D sensor to capture RGB and depth images, estimates skeletal points, and applies deep learning to recognize localized handwashing actions, providing feedback on misrecognized movements without causing mental or physical burden.

Benefits of technology

The device accurately recognizes handwashing behaviors focusing on fingertips, offering real-time feedback to improve hand hygiene without burdening the user, enhancing compliance and reducing infection risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007854380000002
    Figure 0007854380000002
  • Figure 0007854380000003
    Figure 0007854380000003
  • Figure 0007854380000004
    Figure 0007854380000004
Patent Text Reader

Abstract

To provide a manual labor monitoring device and a manual labor monitoring method that can encourage behavioral improvement while monitoring a manual labor behavior without placing a burden on the mind and body of a subject.SOLUTION: A manual labor monitoring device includes a behavior recognition unit that recognizes a manual labor behavior that is a connection of a plurality of local movements from a transition state of a subject's posture on the basis of three-dimensional skeletal information sequentially acquired by a skeletal information acquisition unit, a behavior checking unit that checks the manual labor behavior recognized by the behavior recognition unit while determining whether recognition accuracy is less than a predetermined threshold for each local movement, and a movement improvement suggestion unit that feeds back suggestions for encouraging the subject to improve local movements of manual labor behaviors for which the recognition accuracy is less than the threshold on the basis of results of checking by the behavior checking unit.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a manual handwashing monitoring device and a manual handwashing monitoring method, and aims to propose, for example, a manual handwashing monitoring device and a manual handwashing monitoring method for promoting the improvement of appropriate handwashing behavior in the elderly and those requiring care. [Background technology]

[0002] Elderly people and those requiring long-term care are said to be at high risk of serious injury and illness due to their weakened immune systems, cognitive function, and physical abilities. To ensure the health and safety of those requiring long-term care, long-term care facilities and nursing homes routinely carry out activities to prevent and reduce the risk of injury and illness. These activities include fall prevention, proper posture during meals, oral care, handwashing, disinfection, and ventilation.

[0003] In particular, handwashing and hand sanitization are fundamental to infection control and should be practiced by both caregivers and those receiving care. However, it has been reported that only 50% of nursing care facilities in Japan regularly check the handwashing habits of caregivers, highlighting the need for efforts to make hand hygiene a regular practice.

[0004] Furthermore, caregivers may make mistakes or forget the handwashing procedure due to unexpected events. Additionally, it is conceivable that the person receiving care may become unable to wash their hands properly due to cognitive decline. In such cases, caregivers must stay close to the person receiving care and monitor whether they are washing their hands correctly, but this is difficult in practice considering the burden on caregivers.

[0005] Conventionally, handwashing assistance systems have been proposed to encourage workers to perform exemplary handwashing procedures (see Patent Document 1). This handwashing assistance system displays images of the worker's hands taken during handwashing and exemplary handwashing images representing the exemplary handwashing procedure performed by the worker on the same screen of the display unit.

[0006] This handwashing assistance system determines whether the movements of the worker's fingers, recognized from a hand image, meet the matching conditions for the exemplary handwashing movements shown in the exemplary handwashing image. It displays the result of this determination on the same screen and also displays instruction information to guide the worker's finger movements to meet the matching conditions.

[0007] Furthermore, a handwashing evaluation system has been proposed that allows users to self-recognize their handwashing or disinfecting actions without psychological burden and to improve their motivation (see Patent Document 2). Similar to Patent Document 1 mentioned above, this handwashing evaluation system displays the user's handwashing actions and exemplary actions on the screen, and evaluates the actions by analyzing the user's image information. [Prior art documents] [Patent Documents]

[0008] [Patent Document 1] Japanese Patent Publication No. 2019-219554 [Patent Document 2] Japanese Patent Publication No. 2020-19356 [Overview of the project] [Problems that the invention aims to solve]

[0009] Incidentally, if there were a system that could monitor the handwashing behavior of those requiring care and their caregivers daily, check whether they are washing their hands correctly, and encourage improvements in handwashing behavior as needed, it could be expected to prevent illness and reduce the risk of disease.

[0010] However, the handwashing assistance system described in Patent Document 1 and the handwashing evaluation system described in Patent Document 2 evaluate handwashing behavior based on the degree of agreement between the user's handwashing actions and the exemplary or model actions displayed on the video, which has the problem that it can only recognize rough handwashing actions. For example, it cannot recognize detailed finger movements such as "rubbing fingertips" or "rubbing between fingers," making it difficult to determine whether hands are being washed correctly.

[0011] In recent years, research has progressed on behavioral classification based on human 3D skeletal information as a non-contact method. However, it is difficult to simultaneously measure 3D skeletal information of the whole body and fingers in a non-contact manner, and currently, behavioral classification based only on whole-body skeletal information excluding finger information has been achieved.

[0012] Thus, current non-contact methods for evaluating handwashing behavior have made it difficult to recognize detailed actions, including fingertip movements.

[0013] This invention was made in consideration of the above points, and aims to propose a manual work monitoring device and manual work monitoring method that can monitor manual work behavior and promote behavioral improvement without placing a burden on the subject's mind or body. [Means for solving the problem]

[0014] To solve these problems, the manual monitoring device of the present invention includes an imaging unit that images the upper body of a subject, including both hands and fingers, and sequentially acquires RGB images and depth images; an image coordinate setting unit that estimates the posture of the subject's upper body based on the RGB images sequentially acquired from the imaging unit, and sets multiple representative skeletal points attached to the upper body, centered on the subject's hands and fingers, as image coordinates in a two-dimensional coordinate system; and the image coordinate setting unit sets It is a set of multiple representative points of the skeleton.The system includes: a coordinate group extraction unit that sequentially extracts image coordinate groups related to the subject's manual actions based on the subject's posture transition state from a group of image coordinates; a skeletal information acquisition unit that acquires temporally continuous 3D skeletal information centered on the subject's fingers by temporally synchronizing the image coordinate groups sequentially extracted by the coordinate group extraction unit with depth images sequentially acquired from the imaging unit; an action recognition unit that recognizes manual actions, which are a series of local movements, from the subject's posture transition state based on the 3D skeletal information sequentially acquired by the skeletal information acquisition unit; an action confirmation unit that confirms the manual actions recognized by the action recognition unit by determining whether the recognition accuracy for each local movement is below a predetermined threshold; and an action improvement suggestion unit that, based on the confirmation results from the action confirmation unit, provides feedback to the subject with suggestions to encourage improvement of local movements in the manual actions whose recognition accuracy was below a threshold.

[0015] As a result, the manual work monitoring device can non-contactually recognize the subject's manual work actions, primarily focusing on their fingertips, for each localized movement, and provide feedback with suggestions for improvement as needed. This allows for daily monitoring without placing a burden on the subject's physical or mental state.

[0016] Furthermore, in this invention, the image coordinate setting unit sets the image coordinate group in a two-dimensional coordinate system in which either the left or right shoulder of the subject is the coordinate origin, the line direction connecting the left and right shoulders is the X-axis, and the direction perpendicular to the X-axis is the Z-axis, and the skeletal information acquisition unit uses the depth direction in the depth images sequentially acquired from the imaging unit as the Y-axis, and together with the two-dimensional coordinate system of the image coordinate group sequentially extracted by the coordinate group extraction unit to form a three-dimensional coordinate system and acquire three-dimensional skeletal information.

[0017] As a result, the manual monitoring device clarifies the relationship between the position in 3D space and the position of the captured image for the upper body, focusing on the subject's fingers, making it possible to acquire 3D skeletal information while maintaining high robustness, regardless of the imaging position by the imaging unit.

[0018] Furthermore, in the present invention, the action recognition unit sequentially recognizes the corresponding local actions from the three-dimensional skeleton information sequentially acquired by the skeleton information acquisition unit while referring to a manually created action recognition model constructed by deep learning, using the action recognition pattern set for each local action as teacher data.

[0019] As a result, in the manual monitoring device, the recognition accuracy and recognition speed of local actions based on three-dimensional skeleton information can be significantly improved.

[0020] Furthermore, in the present invention, the action confirmation unit sets a threshold for determining the omission of the action time for each local action, determines whether the action has ended before the threshold corresponding to each local action, and confirms the local action that has ended before the threshold as a misrecognition. The action improvement proposal unit provides feedback with a proposal for promoting the improvement of the local action for the local action confirmed as a misrecognition by the action confirmation unit.

[0021] As a result, in the manual monitoring device, it is possible to determine with relatively high accuracy whether or not there is a misrecognition for a plurality of local actions constituting the manual actions of the subject.

[0022] Furthermore, in the present invention, the action improvement proposal unit provides feedback to the subject that the local action of the manual action that could not be recognized due to the omission of the three-dimensional skeleton information by the action recognition unit cannot be recognized.

[0023] As a result, in the manual monitoring device, even for local actions that could not be recognized due to the omission of three-dimensional skeleton information, such as when the subject's hands are close or the fingertips of both hands overlap, it is possible to provide feedback to the subject about this, and convey in what situations the recognition is impossible.

[0024] Furthermore, the manual monitoring method of the present invention includes a first step of sequentially acquiring RGB images and depth images by imaging the upper body of the subject, including both hands and fingers; a second step of estimating the posture of the subject's upper body based on the RGB images acquired sequentially from the first step, and setting multiple representative skeletal points attached to the upper body, centered on the subject's hands and fingers, as image coordinates in a two-dimensional coordinate system; and setting in the second step It is a set of multiple representative points of the skeleton. The system comprises: a third step of sequentially extracting image coordinate groups related to the subject's manual work actions from a group of image coordinates based on the subject's posture transition state; a fourth step of acquiring temporally continuous three-dimensional skeletal information centered on the subject's fingers by temporally synchronizing the image coordinate groups extracted sequentially in the third step with depth images acquired sequentially from the first step; a fifth step of recognizing manual work actions, which are a series of local movements, from the subject's posture transition state based on the three-dimensional skeletal information acquired sequentially in the fourth step; a sixth step of confirming the manual work actions recognized in the fifth step by determining whether the recognition accuracy for each local movement is below a predetermined threshold; and a seventh step of providing feedback to the subject, based on the confirmation results from the sixth step, with suggestions to encourage improvement of local movements in the manual work actions whose recognition accuracy was below the threshold.

[0025] As a result, the manual monitoring method allows for the non-contact recognition of the subject's manual actions, primarily focusing on their fingertips, for each localized movement, while providing feedback with suggestions for improvement as needed. This enables daily monitoring without placing a burden on the subject's physical or mental state. [Effects of the Invention]

[0026] According to the present invention, it is possible to realize a manual work monitoring device and a manual work monitoring method that can monitor manual work behavior and encourage behavioral improvement without placing a burden on the subject's mind or body. [Brief explanation of the drawing]

[0027] [Figure 1] This is a conceptual diagram illustrating the handwashing monitoring device according to this embodiment. [Figure 2] Figure 1 is a schematic diagram illustrating the procedure for acquiring 3D skeletal information using the handwashing monitoring device. [Figure 3] This is a conceptual diagram illustrating the multiple localized actions that constitute the act of handwashing. [Figure 4] This is a conceptual diagram used to explain the construction of a handwashing behavior recognition model. [Figure 5] This chart shows the recognition rate of each local action in handwashing. [Figure 6] This graph shows an example of the recognition results for a series of local actions that constitute handwashing behavior. [Figure 7] This chart shows the number of frames in which localized movements not performed by the subject were misrecognized. [Figure 8] This graph shows the results of handwashing behavior recognition by a handwashing monitoring device. [Figure 9] This graph shows the results of handwashing behavior recognition by a handwashing monitoring device. [Figure 10] This graph shows the results of handwashing behavior recognition by a handwashing monitoring device. [Modes for carrying out the invention]

[0028] An embodiment of the present invention will be described in detail below with reference to the drawings.

[0029] (1) Configuration of the handwashing monitoring device according to this embodiment Figure 1 shows the overall configuration of the handwashing monitoring device 1 according to this embodiment. This handwashing monitoring device 1 is installed in the handwashing facility in the subject's home and consists of a control unit 2 that controls the entire device, an RGB-D sensor (imaging unit) 3 and a video and audio output unit 4 connected to the control unit 2.

[0030] The RGB-D sensor 3, in addition to its RGB color camera function, has a depth sensor that can measure the distance to an object as seen from the camera, enabling 3D scanning of the object. For example, if the RealSense (a trademark of Microsoft Corporation) LiDAR camera L515 is used as the RGB-D sensor 3, the depth sensor consists of a LiDAR sensor that measures the time it takes for a laser beam to hit an object and bounce back, thereby measuring the distance and direction to the object.

[0031] The RGB-D sensor 3 has a resolution of 640 pixels x 360 pixels and a frame rate of 30 fps, while the depth sensor has a resolution of 1024 pixels x 768 pixels and a frame rate of 30 fps.

[0032] The RGB-D sensor 3 is positioned to capture images of the subject's upper body, including both hands and fingers, based on the subject's standing position in the handwashing facility and the water tap, and sequentially acquires RGB images and depth images as a result of the imaging.

[0033] The control unit 2 includes a control unit 5 consisting of a CPU (Central Processing Unit) that provides overall control of the entire system, and a data storage unit 6 in which various data are stored in a database that can be read and written according to commands from the control unit 5.

[0034] The video and audio output unit 4 has a monitor and a speaker, and displays the necessary video read from the data storage unit 6 on the monitor and outputs the necessary audio from the speaker in accordance with the control unit 5.

[0035] The control unit 5 consists of an image coordinate setting unit 10, a coordinate group extraction unit 11, a skeletal information acquisition unit 12, an action recognition unit 13, an action confirmation unit 14, and an action improvement suggestion unit 15. The image coordinate setting unit 10 estimates the posture of the subject's upper body based on the RGB images sequentially acquired from the RGB-D sensor 3, and sets multiple representative skeletal points attached to the subject's upper body, centered on the subject's fingers, as image coordinates.

[0036] Specifically, the image coordinate setting unit 10 sets the image coordinate set in a two-dimensional coordinate system where either the left or right shoulder of the subject is the coordinate origin, the line connecting the left and right shoulders is the X-axis, and the direction perpendicular to the X-axis is the Z-axis, based on that coordinate origin. In other words, the coordinate origin is moved to the right or left shoulder, and the XZ plane is rotated so that it coincides with the line connecting the left and right shoulders. As a result, it becomes possible to recognize motion with increased robustness regardless of the position of the RGB-D sensor 3.

[0037] The coordinate group extraction unit 11 sequentially extracts image coordinate groups related to the subject's handwashing behavior from the image coordinate group set by the image coordinate setting unit 10, based on the transition state of the subject's posture. In practice, the coordinate group extraction unit 11 needs to extract key areas of the upper body from the handwashing behavior, taking into account changes in the subject's clothing and surrounding environment.

[0038] Therefore, the coordinate group extraction unit 11 employs a lightweight OpenPose architecture (a system that estimates a person's skeleton using deep learning) with a pre-trained model to estimate the posture of both hands and obtain image coordinates of representative skeletal points for those hands. As a result, the coordinate group extraction unit 11 sequentially extracts a total of 60 representative skeletal points as an image coordinate group: 18 points for the upper body torso of the subject and 42 points for both hands (21 points for each hand). Furthermore, by using only stages 1 and 2 of the six stages of the OpenPose architecture, the computational load can be reduced by 32.4%.

[0039] The skeletal information acquisition unit 12 acquires temporally continuous three-dimensional skeletal information centered on the subject's fingers by temporally synchronizing the image coordinate groups sequentially extracted by the coordinate group extraction unit 11 with the depth images sequentially acquired from the RGB-D sensor 3.

[0040] Specifically, the skeletal information acquisition unit 12 uses the depth direction in the depth images sequentially acquired from the RGB-D sensor 3 as the Y-axis, and combines it with the 2D coordinate system of the image coordinate group sequentially extracted by the coordinate group extraction unit 11 to form a 3D coordinate system and acquire 3D skeletal information.

[0041] The conversion formula from the image coordinate system of the two-dimensional coordinate system to the world coordinate system of the three-dimensional coordinate system is expressed as the following formula (1).

Equation

[0042] This K, R, and t are obtained by camera calibration using Zhang's method ("A flexible new technique for camera calibration". IEEE Transactions on Pattern Analysis and Machine Intelligence, 22(11):1330-1334, 2000.). For the calibration to calculate K, a calibration board with circular dots printed in a grid pattern was used. The center coordinates of each dot on the calibration plate were calibrated to 1 / 1000.

[0043] As a result, as shown in FIG. 2, in the handwashing monitoring device 1, regarding the upper body centered on the fingers of the subject, the relationship between the position in the three-dimensional space and the position in the captured image is clarified, and 3D skeleton information can be acquired while maintaining high robustness regardless of the imaging position by the RGB-D sensor 3.

[0044] The action recognition unit 13 recognizes the hand-washing behavior, which is a sequence of multiple local movements, from the transition state of the subject's posture, based on the three-dimensional skeletal information sequentially acquired by the skeletal information acquisition unit 12.

[0045] The handwashing procedures described are based on those recommended by the Japanese Ministry of Health, Labour and Welfare. The appropriate handwashing procedure is as follows: The nine steps recommended are: "Wet your hands with running water," "Apply soap," "Rub your palms well," "Rub the backs of your hands up and down," "Rub your fingertips and nails," "Wash between your fingers," "Wash your thumbs by twisting them in your palms," "Wash your wrists," and "Wet your hands with running water."

[0046] Based on this, we defined eight types of localized actions as handwashing behavior, as shown in Figure 3, from (A) to (H): (A) Wet hands, (B) Apply soap, (C) Rub to palm, (D) Rub back of each hand These are eight types of localized movements: (E) hand, (F) fingertip rub, (G) rub between each finger, (H) thumb rub, and (E) wrist rub.

[0047] The action recognition unit 13 uses the action recognition pattern set for each local action as training data and, while referring to a hand-washing action recognition model constructed by deep learning, sequentially recognizes the corresponding local action from the 3D skeletal information sequentially acquired by the skeletal information acquisition unit 12.

[0048] Specifically, the behavior recognition unit 13 applies a handwashing behavior recognition model consisting of three modules—a convolutional neural network (CNN) layer, a batch normalization layer, and an activation function layer (tanh function)—and a fully connected layer, as shown in Figure 4, in order to recognize the eight types of local actions described above.

[0049] The handwashing behavior recognition model takes the skeletal information of the current frame and the frame two frames prior as input and can calculate the likelihood for each of the nine types of actions, including eight types of local actions related to handwashing and other actions (actions other than handwashing).

[0050] Furthermore, the action recognition unit 13, as a recognition result, post-processes the data by selecting the most frequent action from the actions observed in 15 frames (the current frame and the previous 14 frames) as the current local action. In this way, the action recognition unit 13 learned a handwashing action recognition model using supervised learning with its own dataset. For optimization, a cross-entropy loss function and Adam (adaptive movement estimation) with a learning rate of 0.001 were used.

[0051] As a result, the handwashing monitoring device 1 can significantly improve the recognition accuracy and speed of local movements based on three-dimensional skeletal information.

[0052] The action confirmation unit 14 confirms the handwashing behavior recognized by the action recognition unit 13, determining whether the recognition accuracy for each local action is below a predetermined threshold.

[0053] Experiments using the handwashing monitoring device 1 revealed that the recognition rate (Accuracy) of each local action in handwashing behavior is shown in the table in Figure 5. "Wet hands" was 97.5% [%], "Apply soap" was 95.9% [%], "Rub palm to palm" was 98.7% [%], "Rub back to palm" was 67.0% [%], "Rub fingertips" was 94.0% [%], "Rub between each finger" was 97.2% [%], "Rub thumb" was 91.9% [%], "Rub wrist" was 94.1% [%], and "else" was 90.6% [%]. The average recognition rate for all local actions was 91.1 [%]. It was confirmed that the recognition rate for localized movements other than "rubbing the backs of both hands" was 90.0% or higher.

[0054] Figure 6 shows an example of the recognition results of a series of local actions constituting handwashing behavior using this handwashing monitoring device 1. In Figure 5, the horizontal axis represents time and the vertical axis represents the series of local actions. The dashed line shows the annotation label that serves as training data, and the solid line shows the action recognition result (system recognition) by the handwashing monitoring device 1.

[0055] The experiment showed that the recognition results of a series of localized actions by the handwashing monitoring device 1 gradually became inaccurate. As a result, there were 3 misrecognitions for the action of "rubbing the backs of both hands," 2 misrecognitions for the action of "rubbing between the fingers," 1 misrecognition for the action of "rubbing the wrists," and 1 misrecognition each for the actions of "rubbing the fingertips" and "rubbing the wrists."

[0056] Thus, among the series of localized actions performed by the handwashing monitoring device 1, the recognition rate of the "rubbing the backs of both hands" action was 67.0% [%], which was lower than the recognition rates of the other localized actions. This is thought to be because the "rubbing the backs of both hands" action is similar to the "rubbing between the fingers of both hands" action in handwashing.

[0057] In particular, it is difficult to correctly classify local movements when switching between them or when transitioning between the left and right hands. By annotating them as "other (else)" that do not belong to any specific local movement, it is possible to improve the accuracy of local movement recognition.

[0058] Based on the confirmation results from the behavior confirmation unit 14, the action improvement suggestion unit 15 provides feedback to the subject regarding local actions in handwashing where the recognition accuracy was below a threshold, offering suggestions to encourage improvement of those local actions.

[0059] The video and audio output unit 4, which serves as a feedback mechanism, outputs a video to the monitor's display screen that encourages improvement of the local operation targeted by the operation improvement suggestion unit 15, and also outputs audio from the speaker that encourages improvement of the local operation.

[0060] In the above configuration, the handwashing monitoring device 1 can non-contactually recognize the handwashing behavior of the subject, focusing on their fingertips, for each localized movement, and provide feedback with suggestions to encourage improvement as needed, enabling daily monitoring without placing a burden on the subject's mind or body.

[0061] (2) Other embodiments As described above, in this embodiment, the action confirmation unit 14 confirms the handwashing action recognized by the action recognition unit 13 by determining whether the recognition accuracy for each local action is below a predetermined threshold. However, the present invention is not limited to this, and a threshold for determining the absence of action time may be set for each local action, and it may be determined whether the action has ended below the corresponding threshold for each local action, and local actions that have ended below the threshold may be confirmed as misrecognized.

[0062] The action improvement suggestion unit 15 then provides feedback to the action confirmation unit 14 regarding local actions that have been identified as misrecognized, offering suggestions to encourage improvement of those local actions. As a result, the handwashing monitoring device 1 can determine with relatively high accuracy whether or not multiple local actions constituting the subject's handwashing behavior are misrecognized.

[0063] As a prerequisite, five experiments were conducted to detect errors in handwashing behavior using the handwashing monitoring device 1. As shown in Figure 7, the number of frames in which the device misrecognized local actions not performed by the subject is indicated. According to the results of the five trials, in Trial 1 and Trial 4, the device did not misrecognize local actions not performed by the subject, but in Trial 2, Trial 3, and Trial 5, it was confirmed that misrecognition of local actions not performed by the subject occurred.

[0064] In Trial 2, 92 frames of misrecognition occurred for the action of "rubbing between each finger," which the subject was not performing. In Trial 3, 6 frames of misrecognition occurred for the action of "rubbing the thumb." In Trial 5, 24 frames of misrecognition occurred for the action of "rubbing between each finger." The recognition results of handwashing behavior by the handwashing monitoring device 1 for Trials 2, 3, and 5, in which misrecognition occurred, are shown in Figures 8, 9, and 10, respectively.

[0065] Thus, in trials 2, 3, and 5, local actions that were not performed by the subject were misrecognized. This is thought to be due to limitations in the classification accuracy of the handwashing behavior recognition model of the handwashing monitoring device 1 according to this embodiment. In this error detection experiment, the 92 frames of misrecognition in trial 2, which had the most misrecognitions, amounted to 3.0 seconds, which was shorter than the required time for each local action.

[0066] Therefore, as described above, by setting a threshold for determining the absence of operation time for each local operation, and by determining whether the operation has finished below the corresponding threshold for each local operation, it is possible to identify local operations that have finished below the threshold as misrecognitions (omissions of operation), and to issue appropriate instructions or suggestions.

[0067] In this embodiment, the action improvement suggestion unit 15 provides feedback to the subject regarding local actions in the handwashing behavior whose recognition accuracy was below a threshold, based on the confirmation results by the action confirmation unit 14. However, the present invention is not limited to this, and the action recognition unit 13 provides feedback that local actions could not be recognized due to the lack of 3D skeletal information.

[0068] As a result, the handwashing monitoring device 1 can provide feedback to the subject about local movements that could not be recognized due to the lack of 3D skeletal information, such as when the subject's hands are close together or when the fingertips of both hands are overlapping, thereby communicating the circumstances under which recognition becomes impossible.

[0069] Furthermore, although this embodiment describes a case where a handwashing monitoring device is applied as the manual work monitoring device, it can be broadly applied to various manual work monitoring devices as long as it is possible to monitor manual work behavior and encourage behavioral improvement without placing a burden on the physical and mental state of the subject.

[0070] For example, the interactive information transmission system described in Japanese Patent Registration No. 7157424 by the present inventor can be applied to the manual work monitoring device according to the present invention. This interactive information transmission system allows a skilled worker and a collaborator to exchange information with each other via a network, and the skilled worker can instruct the collaborator on their skills regarding a specific task.

[0071] In this interactive information transmission system, the expert observes the same video as the video centered on the object being handled by the collaborator, while simultaneously transmitting the position of their gaze to the collaborator in real time. At the same time, they instruct the collaborator to transmit the three-dimensional movement of their fingertips as force feedback to each of their fingers in real time, and then feed back the results of this transmission to the expert. As a result, the collaborator can share a sense of presence with the expert, who is located remotely, while performing their work, and indirectly receive real-time instruction on the expert's tacit knowledge and techniques. Furthermore, by sensing the results of the force feedback transmitted to the collaborator, the expert can perceive the gap between their instruction and the collaborator's response in real time.

[0072] The manual work monitoring device uses the fingertip-centered manual work behavior of skilled workers in an interactive information transmission system as a baseline (training data). It then non-contactually recognizes the fingertip-centered manual work behavior of collaborators for each localized movement, and provides feedback with suggestions for improvement as needed. As a result, it is possible to realize a manual work monitoring device that can be used on a daily basis without placing a burden on the physical or mental state of the collaborators.

[0073] Furthermore, the upper limb movement support device described in Japanese Patent Registration No. 6763968 by the present inventor can be applied to the manual work monitoring device according to the present invention. This upper limb movement support device enables a robotic arm mounted on a table to operate in coordination with the operator's hand.

[0074] In this upper limb movement support device, the control unit refers to the recognition content from the upper limb movement recognition unit and appropriately controls the robot arm (articulated arm and end effector) so that the imaging unit alternately captures the operator's face and the area beyond the operator's line of sight at desired switching timings, thereby coordinating the robot arm's movement with the operator's upper limb movements. As a result, the upper limb movement support device can recognize objects beyond the operator's line of sight in real time and operate the articulated arm and end effector in accordance with the operator's intentions and in coordination with the operator's hand.

[0075] The manual work monitoring device uses the operator's manual actions as a baseline (training data) in the upper limb movement support device to non-contact recognize manual actions centered on the fingertips (end effectors) of the robot arm for each localized movement, and provides feedback with suggestions for improvement as needed. As a result, a manual work monitoring device that can routinely monitor the coordinated movements of the robot arm can be realized. [Explanation of Symbols]

[0076] 1...Handwashing monitoring device, 2...Control unit, 3...RGB-D sensor (imaging unit), 4...Video and audio output unit, 5...Control unit, 6...Data storage unit, 10...Image coordinate setting unit, 11...Coordinate group extraction unit, 12...Skeletal information acquisition unit, 13...Action recognition unit, 14...Action confirmation unit, 15...Motion improvement suggestion unit.

Claims

1. An imaging unit that captures images of the subject's upper body, including both hands and fingers, and sequentially acquires RGB images and depth images, An image coordinate setting unit estimates the posture of the subject's upper body based on the RGB images sequentially acquired from the imaging unit, and sets multiple representative skeletal points attached to the subject's upper body, centered on the subject's fingers, as image coordinates in a two-dimensional coordinate system. A coordinate group extraction unit sequentially extracts image coordinate groups related to the subject's manual work actions based on the subject's posture transition state from the image coordinate group which is a set of multiple skeletal representative points set by the image coordinate setting unit, A coordinate group extraction unit sequentially extracts image coordinate groups related to the subject's manual work actions based on the transition state of the subject's posture from the image coordinate group set by the image coordinate setting unit, A skeletal information acquisition unit acquires temporally continuous three-dimensional skeletal information centered on the fingers of the subject by sequentially synchronizing the image coordinate group extracted by the coordinate group extraction unit with the depth image acquired sequentially from the imaging unit. Based on the three-dimensional skeletal information sequentially acquired by the skeletal information acquisition unit, the behavior recognition unit recognizes manual actions, which are a series of local movements, from the transitional states of the subject's posture. The action confirmation unit confirms the manual action recognized by the action recognition unit, while determining whether the recognition accuracy for each local action is below a predetermined threshold, Based on the confirmation results by the aforementioned action confirmation unit, the action improvement suggestion unit provides feedback to the subject regarding the local actions among the manual actions whose recognition accuracy was below the threshold, offering suggestions to encourage improvement of those local actions. A manual monitoring device characterized by comprising the following features.

2. The image coordinate setting unit sets the image coordinate group in a two-dimensional coordinate system in which either the left or right shoulder of the subject is the coordinate origin, the direction of the line connecting the left and right shoulders is the X-axis, and the direction perpendicular to the X-axis is the Z-axis, based on the coordinate origin. The skeletal information acquisition unit uses the depth direction in the depth images sequentially acquired from the imaging unit as the Y-axis, and combines it with the two-dimensional coordinate system of the image coordinate group sequentially extracted by the coordinate group extraction unit to form a three-dimensional coordinate system, thereby acquiring the three-dimensional skeletal information. The manual monitoring device according to feature 1.

3. The action recognition unit uses the action recognition pattern set for each local action as training data, and while referring to a manual action recognition model constructed by deep learning, it sequentially recognizes the corresponding local action from the three-dimensional skeletal information sequentially acquired by the skeletal information acquisition unit. A manual monitoring device according to feature 1 or 2.

4. The aforementioned action confirmation unit sets a threshold for determining the absence of operation time for each local operation, determines whether the operation has ended below the corresponding threshold for each local operation, and confirms that the local operation that has ended below the threshold is a misrecognition. The aforementioned operation improvement suggestion unit provides feedback to the aforementioned local operation that has been identified as a misrecognition by the aforementioned action confirmation unit, offering suggestions to encourage improvement of the said local operation. A manual monitoring device according to feature 1 or 2.

5. The aforementioned motion improvement suggestion unit provides feedback to the action recognition unit regarding the local movements among the manual actions that could not be recognized due to the lack of three-dimensional skeletal information, indicating that the local movements could not be recognized. A manual monitoring device according to feature 1 or 2.

6. The first step involves imaging the upper body of the subject, including both hands and fingers, and sequentially acquiring RGB images and depth images. A second step involves estimating the posture of the subject's upper body based on the RGB images acquired sequentially from the first step, and setting multiple representative skeletal points attached to the subject's upper body, centered on the subject's fingers, as image coordinates in a two-dimensional coordinate system. A third step involves sequentially extracting image coordinate groups related to the subject's manual work actions based on the transition state of the subject's posture from the image coordinate group which is a set of multiple skeletal representative points set in the second step, A fourth step involves synchronizing the image coordinate group extracted sequentially in the third step with the depth images acquired sequentially from the first step, thereby acquiring temporally continuous three-dimensional skeletal information centered on the fingers of the subject. A fifth step involves recognizing a manual action, which is a sequence of multiple local movements, from the transitional state of the subject's posture based on the three-dimensional skeletal information acquired sequentially in the fourth step, The sixth step involves confirming the manual action recognized in the fifth step by determining whether the recognition accuracy for each local action is below a predetermined threshold, Based on the results of the verification in step 6, step 7 involves providing feedback to the subject regarding the local actions among the manual actions whose recognition accuracy was below the threshold, with suggestions to encourage improvement of those local actions. A manual monitoring method characterized by comprising the following features.

7. The second step involves setting the image coordinate group in a two-dimensional coordinate system where either the left or right shoulder of the subject is the coordinate origin, the line direction connecting the left and right shoulders is the X-axis, and the direction perpendicular to the X-axis is the Z-axis, The fourth step involves using the depth direction in the depth images acquired sequentially from the first step as the Y-axis, and combining it with the two-dimensional coordinate system of the image coordinate group extracted sequentially in the third step to form a three-dimensional coordinate system, thereby acquiring the three-dimensional skeletal information. The manual monitoring method according to feature 6.

8. The fifth step sequentially recognizes the corresponding local movements from the three-dimensional skeletal information acquired sequentially in the fifth step, while referring to a manual motion recognition model constructed by deep learning, using the motion recognition pattern set for each local movement as training data. The manual monitoring method according to claim 6 or 7.

9. The sixth step involves setting a threshold for determining the absence of operation time for each local operation, determining whether the operation was completed within the corresponding threshold for each local operation, and confirming that the local operation was completed within the threshold as a misrecognition. The seventh step provides feedback regarding the local operation that was identified as a misrecognition in the sixth step, indicating that the local operation could not be recognized. The manual monitoring method according to claim 6 or 7.

10. The seventh step provides feedback on the local movements that could not be recognized in the fifth step due to the lack of three-dimensional skeletal information, offering suggestions to encourage improvement of those local movements. The manual monitoring method according to claim 6 or 7.

Citation Information

Patent Citations

  • Dexterous hand teleoperation control method based on Kinect human hand motion capturing

    CN104589356A

  • Hand-washing assist system, hand-washing assist method, and hand-washing assist device

    JP2019219554A

  • Control device for vehicle

    JP2020019356A

  • Information processing system, information processing method and learning model generation method

    JP2021152759A

  • Hand-washing evaluation device and hand-washing evaluation program

    JP2021174488A