Dysphagia risk determination device, method, and program
The device estimates head orientation and neck hyperextension from front video images during meals to detect aspiration risk, addressing the challenge of identifying risky postures during eating and providing real-time alerts to prevent aspiration.
Patent Information
- Application Number
- JP2022153039
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-09-26
- Publication Date
- 2025-05-26
- Estimated Expiration
- 2042-09-26
AI Technical Summary
Existing technologies are unable to effectively and simply detect situations where aspiration risk occurs during meals, particularly for the elderly who may assume a hunched-back posture leading to neck hyperextension and increased aspiration risk.
A device and method that estimate the head orientation and neck hyperextension angle from a front video image captured during meals, determining the aspiration risk based on the neck hyperextension angle, which can be implemented using a smartphone or tablet with a camera.
Enables simple and effective detection of aspiration risk during meals, allowing for real-time alerts to prevent aspiration pneumonia, which is a significant health concern for the elderly.
Smart Images

Figure 0007682838000002 
Figure 0007682838000003 
Figure 0007682838000004
Abstract
Description
Technical Field
[0001] The present invention relates to an aspiration risk determination device, method, and program.
Background Art
[0002] As a method for automating monitoring for watching over whether a subject is in a safe state or the like, for example, there are the methods of Patent Documents 1 to 3.
[0003] Patent Document 1 relates to posture determination on a bedding such as a bed, and determines the sleeping posture of a user present on the bedding using a load sensor. Patent Document 2 estimates the human body posture from an image sensor composed of an infrared sensor, and the determinable postures are lying position on the bed, sitting position on the bed, sitting position at the edge of the bed, standing position, falling, staying in the room / not staying. Patent Document 3 relates to a posture detection method, and determines whether or not it is a predetermined posture based on the size, height, position, orientation of the head image detected by acquiring an image taken from above the user's head, and the positional relationship between the head and the torso.
[0004] In Patent Document 1, pressure sensors are arranged in a matrix on the human body support surface of the bedding, and in Patent Document 2, an infrared sensor is used. In both cases, since it is necessary to provide special equipment, there is a difficulty in simply realizing the automation of monitoring. In this regard, in Patent Document 3, since it can be determined from an image that can be taken with a camera or the like that is standardly provided in a mobile terminal such as a recent smartphone and can be generally used, the monitoring can be automated simply to some extent.
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Patent Document 2
Patent Document 3
Non-Patent Literature
[0006]
Non-Patent Literature 1
Non-Patent Literature 2
Non-Patent Literature 3
Non-Patent Literature 4
Non-Patent Literature 5
Non-Patent Literature 6
Summary of the Invention
Problems to be Solved by the Invention
[0007] However, in the prior art, as a matter to be monitored, it was not possible to simply deal with the fact that "one cannot tell oneself whether one's posture during a meal is in a posture that is prone to aspiration".
[0008] That is, Patent Document 1 determines the lying posture and does not deal with the meal time when taking a sitting posture. Similarly, Patent Document 2 also does not deal with the sitting posture during a meal. Furthermore, Patent Document 3 also determines the posture of monitoring targets such as falls and drops in nursing facilities and does not deal with the sitting posture during a meal. Also, although Patent Document 3 uses a general-purpose camera, since it takes pictures from above the head, it is necessary to install a camera on the ceiling or the like, and installation costs are also required.
[0009] Regarding aspiration and eating behaviors and postures that can cause aspiration, for example, they are studied in Non-Patent Documents 1 to 5 etc.
[0010] The elderly are prone to kyphosis (a posture where the thoracolumbar spine is overly lordotic, hunchback) due to spinal deformities caused by osteoporosis and decreased trunk muscle strength, which has an adverse effect on daily life. When kyphosis occurs during a meal, the head tends to tilt backward, and as a result, the neck is overly hyperextended, and the food and drink are more likely to enter the trachea on the anterior wall side instead of the esophagus on the posterior wall side of the pharynx, which becomes a factor in aspiration. Hyperextension of the neck is anatomically such that food and drink are likely to directly enter the bronchus when swallowing, and conversely, forward flexion during swallowing is said to be safe (Non-Patent Document 6).
[0011] The elderly living alone may assume a hunched-back posture during solitary meals and unknowingly adopt a feeding posture that makes them prone to aspiration, which may lead to aspiration pneumonia. In particular, it is said that there are many cases of solitary meals among the increasing number of elderly people living alone in recent years. Also, when there are meal assistants or nurses, they assist with positioning to correct the feeding posture. However, it is expected that the shortage of nursing staff will accelerate at the care site, and there is a possibility that sufficient attention cannot be paid to each individual elderly person.
[0012] If aspiration occurs, aspiration pneumonia is something that must be noted. Pneumonia is the leading cause of death among people aged 80 and above in Japan, and 30% of them are due to aspiration pneumonia in those requiring care. Since aspiration pneumonia is a dangerous disease that can lead to life-threatening situations for the elderly, prevention of aspiration is important.
[0013] As described above, prevention of aspiration is important, but the prior art has not been able to simply and effectively detect situations in which aspiration can occur.
[0014] In view of the problems of the above prior art, an object of the present invention is to provide an aspiration risk determination device, method, and program that can simply and effectively detect the aspiration risk.
Means for Solving the Problems
[0015] To achieve the above object, the present invention is an aspiration risk determination device, which includes a first process of estimating the head orientation of a user from a front image that is a frame image at each time of a front video obtained by photographing the user during feeding from the front side, a second process of estimating the neck hyperextension angle of the user as at least interlocking with the head orientation, and a third process of determining that the user is in a state where aspiration may occur based at least on a determination that the neck hyperextension angle has become large. Also, it is characterized by being a method and program corresponding to the device.
Effects of the Invention
[0016] According to the present invention, by analyzing a front image that can be captured with a simple camera arrangement and determining that the user is in a state where aspiration may occur based at least on the determination that the neck flexion angle has increased, the aspiration risk can be determined simply and effectively.
Brief Description of the Drawings
[0017]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Embodiments for Carrying Out the Invention
[0018] FIG. 1 is a functional block diagram of the aspiration risk determination device 10 according to an embodiment. As shown in the figure, the aspiration risk determination device 10 includes a face-related data extraction unit 31 including a photographing unit 1, a face feature point extraction unit 2, a neck backward bending angle estimation unit 3, an aspiration risk determination unit 4, an aspiration alert output unit 5, a front image analysis unit 311, and a side image analysis unit 312, a face-related data storage unit 32, a set value calculation unit 41, and an aspiration risk management table storage unit 42. The aspiration risk determination device 10 can be realized, for example, as a mobile terminal such as a smartphone having a camera. The photographing unit 1 can be configured as a camera that performs photographing as hardware.
[0019] FIG. 2 is a flowchart of the operation of the aspiration risk determination device 10 according to an embodiment. Hereinafter, while explaining each step S1 to S3 in FIG. 2, the details of each functional unit of the aspiration risk determination device 10 in FIG. 1 will be described.
[0020] In step S1, as the first pre-registration process related to aspiration of the user who uses the aspiration risk determination device 10, the following processes 1 and 2 are performed. (Process 1) By photographing the front image and the side image of the user's face with the photographing unit 1, or by obtaining these images separately prepared by photographing, the face-related data extraction unit 31 receives the input of the front image and the side image. (Process 2) By analyzing these front image and side image in the face-related data extraction unit 31, the face feature points and the neck angle are extracted as the face-related information of the user, and the extracted face feature points and neck angle are stored in the face-related data storage unit 32 as the pre-registration information of the user.
[0021] Note that the face-related data storage unit 32 can be realized as a storage medium as hardware, stores the pre-registration information, and refers to the pre-registration information for the processing of the face feature point extraction unit 2 and the neck backward bending angle estimation unit 3 when analyzing the video of the user during eating described later.
[0022] FIG. 3 schematically shows the arrangement when the imaging unit 1 (cameras CAM1 and CAM2 as hardware) of the dysphagia risk determination device 10 in the present embodiment images the user as a schematic example of the imaging method in process 1 in step S1. In FIG. 3, the state where the user U is sitting on the chair CH in front of the table TB in the same sitting posture as during eating is imaged by the cameras of the imaging unit 1 in step S1, and is schematically shown by a front view FV (a front view FV from slightly above rather than from the horizontal front of the user U to clarify the spatial arrangement), a side view SV (a side view SV from the right hand side where the user U holds a cup in the figure), and a plan view PV. In each of the views FV, SV, and PV, the three-dimensional space coordinates (x, y, z) are such that the xy plane is a horizontal plane parallel to the floor surface or the surface of the table TB, y is the direction facing the front of the user U's face (the depth direction at the front of the face), the x-axis is the direction facing the right hand side where the user U holds a cup in the figure, and the z-axis is the vertical direction (height direction). (Hereinafter, this three-dimensional space coordinate (x, y, z) will be described as the world coordinate.)
[0023] As shown in FIG. 3, in step S1, the cameras CAM1 and CAM2 are respectively arranged so that the portion BP from above the user U's chest to the face (the bust-up portion BP of the user U) is imaged in the front image and the side image. That is, the camera CAM1 is arranged in front of the user U's face and at the same height as the face (the camera CAM1 is arranged on the table TB, for example, to be at this height), and a front image is obtained by imaging from the true front direction D1 (y-axis direction) of the face. At the same time, the camera CAM2 is arranged on the side of the user U's face (either the right side or the left side, and the left side is used in this example) and at the same height as the face (the camera CAM2 is arranged on a pedestal PD separate from the table TB to be at this height), and a side image can be obtained by imaging from the true side direction D2 (x-axis direction) of the face.
[0024] In the process 1 of step S1, a front image and a side image may be obtained by arranging one camera in two ways, i.e., the cameras CAM1 and CAM2 in FIG. 3, and performing two shootings. When performing the two shootings, it is desirable that the user U to be photographed stands still without moving his / her posture so that the front image and the side image are taken in the same posture state of the user U as much as possible. Alternatively, in the process 1 of step S1, two cameras may be arranged like the cameras CAM1 and CAM2 in FIG. 3, and one shooting may be performed with each camera to obtain a front image and a side image. Also in this case, it is desirable that the front image and the side image are taken in the same posture state of the user U as much as possible by shooting simultaneously with the two cameras or the like.
[0025] In the process 2 of step S1, in the face-related data extraction unit 31, the front image analysis unit 311 extracts face feature points configured as a three-dimensional coordinate point group by analyzing the front image, and the side image analysis unit 312 estimates the neck angle by analyzing the side image, and saves these results in the face-related data storage unit 32 in a linked form. The details of the processing of each of the units 311 and 312 of the face-related data extraction unit 31 are as follows.
[0026] The front image analysis unit 311 extracts face feature points (a face feature point group FP as a plurality of face feature points whose coordinates in a three-dimensional space are determined) from the front image by using existing methods such as Non-Patent Documents 7 to 9 below, and sets the head angle θ as the orientation of the face feature point group FP in the three-dimensional space. H to set. [Non-Patent Document 7] Eye-Tack Solutions Co., Ltd., "Estimating the orientation of the face using head-pose-estimation", [searched on August 12, Reiwa 4], Internet <URL:https: / / www.itd-blog.jp / entry / peep-prevention-1> [Non-Patent Document 8] Arnaldo Gualberto, "Real-Time Face Pose Estimation with Deep Learning", [Searched on August 12, Reiwa 4], Internet <URL:https: / / medium.com / analytics-vidhya / face-pose-estimation-with-deep-learning-eebd0e62dbaf> [Non-Patent Document 9] Kazemi, V., & Sullivan, J. (2014). One millisecond face alignment with an ensemble of regression trees. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 1867-1874).
[0027] In the existing methods such as Non-Patent Documents 7 to 9, facial feature points are extracted by machine learning or deep learning networks. Since the accuracy is improved by using a large number of training images, for example, training images can be prepared from the facial image dataset "LFW" (Labeled Faces in the Wild), which is a standard dataset used worldwide, and the model parameters and the like can be pre-trained. The frontal image analysis unit 311 can extract facial feature points from the frontal image by using this pre-trained model.
[0028] Fig. 4 schematically shows the facial feature points and the head angle obtained by the processing of the frontal image analysis unit 311. As shown in the figure, a frontal image is obtained by the camera CAM photographing the face including the head H of the user U from the front. From this frontal image, as shown by white circles in the figure, as a predetermined facial feature point group FP on the face, for example, a plurality of points each of the eyebrow part, the contour part of the eyes, the nasolabial fold and the lower end of the nose, the contour part of the mouth, and the jaw part can be extracted as three-dimensional coordinate points in the three-dimensional camera coordinates (x c , y c , z c ) of the camera CAM.
[0029] From the group of face feature points FP extracted as three-dimensional coordinate points, the head angle θ H representing the orientation within the three-dimensional camera coordinates (x c , y c , z c ) of the camera CAM of the head H of the user U can be determined. Here, for the camera CAM in FIG. 4, by setting the arrangement when taking a front image to be the same as that of the camera CAM1 in FIG. 3, i.e., the true front of the face of the user U, the three-dimensional camera coordinates (x H c , y c , z c ) in FIG. 4 can be made to coincide with the three-dimensional world coordinates (x, y, z) in FIG. 3 (by considering that the directions of the respective coordinate axes coincide, the two coordinate systems are in a state where they can be considered to coincide with each other), and the head angle θ H can be determined as the orientation in the three-dimensional world coordinates (x, y, z) in FIG. 3.
[0030] The head angle θ H can be expressed, for example, as the following three angles (X, Y, Z). (In FIG. 4, the rotation axes for defining these angles X, Y, Z are drawn as X, Y, Z.) X: Angle in the nodding direction (pitch) Y: Angle when shaking the head left and right (roll) Z: Angle when tilting the head (yaw)
[0031] In the present embodiment, the head angle θ H in the front image analyzed by the front image analysis unit 311 in step S1 is set as the state in which the head H of the user U (the direction toward the top of the head) is vertical and the face is facing forward. That is, this head angle θ H is, when expressed by the above angles (X, Y, Z), θ H = (X, Y, Z) = (0, 0, 0).
[0032] By setting it in this way, the head angle θ H with the head vertically oriented and the face facing forward within the three-dimensional world coordinates (x, y, z) H =(X, Y, Z) = (0, 0, 0), the facial feature point group FP = FP(0, 0, 0) can be recorded in the face-related data storage unit 32. Since this facial feature point group FP(0, 0, 0) has the meaning of the origin as a reference when the user U's head is upright in the vertical direction and there is no inclination of the face, it is appropriately called the reference facial feature point group FP ref as such.
[0033] The side image analysis unit 312 measures the neck angle θ N by analyzing the side image. Specifically, for the neck angle θ N , after detecting the neck joint point a and the shoulder joint point b (either left or right) from the side image, the orientation of the straight line connecting these points a and b within the side image of the side image is measured as the neck angle θ N .
[0034] FIG. 5 schematically shows the processing of the side image analysis unit 312 as examples EX51 and EX52. Example EX51 is a schematic example of the joint extraction process for an image for detecting the neck joint point a and the shoulder joint point b as three-dimensional space coordinate points from the side image.
[0035] Any existing method for extracting skeletal joints as key points from a person image can be used for the joint extraction process. This key point extraction may be extracted as three-dimensional world coordinates or as two-dimensional image coordinates. Among the various joints extracted by key point extraction as in Example EX51, the neck joint point a and the shoulder joint point b may be obtained. In Example EX51, various joints are shown as being visible from the side of the front camera CAM1 in FIG. 3. However, as already described, in the side image analysis unit 312, as shown in Example EX52, from the side image taken by the side camera CAM2 in FIG. 3, the neck joint point a and the shoulder joint point b are obtained, and the orientation of the straight line Lab connecting a and b is set as the neck angle θ N . (Note that the neck joint point a and the shoulder joint point b are defined in advance in the key point extraction method used and do not necessarily coincide with medical joint points.)
[0036] As shown in Example EX52, the straight line Lab is defined in the yz plane parallel to the side surface of the user U as projected in the x-axis direction (ignoring the coordinate value in the x-axis direction) in the world coordinates (x, y, z), and the direction of the straight line Lab on this yz plane is defined as the neck angle θ N can be determined. That is, similar to FIG. 4 described for the front image, for the side image as well, like the camera CAM2 in FIG. 3, the relationship between the camera coordinates (x c , y c , z c ) and the world coordinates (x, y, z) is known (in the side image captured by the camera CAM2 in FIG. 3, the relationship is (x, y, z) = (y c , -x c , z c ). Therefore, the detection positions of the joint points a and b in the camera coordinates (x c , y c , z c ) can be obtained as the positions (x a , y a , z a ), (x b , y b , z b ) in the world coordinates respectively, and the direction of the straight line Lab connecting the point a (y a , z a ) and the point b (y b , z b ) on the yz plane can be determined as the neck angle θ N .
[0037] The above is the case where 3D keypoint extraction is applied. However, when keypoint extraction is applied as 2D image coordinates in the image coordinates (u, v) to obtain the joint points a and b, similarly, the direction of the straight line Lab can be determined as the neck angle θ N . That is, the u-axis (horizontal direction) of the image coordinates is made to coincide with the direction of the y-axis of the world coordinates, and the v-axis (vertical direction) is made to coincide with the direction of the z-axis of the world coordinates. The direction of the straight line Lab connecting the point a (u a , v a ) and the point b (u b , v bThe direction of the straight line Lab connecting them is defined as the direction in the uv plane (and also the direction in the yz plane), which is the neck angle θ N It can be determined as such. When extracting the key points in the two-dimensional image coordinates in this way and measuring the angle θ N from the straight line Lab, it is necessary to use the side image. On the other hand, when extracting the key points from the image in three-dimensional coordinates and measuring the angle θ N from the straight line Lab, it is not necessarily required to use the side image. However, even when extracting the three-dimensional coordinates, from the perspective of ensuring the measurement accuracy of the direction of the straight line Lab by accurately photographing the user from the side-facing direction through adjusting the camera arrangement during shooting, it is desirable to use the side image.
[0038] Regarding the neck angle θ N as shown in Example EX52, when the straight line Lab coincides with the vertical direction Lz (z-axis direction), it is set as 0° (the origin), and the neck angle θ N >0 in the direction of forward flexion from this vertical direction Lz can be defined accordingly.
[0039] The neck angle θ N defined in this way is stored in the face-related data storage unit 32 as the direction in the xz plane, which is the plane of the side of the user U, for use in estimating the neck hyperextension angle that is a factor in dysphagia.
[0040] FIG. 6 is a diagram showing exemplary patterns of the user in various states that are the recording and analysis targets for performing dysphagia determination by the dysphagia risk determination device 10. In FIG. 6, the user states are schematically drawn in the yz plane, which is the plane of the side of the user, and these correspond to examples of side images when taken with the side-mounted camera CAM2 in FIG. 3.
[0041] In FIG. 6, in Examples EX1A and EX1B, the head upright state of the first user U1 (the state in the front image and the side image acquired in step S1) and the state in which the head faces backward from this upright state (the head tilts backward in the direction of the pitch angle X), resulting in a state where hyperextension of the neck has occurred. Similarly, in Examples EX2A and EX2B, the head upright state of the second user U2 (the state in the front image and the side image acquired in step S1) and the state in which the head faces backward from this upright state (the head tilts backward in the direction of the pitch angle X), resulting in a state where hyperextension of the neck has occurred.
[0042] In each example of FIG. 6, point c is the position of the user's head top, and point d is the position of the user's external auditory meatus. When the straight line Lcd connecting these points c and d faces the vertical direction Lz, the user's head is in an upright state in the vertical direction. However, in this embodiment, instead of detecting the user's head top c and external auditory meatus d from the image, the reference face feature point group FP obtained by the front image analysis unit 311 in step S1 ref is regarded as the state where the user's head is in an upright state in the vertical direction when the user's head takes the posture. In each example of FIG. 6, for the sake of clearly expressing the tilt of the user's head, the head angle θ H of the user, a straight line Lcd is drawn. However, it should be noted that the head angle θ H of the user is obtained not by detecting points c and d, but as the posture of the face feature point group FP. (Regarding the case where the head angle θ H changes from 0° in the vertical state as in Examples EX1B and EX2B, it will be described in the video analysis process during eating, which will be described later.)
[0043] As schematically shown in FIG. 6, the first user U1 has no tendency of kyphosis, and the second user U2 has a tendency of kyphosis. Let the value of the neck angle θ N of the first user U1 be θ N [U1], and the value of the neck angle θ N of the second user be θ N [U2]. Then, "θ N [U1]<θ N [U2]", and the more the tendency of kyphosis, the larger the neck angle θ NThe value is assumed to be calculated as a large value. In the video analysis process during eating (Examples EX1B and EX2B in FIG. 6 are referred to in the explanation of this process), the risk of aspiration is determined. However, for the neck angle θ N the larger the value of NB the more the neck backward bending angle θ N is estimated to be large, and the information on the neck angle θ registered in step S1 is utilized in such a way that the tendency to be determined as having an aspiration risk becomes higher.
[0044] The above explains step S1 in FIG. 2. Next, in step S2, as the second pre-registration process, the input of basic information regarding the user's health condition is accepted. After calculating the set value from the basic information in the set value calculation unit 41, the basic information and the set value are stored in the aspiration risk management table storage unit 42 as the aspiration risk management table.
[0045] FIG. 7 is a schematic example of the aspiration risk management table TBL registered for each user. The input basic information is, for example, seven explanatory variables X = (x 1 , x 2 , …, x 7 ) regarding the user's health condition, where x 1 indicates whether the user is 65 years old or older, x 2 indicates the presence or absence of underlying diseases such as respiratory diseases, x 3 indicates the history of aspiration pneumonia, x 4 indicates the presence or absence of a kyphotic posture, x 5 indicates the presence or absence of dementia, x 6 indicates the presence or absence of dentures, and x 7 indicates the presence or absence of a tendency to eat quickly. Here, for any of the seven variables x i (i = 1, 2, …, 7), x i = 0 means "none," x i = 1 means "present," and it is an evaluation item that becomes "present" when the health condition tends to be poor. Other items similar regarding the health condition may be added to the explanatory variable X for accepting the input as basic information, or only a part of the seven items illustrated in FIG. 7 may be used.
[0046] In the set value calculation unit 41, as a function of the input explanatory variable X, the objective variable Y = f(X) is calculated, and the value of the objective variable Y is used as the set value. In the present embodiment, the three-item objective variable Y = (y 1 , y 2 , y 3 ) is calculated, where y 1 is the neck hyperextension angle threshold, y 2 is the alert output delay time, and y 3 is the alert duration time. Initial values B 1 , B 2 , B 3 can be set in advance for each of them. For the function f for calculating the objective variable Y, for example, a weighted sum of the following explanatory variables can be used.
[0047]
Equation
[0048] Regarding the set value of the neck hyperextension angle threshold y 1 when j = 1, for each explanatory variable x i , the weight coefficient W 1i (1 ≤ i ≤ 7) is set to a predetermined negative value, so that the worse the user's health condition is in the explanatory variable x i input as basic information, the value of the neck hyperextension angle threshold y 1 can be calculated to be smaller. For the initial value B 1 of the neck hyperextension angle threshold y 1 , for example, a value such as B 1 = 20° is set in advance.
[0049] Regarding the set value of the alert output delay time y 2 when j = 2, for each explanatory variable x i , the weight coefficient W 2i (1 ≤ i ≤ 7) is set to a predetermined negative value, so that the worse the user's health condition is in the explanatory variable x i input as basic information, the value of the alert output delay time y 2 becomes smaller (the alert output delay time y 2can be calculated so as to be shortened. The alert output delay time y 2 For the initial value B 2 for example, set a value such as B 2 = 3 seconds in advance.
[0050] Note that for the alert output delay time y 2 the weight coefficients W 2i (1 ≤ i ≤ 7) may be set so that the set value is 0 seconds or more. Alternatively, when the set value is calculated to be less than 0 seconds, additional processing may be performed to overwrite the set value with 0 seconds.
[0051] For the set value of the alert duration y when j = 3 3 for each explanatory variable x i the weight coefficient W 3i (1 ≤ i ≤ 7) is set to a predetermined positive value, so that the greater the health state of the user is on the bad side in the explanatory variable x i input as basic information, the greater the value of the alert duration y 3 (the alert duration y 3 becomes longer) can be calculated. For the initial value B 3 of the alert duration y 3 for example, set a value such as B 3 = 5 seconds in advance.
[0052] The aspiration risk management table storage unit 42 can be realized as a storage medium as hardware, and stores the basic information as these pre-registered information and the set values calculated from the basic information, and refers to the pre-registered information for the processing of the neck flexion angle estimation unit 3 and the aspiration risk determination unit 4 when performing video analysis during the user's eating described later.
[0053] Note that in the above, it is assumed that the input of the basic information regarding the user's health state is accepted, but this input may be omitted. When omitted, for each set value, the initial value B 1 , B 2 , B 3As something that corresponds thereto, it may be saved in the dysphagia risk management table storage unit 42, and for the basic information, it may be saved that there was no input.
[0054] Also, for the function that calculates from the basic information X with the set value as Y = f(X), in addition to the above weighted sum, a machine learning model or a deep learning network that has been learned in advance using a large number of learning data may be used.
[0055] The above has explained step S2 in FIG. 2. Next, in step S3, by using the information pre-registered in steps S1 and S2 respectively, with the video of the user during eating (the same user who obtained the pre-registered information in steps S1 and S2) taken from the front as the input, the dysphagia risk determination device 10 performs the process of determining the dysphagia risk of the user in real time, and the flow in FIG. 2 ends.
[0056] FIG. 8 is a flowchart showing a detailed example of step S3 in FIG. 2.
[0057] In step S31, after installing the camera, start shooting the video of the user during eating, and proceed to step S32. Specifically, by arranging the camera as the hardware constituting the imaging unit 1 in front of the user U during eating in the same way as the camera CAM1 in FIG. 3, a front video is shot.
[0058] For the sake of explanation, let the frame image (front image) at time t of the front video shot in real time at each time t = 1, 2, 3... in the imaging unit 1 be F(t). The imaging unit 1 outputs the frame image F(t) to the face feature point extraction unit 2 at each time t, so that the processes by the face feature point extraction unit 2, the neck hyperextension angle estimation unit 3, the dysphagia risk determination unit 4, and the dysphagia alert output unit 5 are repeatedly performed at each real-time time t. In FIG. 8, steps S32 to S37 represent the processes repeated at each time t, and step S38 represents the process of updating the real-time time t to the next time t + 1 (the process of managing the processing timing).
[0059] In step S32, after the frame image F(t) at the current time t obtained by the imaging unit 1 is read by the face feature point extraction unit 2, the process proceeds to step S33.
[0060] In step S33, after the face feature point extraction unit 2 extracts face feature points from the frame image F(t) and outputs them to the neck hyperextension angle estimation unit 3, the process proceeds to step S34.
[0061] In step S34, the neck hyperextension angle estimation unit 3 estimates the neck hyperextension angle θ NB (t) at the current time t from the difference (difference in orientation) between the face feature points of the frame image F(t) and the pre-registered face feature points (reference face feature point group), and outputs the neck hyperextension angle θ NB (t) to the dysphagia risk determination unit 4, and then the process proceeds to step S35.
[0062] The processing of these steps S33 and S34, similar to the processing of the front image analysis unit 311, estimates the value θ H of the head angle θ H (t) at the current time t by the methods of Non-Patent Documents 7 to 9 and the like, and then adds it to the neck angle θ N as a fixed value obtained by the side image analysis unit 312, so that the neck hyperextension angle θ NB (t) can be estimated as shown in Equation (1). θ NB (t)=θ N +θ H (t) …(1)
[0063] Examples EX1B and EX2B in FIG. 6 are schematic examples of Equation (1) when the head angle θ H (t) of each user is located in the hyperextension direction. As already described, the neck angle θ N as a fixed value is defined as positive on the side that bends forward from the vertical direction Lz. Also, the head angle θ H (t) is defined as positive on the side that bends backward from the vertical direction Lz. Equation (1) estimates the neck hyperextension angle θ NB (t) by adding based on these definitions, but the neck hyperextension angle θ NB(t) means the difference in orientation (the angle formed by the straight line Lab and the straight line Lcd) between the straight line Lab representing the orientation of the neck and the straight line Lcd representing the orientation of the head, as shown in Examples EX1B and EX2B.
[0064] In the process of the front image analysis unit 311, as shown in Examples EX1A and EX2A, the head angle θ H corresponding to the origin (0°) of (t) is the reference face feature point group FP ref is registered. In step S33, after similarly extracting the face feature point group FP(t) at the current time t from the frame image F(t), the angle change of the reference face feature point group FP in the face feature point group FP(t) at the current time t ref from the reference face feature point group FP H can be estimated as the head angle θ H (t) at the current time t. That is, if the orientation conversion by pitch, roll, and yaw of the head angle θ H (t)=(X(t),Y(t),Z(t)) is also denoted as "θ FP(t)=θ H (t)·FP ref
[0065] In actuality, due to noise and the like when detecting face feature points from the image, FP(t) and FP ref do not completely match in the orientation conversion as described above. However, as a method to minimize the conversion error, the head angle θ H (t)=(X(t),Y(t),Z(t)) can be estimated by the methods of Non-Patent Documents 7 to 9 and the like.
[0066] Regarding θ H (t) in Equation (1), among pitch, roll, and yaw (X(t),Y(t),Z(t)), by using the value of the pitch X(t) component, as a one-dimensional angle in the plane of the side view of the user (head), the neck flexion angle θ NB(t) may be estimated. The schematic example in FIG. 6 assumes that roll and yaw are 0°, or can be ignored as small values. However, even when roll and yaw values exist to a certain extent, the roll and yaw values can be ignored and only the pitch component can be focused on, and the calculation according to Equation (1) can be performed in the same manner as in FIG. 6.
[0067] In Equation (1), the neck angle θ N is treated as a constant value. In reality, the neck angle θ N also θ N = θ N (t) is assumed to change with time, but in this embodiment, on the premise that this time change is small, the neck angle θ N is treated as a constant value.
[0068] In step S35, the neck hyperextension angle θ NB (t) obtained in step S34 is referred to, including the history {θ NB (k)|k = 1, 2, …, t} up to the current time t, and compared with the threshold value of the aspiration risk management table pre-registered in the aspiration risk management table storage unit 42. Then, the aspiration risk determination unit 4 determines whether there is an aspiration risk for the user during eating at the current time t, and then proceeds to step S36.
[0069] Specifically, in step S35, among the aspiration risk management tables, the set value of the neck hyperextension angle threshold y 1 (denoted as θ NB閾値 to clarify that it is a threshold value) and the set value of the alert output grace time y 2 (when converted to the number of processing steps at each time t in FIG. 8, this set value y 2 is assumed to be the time length of N steps) are referred to, and when the following holds, it can be determined that there is an aspiration risk. “θ NB (t - k) ≥ θ NB閾値 ” holds for all k = 0, 1, 2, …, N - 1.
[0070] In other words, the alert output grace time y 2When the estimated value of the neck flexion angle remains above the threshold for a length equal to the set value (for N processing steps), it can be determined that there is a risk of aspiration.
[0071] Note that the alert output delay time y 2 The set value is calculated as a shorter value as the user's health condition is evaluated to be on the worse side as described above. For users with poor health and a high risk of aspiration, even if the state where the neck flexion angle is determined to be large continues for a short time, it can be determined that there is a risk of aspiration.
[0072] Also, the alert output delay time y 2 If the set value is 0 seconds, by setting N = 1 above, at the time t when "θ NB (t) ≧ θ NB閾値 ", it can be immediately determined that there is a risk of aspiration.
[0073] In step S36, it is determined whether (condition 1) the current time t already corresponds to a state where the aspiration alert output unit 5 described later has been continuously outputting an alert, or (condition 2) the current time t is in a state where the aspiration alert output unit 5 is not outputting an alert, and the determination result in step S35 is that there is a risk of aspiration. If it does not meet either of these conditions 1 and 2, the process proceeds from step S36 to step S38. If it meets either of these conditions 1 and 2, the process proceeds to step S37.
[0074] In step S37, after the aspiration alert output unit 5 outputs an aspiration alert, the process proceeds to step S38. Note that the aspiration alert output may start at the current time t or may have already started at a past time t - M (M > 1) and continue to be output at the current time t.
[0075] The aspiration alert output unit 5 refers to the alert duration y 3 in the aspiration risk management table, and from the time when the aspiration risk determination unit 4 determines that there is a risk of aspiration, this alert duration y 3Continuously output aspiration alerts over time. When the time y 3 has elapsed after the alert output, the aspiration alert output unit 5 may stop the alert output.
[0076] The alert output may be, for example, to generate a predetermined warning sound such as a buzzer sound or a siren sound through a speaker to give an auditory stimulus to the user, or to vibrate a vibration element such as a vibrator to give a stimulus such as an auditory stimulus to the user, or to give a visual stimulus to the user by flashing a display screen or a light.
[0077] When using an alert sound for the alert output, for example, use a voice that pronounces text such as "Caution for choking! Your face is tilted backward when swallowing. Please correct your posture." (Appropriately repeat the voice so that it continues until the alert duration y 3 ). Alternatively, use a display display output for the alert output, and display the text as a warning display on the display so that the user can see the warning text. You may use an alert output that combines these various warning means.
[0078] Note that the set value of the alert duration y 3 is calculated as a longer value as the user's health condition is evaluated to be on the worse side as described above. For users with a poor health condition and a high risk of aspiration, by continuing the alert output for a longer time, aspiration prevention can be made more reliable.
[0079] Note that in the flow of FIG. 8, when the aspiration alert output unit 5 is continuing the alert output, the processes of steps S32 to S35 may be omitted.
[0080] In step S38, the time t is updated to the real-time time t+1 which is the next processing timing, and the process returns to step S32, and the same process is repeated for the next time t+1. Although not described in FIG. 8, when the user finishes eating, the process of the aspiration risk determination device 10 may be terminated by receiving an input from the user to terminate the aspiration risk determination process.
[0081] As described above, according to the aspiration risk determination device 10 of the present embodiment, it is possible to simply perform aspiration risk determination by analyzing the front image of the user during eating, and further, when it is determined that there is an aspiration risk, an alert output is performed, so that it is possible to effectively prevent the user from aspirating. Furthermore, the following effects can also be achieved.
[0082] · According to the present embodiment, compared with the conventional one, the user himself / herself can take a video and place it on the dining table, and only the existing device owned by the user, such as a tablet or a smartphone on which the application of the present embodiment operates, needs to be prepared. Even without a meal assistant, it is possible to expect the effect of being aware of the aspiration risk (prevention of choking) and correcting the posture in real time at low cost. · Since it is determined in real time with a video instead of judging with a single photo, it is possible to cope with the entire eating behavior and its chronological changes. · By accumulating and analyzing the actual performance data of the aspiration alert output, the setting value of the "aspiration risk management table" can be changed to a value suitable for the characteristics of each user, and a greater effect on improving the eating posture can be expected.
[0083] Hereinafter, various supplementary examples, modification examples, addition examples, etc. will be described.
[0084] According to the aspiration risk determination device 10 of the present embodiment, since it is possible to prevent the user from aspirating during eating by alarm output, it is possible to contribute to Goal 3 of the Sustainable Development Goals (SDGs) led by the United Nations, "Ensure healthy lives and promote well-being for all people of all ages."
[0085] Taking the embodiments described above as basic examples, various modifications of these basic examples will be described.
[0086] <Modification Example 1> It may be possible to omit obtaining the side image and measuring the neck angle θ by the side image analysis unit 312. In this case, the following formula (1a) may be used instead of formula (1). In the basic example, the neck angle θ N is treated as 0°, which corresponds to treating the orientation of the head as coinciding with the neck retroflexion angle. N =0° and treating the orientation of the head as coinciding with the neck retroflexion angle. θ NB (t)=θ H (t) …(1a)
[0087] <Modification Example 2> In the basic example, the neck angle θ N was measured only once in the pre-registration and treated as a fixed value. However, by using the value θ N (t) that can change at each real-time time t for the neck angle θ N , the following formula (1b) may be used instead of formula (1). θ NB (t)=θ N (t)+θ H (t) …(1a)
[0088] In this case, two cameras, one for taking front images and one for taking side images, are used. In addition to analyzing the front video by the camera CAM1 in the arrangement of FIG. 3 in the same manner as in the basic example, the side video by the camera CAM2 is also analyzed in the same manner as the method of the side image analysis unit 312 at the time of pre-registration, so that the neck angle θ N (t) at each time t can be obtained. That is, by using the frame image F(t) at each time t of the side video as the user's side image and applying the same processing as the side image analysis unit 312 performs on the side image at the time of pre-registration, the neck angle θ N (t) at each time t can be obtained.
[0089] Note that in both the basic example and Modification Examples 1 and 2, the neck retroflexion angle θ NB (t) is at least the head angle θ HIt is estimated to be associated with (t).
[0090] <Modification 3> Regarding the arrangement of the cameras in the basic example In the basic example, as described with reference to FIG. 3 and the like, both the front image and the side image are captured by cameras arranged almost directly in front of and directly to the side of the user's face. However, if the error is within an acceptable range, arrangements that are generally in the front and generally to the side may also be acceptable. Also, even if the error is outside the acceptable range (assuming that the front and side of the user's face can be captured), the camera posture can be estimated by any existing method such as posture measurement using a posture sensor such as a gyro installed in the camera, or detecting a square marker or the like used in augmented reality display from the image to estimate the external parameters (camera posture) of the camera (if the camera is not fixedly installed and moves, it may be estimated in real time). By obtaining the conversion relationship between the camera coordinates (x c , y c , z c ) and the world coordinates (x, y, z), various angles and the like can be estimated in the world coordinates (x, y, z) in the same manner as when captured by cameras arranged directly in front of and directly to the side of the face.
[0091] <Hardware Configuration> FIG. 9 is a diagram showing an example of the hardware configuration in a general computer. The aspiration risk determination device 10 can be realized as one or more computer devices 70 having such a configuration. When the aspiration risk determination device 10 is realized by two or more computer devices 70, information necessary for processing may be transmitted and received via a network. The computer device 70 includes a CPU (Central Processing Unit) 71 that executes a predetermined instruction, a GPU (Graphics Processing Unit) 72 as a dedicated processor that executes some or all of the execution instructions of the CPU 71 instead of or in cooperation with the CPU 71, a RAM 73 as a main storage device (memory) that provides a work area for the CPU 71 (and the GPU 72), a ROM 74 as an auxiliary storage device (storage), a communication interface 75, a display 76, a mouse, a keyboard, an input interface 77 that receives user input by a touch panel or the like, a camera 81, a speaker 82, a light 83, a vibration element 84, and a bus BS for exchanging data between these components.
[0092] Each functional unit of the aspiration risk determination device 10 can be realized by the CPU 71 and / or the GPU 72 that reads and executes a predetermined program corresponding to the function of each unit from the ROM 74. Note that both the CPU 71 and the GPU 72 are a type of arithmetic unit (processor). Here, when display-related processing is performed, the display 76 operates in conjunction, and when communication-related processing related to data transmission and reception is performed, the communication interface 75 operates in conjunction.
[0093] When visually outputting an alert in the aspiration alert output unit 5, the display on the display 76, the lighting of the light 83, or the like may be used. When audibly outputting an alert, the speaker 82 may be used. When outputting an alert by generating vibration, the vibration element 84 composed of a vibrator or the like may be used. The face-related data storage unit 32 and the aspiration risk management table storage unit 42 can be realized in the ROM 74 or the RAM 73.
Description of Reference Numerals
[0094] 10… Aspiration risk determination device, 1… Imaging unit, 2… Facial feature point extraction unit, 3… Neck flexion angle estimation unit, 4… Aspiration risk determination unit, 5… Aspiration alert output unit, 31… Facial-related data extraction unit, 311… Frontal image analysis unit, 312… Lateral image analysis unit, 32… Facial-related data storage unit, 41… Set value calculation unit, 42… Aspiration risk management table storage unit
Claims
1. A first process of estimating the head orientation of the user from a front image which is a frame image at each time of a front video obtained by photographing the user during eating from the front side; A second process of estimating the neck hyperextension angle of the user as at least being interlocked with the head orientation; A third process of determining that the user is in a state where aspiration may occur, based at least on the determination that the neck hyperextension angle has become large. A device for determining aspiration risk, characterized by executing the above processes.
2. In the first process, a three-dimensional face feature point group of the user is extracted from the front image, and the orientation of the three-dimensional face feature points is determined with respect to the orientation of a reference three-dimensional face feature point group extracted from a captured image obtained by photographing the user from the front side in a state where the user's head is upright in the vertical direction, which has been registered in advance for the user. The device for determining aspiration risk according to Claim 1, characterized by estimating the head orientation as the change amount of the orientation of the three-dimensional face feature points.
3. In the first process, the head orientation is set to 0° when the user's head is upright in the vertical direction, and the head orientation is estimated such that the head orientation increases to a positive value when the user's head changes from the upright state to the hyperextension direction. In the second process, the device for determining aspiration risk according to Claim 1, characterized by estimating the neck hyperextension angle as a value equal to the value of the estimated head orientation.
4. The head orientation is an angle with the hyperextension side being positive. In the second process, the device for determining aspiration risk according to Claim 1, characterized by estimating the neck hyperextension angle as the sum value of a neck angle, which is an angle formed by the user's cervical joint and shoulder joint in a side view of the user and has the flexion side being positive and is registered in advance as a fixed value for the user, and the head orientation.
5. In the first process, further, from a side image which is a frame image at each time of a side video obtained by photographing the user during eating from the side, the neck angle, which is an angle formed by the user's cervical joint and shoulder joint in a side view of the user, is measured such that the flexion side is positive. The head orientation is an angle with the hyperextension side being positive. In the second process, the device for determining aspiration risk according to Claim 1, characterized by estimating the neck hyperextension angle of the user as the sum value of the neck angle and the head orientation.
6. In the third process, when the state where the neck backward bending angle exceeds the threshold value continues for a certain period, or when the neck backward bending angle exceeds the threshold value, it is determined that the user is in a state where aspiration may occur. The aspiration risk determination device according to claim 1, characterized in that.
7. Based on the values of one or more evaluation items regarding the health state of the user, the threshold value is set to be smaller as the health state of the user is evaluated to be on the worse side. The aspiration risk determination device according to claim 6, characterized in that.
8. Based on the values of one or more evaluation items regarding the health state of the user, the certain period is set to be shorter as the health state of the user is evaluated to be on the worse side. The aspiration risk determination device according to claim 6, characterized in that.
9. When it is determined in the third process that the user is in a state where aspiration may occur, a fourth process of generating an alarm by visual stimulation, auditory stimulation, or vibration generation for the user is further executed. The aspiration risk determination device according to claim 1, characterized in that.
10. In the fourth process, the alarm is continuously generated for a certain period, Based on the values of one or more evaluation items regarding the health state of the user, the certain period is set to be longer as the health state of the user is evaluated to be on the worse side. The aspiration risk determination device according to claim 9, characterized in that.
11. A first procedure for estimating the head orientation of the user from a front image which is a frame image at each time of a front video obtained by photographing the user during eating from the front side, A second procedure for estimating the neck backward bending angle of the user as at least interlocking with the head orientation, Based on at least the determination that the neck backward bending angle has become large, a third procedure for determining that the user is in a state where aspiration may occur. An aspiration risk determination method, characterized in that a computer executes the method.
12. An aspiration risk determination program, characterized in that a computer is caused to function as the aspiration risk determination device according to any one of claims 1 to 10.
Citation Information
Patent Citations
System and method for measuring ingesting action
JP2013031650A
Head and neck posture measurement device and method
JP2013150743A
Posture estimation device, posture estimation system, posture estimation method, posture estimation program, and computer-readable recording medium recording posture estimation program
JP2015231517A
Living body movement identification system and living body movement identification method
JP2018000871A
JPP6903368B