Care system, method for monitoring, and control program
The nursing care system uses video and audio analysis to generate captions explaining care recipient actions, addressing the limitations of existing systems and improving care staff efficiency.
Patent Information
- Application Number
- JP2024061377
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-05
- Publication Date
- 2025-10-17
AI Technical Summary
Existing care record systems struggle to accurately determine the purpose or cause of a care recipient's actions, as they rely on general-purpose words and cannot analyze consecutive images in chronological order, leading to increased burden on care staff.
A nursing care system that uses video data captured by an image capturing unit, determines specific situations through a trained model, and generates captions explaining the causes of these situations, including features like stereo cameras for height detection and audio analysis.
Facilitates easy understanding of care recipient situations by care staff, reducing the burden of manual record-keeping and enhancing the accuracy of care plans.
Smart Images

Figure 2025158635000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a care system, a monitoring method, and a control program. [Background technology]
[0002] Japan has seen a remarkable increase in life expectancy due to improvements in living standards, sanitary conditions, and medical standards that came with the rapid economic growth after the war. This, coupled with a declining birthrate, has led to an aging society with a high aging rate. In such an aging society, an increase in the number of people requiring care due to illness, injury, aging, and other factors is expected. In nursing care facilities such as hospitals and elderly welfare facilities (hereinafter simply referred to as "facilities"), care is provided to those receiving care by caregivers and nurses (hereinafter referred to as "care staff").
[0003] In facilities, care staff are required to record the care status of care recipients in care records on a daily basis in order to understand the health status and daily living situation of the care recipients and to use this information to create care plans. Recording in care records is important for providing adequate care to care recipients, but as the number of care recipients increases, the burden on care staff increases, which is an issue.
[0004] In this regard, Patent Document 1 discloses a care record creation system that aims to reduce the burden on caregivers involved in filling out care records or medical records. This care record creation system uses a sensor to detect the behavior or condition of the care recipient, converts the sensing information obtained from the sensor into information used in the care record, and records the information converted by the conversion means as text information in the care record.
[0005] Patent Document 2 also discloses a photobook creation system that facilitates editing of photobooks. This system includes a comment caption generation unit that generates and adds comments to selected images based on objects or text detected by an image analysis unit, and generates photobook data using the selected images and comments. [Prior art documents] [Patent documents]
[0006] [Patent Document 1] Japanese Patent Publication No. 2021-162927 [Patent Document 2] Japanese Patent Application Publication No. 2019-160186 Summary of the Invention [Problem to be solved by the invention]
[0007] The technology in Patent Document 1 determines and records text information (care record words) from a combination of characteristic information detected by sensors such as microphones and cameras, and requires that the care record words be set in advance from the combination of characteristic information. Only general-purpose words can be registered as care record words, and it is not possible to record the purpose or cause of the care recipient's actions.
[0008] Furthermore, the technology in Patent Document 2 extracts text from still images (photos), and does not analyze consecutive images in chronological order, and therefore cannot determine the intention behind the actions of people in the still images.
[0009] The present invention has been made in consideration of the above circumstances, and aims to easily grasp the situation of a care recipient who is being watched over.
[0010] The above object of the present invention can be achieved by the following means.
[0011] (1) an acquisition unit that acquires video data obtained by capturing an image of the care recipient in the room using an image capturing unit; a determination unit that determines whether a specific situation of the care recipient occurs in the room; a caption generation unit that generates a caption including a cause for the specific situation by inputting the video data acquired by the acquisition unit into a trained model that has been trained based on past data related to care; A care system comprising:
[0012] (2) When the determination unit determines that the specific situation has occurred, The nursing care system described in (1) above inputs video data of a predetermined time range including the time when the specific situation occurred into the trained model.
[0013] (3) The nursing care system described in (1) above, wherein the trained model is trained using training data that pairs video data containing a specific situation with captions that explain the specific situation for the video data.
[0014] (4) The care system according to (3) above, wherein the captions of the correct labels used in the training data are created by a care professional.
[0015] (5) The nursing care system described in (1) above, wherein the determination unit determines that the specific situation has occurred based on the video data.
[0016] (6) The care system according to (5) above, wherein the specific situation includes at least one of the care recipient getting out of bed, getting up, and falling.
[0017] (7) The care system described in (5) above, wherein the determination unit determines that a specific situation has occurred when it detects that the care recipient has been standing and has been stopped for a predetermined period of time or more.
[0018] (8) The photographing unit is a stereo camera arranged above the room, The care system described in (1) above, wherein the judgment unit extracts height information from the parallax of two video data, and determines that the specific situation has occurred when the time rate of change in the height of the care recipient is greater than or equal to a predetermined value.
[0019] (9) The nursing care system described in (1) above, wherein the judgment unit judges that the specific situation has occurred when a volume greater than a predetermined value is detected from audio data obtained by a sound collection unit that collects sound within the room.
[0020] (10) Further, an output unit for generating a care list that lists information on a specific situation including the caption is provided. Each time the specific situation occurs, video data is input to the trained model of the caption generation unit to generate the caption; and The care system according to (1) above, wherein the output unit adds information about the newly occurring specific situation to the care list.
[0021] (11) A care system as described in (10) above, wherein information on a specific situation in the care list is linked to video data for a predetermined time range including the time when the specific situation occurred.
[0022] (12) a step (a) of acquiring video data obtained by photographing the care recipient in the room with a photographing unit; (b) determining whether a specific situation of the care recipient occurs in the room; Step (c) of generating a caption including the cause of the specific situation by inputting the video data acquired in step (a) into a trained model trained based on past data related to caregiving; Monitoring methods, including:
[0023] (13) a step (a) of acquiring video data obtained by photographing the care recipient in the room with a photographing unit; (b) determining whether a specific situation of the care recipient occurs in the room; Step (c) of generating a caption including the cause of the specific situation by inputting the video data acquired in step (a) into a trained model trained based on past data related to caregiving; A control program for causing a computer to execute a process including the above. [Effects of the Invention]
[0024] The nursing care system of the present invention includes an acquisition unit that acquires video data obtained by capturing video of a care recipient in a room using a capturing unit, a determination unit that determines whether a specific situation has occurred for the care recipient in the room, and a caption generation unit that generates a caption including the cause of the specific situation by inputting the video data acquired by the acquisition unit into a trained model that has been trained based on past data related to care. This allows care staff and others to easily understand the situation of the care recipient they are watching over. [Brief explanation of the drawings]
[0025] Advantages and features provided by one or more embodiments of the present invention will be more fully understood from the following detailed description and the accompanying drawings, which are for purposes of illustration only and are not intended to be limiting. [Figure 1] FIG. 1 is a schematic diagram illustrating an information processing apparatus according to an embodiment of the present invention. [Figure 2] FIG. 2 is a block diagram showing the configuration of a detection unit. [Figure 3] FIG. 2 is a block diagram showing the configuration of a control device. [Figure 4] FIG. 10 is a schematic diagram showing a person / object detection process and a joint point detection process performed based on photographic data. [Figure 5] FIG. 1 is a block diagram showing a configuration of an information processing device. [Figure 6] 10 is a flowchart showing a process of registering a detected specific situation in a care list. [Figure 7A] 10 is a subroutine flowchart showing the process of detecting a specific situation in step S02 in the first process. [Figure 7B] 10 is a subroutine flowchart showing the process of detecting a specific situation in step S02 in the second process. [Figure 7C] 10 is a subroutine flowchart showing the process of detecting a specific situation in step S02 in the third process. [Figure 7D] 10 is a subroutine flowchart showing the process of detecting a specific situation in step S02 in the fourth process. [Figure 8] This is an example of a care list. [Figure 9] 10 is a flowchart showing additional learning of a trained model. [Figure 10] FIG. 10 is a diagram illustrating an example of a pair of video data and a correct label. [Figure 11] 10 is a flowchart showing a process for automatically generating captions from input video data. [Figure 12] FIG. 4 is a diagram showing an example of an operation screen displayed on a display unit. DETAILED DESCRIPTION OF THE INVENTION
[0026] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, the scope of the present invention is not limited to the disclosed embodiments. In the description of the drawings, the same elements are denoted by the same reference numerals, and duplicate explanations will be omitted. Furthermore, the dimensional proportions in the drawings are exaggerated for the convenience of explanation and may differ from the actual proportions.
[0027] 1 is a schematic diagram showing an information processing system 1000. The information processing system 1000 includes an information processing device 10, a detection unit 20, and a control device 30, and each device is communicatively connected to each other via a network 50. The information processing system 1000 functions as a care system.
[0028] The information processing device 10 is a server or a personal computer (PC). When the information processing device 10 is configured as a server, it may be an on-premise server installed in a nursing care facility or a cloud server using a commercial cloud service. The control device 30 is a PC operated by users such as nursing care staff (hereinafter simply referred to as staff) or administrators at the nursing care facility. The control device 30 is installed in the nursing care facility and functions as an edge server. The detection unit 20 will be described first, followed by the control device 30 and the information processing device 10. In the information processing system 1000 shown in FIG. 1, a detection unit 20 and a control device 30 are installed in multiple rooms in the nursing care facility. Each device in the nursing care facility is connected to the information processing device 10 via a network 50. That is, the information processing device 10 is connected to each device 20 and 30 in one nursing care facility. However, this is not limited thereto, and one information processing device 10 may be connected to multiple nursing care facilities, each of which has a control device 30 and multiple detection units 20.
[0029] (Detection unit 20) FIG. 2 is a block diagram showing the configuration of the detection unit 20. Referring to FIGS. 1 and 2, the detection unit 20 is configured to monitor the movement (behavior) of the care recipient 71 within the imaging area (room) of a room in a nursing care facility where the care recipient 71 is located, with the room being used as an imaging area. The care recipient 71 is a person receiving care and a resident of the nursing care facility (hereinafter also referred to as a "resident"). Care or nursing care is provided to the care recipient 71 by staff. The room is, for example, one room in a nursing care facility where multiple care recipients 71 reside. Furthermore, the information processing device 10 is connected to detection units 20 provided in each of multiple rooms as shown in FIG. 1. The device ID of each detection unit 20 and the subject ID of the care recipient 71 to be monitored are associated and stored in a memory unit (memory unit 32 or memory unit 12 described below).
[0030] As shown in FIG. 2 , the detection unit 20 includes a control unit 21, a communication unit 22, an imaging unit 23, a sensor 24, a care call unit 25, a microphone 26, and the like. The detection unit 20 generates sensing data (image data, audio data, etc.) by continuously sensing the care recipient 71 in the imaging area (room) in real time, 24 hours a day, using a plurality of various sensors. The control unit 21 may also include a large-capacity memory. The imaging unit 23, the sensor 24, and the microphone 26 (hereinafter simply referred to as the microphone 26) are disposed in a main body of the detection unit 20. The main body is disposed on the ceiling or upper part of a wall of the room. The care call unit 25 is disposed separately from the main body and is communicatively connected to the main body via wired or short-range wireless communication. Multiple microphones 26 may be disposed not only on the main body near the ceiling of the detection unit 20 but also next to the bed 81, for example. For example, the microphone 26 may also be provided inside the care call unit 25.
[0031] The control unit 21 is composed of a CPU, RAM, ROM, etc., and controls each part of the detection unit 20 and performs calculation processing according to a program. The communication unit 22 is an interface circuit (for example, a LAN card, etc.) for communicating with other devices such as the control device 30 via the network 50.
[0032] The photographing unit 23 is a camera composed of optical elements such as lenses and an imaging element such as a CMOS. For example, it is placed on the ceiling or the upper part of a wall of a room and photographs the interior of the room directly below as a photographing area from above. The photographing unit 23 is an example of a sensor that detects the movement of the care recipient 71 within the photographing area. Objects such as a bed 81 and a wheelchair 82 are present in the room. The photographing unit 23 photographs the care recipient 71, who is a person, as well as surrounding objects, and outputs photographed data (also referred to as video data). The video data is composed of a continuous time series of frame images (still image data). For example, the video data is a video at 30 FPS. The photographing unit 23 is a near-infrared camera, but a visible light camera may be used instead, or both may be used. The photographing unit 23 may also be a wide-angle camera.
[0033] As another example, the image capturing unit 23 may be a stereo camera. In this case, the image capturing unit 23 is composed of a first camera and a second camera. Video data obtained from one of these cameras is used for various processes described below. The first and second cameras are positioned to capture an image of the room in which the care recipient 71 resides, with their optical axes extending in a substantially vertical direction, providing a bird's-eye view. The optical axes of the first and second cameras are parallel to each other. The two optical axes are positioned a predetermined distance (hereinafter also referred to as the optical axis distance) apart in a direction perpendicular to the optical axes. The optical axis distance is several tens of centimeters, for example, 20 cm. A distance image is generated from two sets of image data captured simultaneously by the first and second cameras, with the distance value from the subject to the camera calculated based on parallax information stored in each pixel. Because this distance image is obtained from image data from a stereo camera with a bird's-eye view, it is also referred to as "height information" below. The process of generating a distance image from two sets of image data from the stereo camera can be performed using known technology. The conversion from the distance image to height information is performed using the following method. While distance values are defined by the distance from the camera, height information is the distance from a reference surface (floor). Therefore, height information is converted by subtracting the distance value of each pixel from the distance to the floor set for each pixel position (each direction) in the distance image. Regarding the distance to the floor, the information processing device 10 assumes that the room is a rectangular parallelepiped or a three-dimensional object made up of a combination of rectangular parallelepipeds. That is, it assumes that the floor is flat and perpendicular to the side walls, and estimates and stores the distance from the visible floor and side wall surfaces to the entire range of the floor in the X and Y directions on the floor plane, as well as the distance to the floor in each pixel direction in the distance image.
[0034] The sensor 24 detects the movement of the care recipient 71. The sensor 24 includes various sensors other than the imaging unit 23 and the microphone 26. For example, the sensor 24 may include at least one of a body movement sensor, a bed sensor, a mat sensor, a thermal sensor, and an infrared sensor. The "body movement sensor" may be a Doppler shift sensor that transmits and receives microwaves to and from the bed 81 and detects the Doppler shift of the microwaves caused by the body movement (e.g., breathing) of the care recipient 71. This body movement sensor detects chest movement (up and down movement of the chest) associated with the breathing of the care recipient 71. The period and amplitude of the body movement can be used to determine the sleep state or abnormal micro-movements (e.g., due to cardiac arrest). The "bed sensor" may be a sensor attached to the bed. For example, the bed sensor may detect weight and be placed on the bed 81 or on the floor at the exit of the bed 81 as an observation area to detect whether a person is standing on the sensor or to detect the sleep state. The "mat sensor" has the same function as the bed sensor, detecting the presence of a person in each divided area of the floor. The "infrared sensor," also known as a human presence sensor, detects whether a person is present in the room. For example, the infrared sensor is placed throughout the room or on the bed 81 as its observation area.
[0035] The care call unit 25 includes a push-button switch, and detects a care call when the switch is operated by the care recipient 71. The microphone 26 functions as a sound collection unit, collects sounds in the room where the care recipient 71 resides, and outputs audio data. The microphone 26 is an example of a sensor for determining a specific situation of the care recipient 71 within the observation area.
[0036] (Control device 30) FIG. 3 is a block diagram showing the configuration of the control device 30. As shown in FIG. 3, the control device 30 includes a control unit 31, a memory unit 32, a communication unit 33, a display unit 34, and an operation input unit 35. The control unit 31 is composed of a CPU, RAM, ROM, etc., and controls and processes each unit of the control device 30 according to a program. The control unit 31 functions as a determination unit 311. The memory unit 32 is composed of a hard disk, etc., which stores various programs and data. The communication unit 33 is an interface circuit for communicating with other devices via the network 50. The display unit 34 is composed of an LCD display, a touch sensor, etc., and displays various information and operation screens. The operation input unit 35 is an input device, such as a keyboard or mouse, and accepts input from staff, staff leaders of units (also referred to as areas or groups), and facility managers (hereinafter collectively referred to as managers, etc.). Furthermore, as described below, managers, etc. can check the care list via the control device 30. The care list includes a classification of a specific situation automatically generated by the information processing system 1000, the circumstances under which it occurred (date and time, name of the person receiving care, etc.), video data, and captions for that video data (see Figure 12 described below).
[0037] (Judgment unit 311) The determination unit 311 determines whether a particular situation has occurred based on various types of multimodal information such as video data, altitude information, and audio data.
[0038] The determination unit 311 determines whether a specific situation has occurred by the following first to fourth processes (FIGS. 7A to 7D, which will be described later). The first to third processes are for determining that a specific situation has occurred based on video data. The fourth process is for determining that a specific situation has occurred based on audio data. Furthermore, when the determination unit 311 determines that a specific situation has occurred in a room, it extracts video data of the room taken over a predetermined period including the time when the specific situation occurred (the time of taking the video or the time of recording the video), and registers this data in association with the specific situation.
[0039] The first process is a process of determining whether the care recipient 71 has woken up, gotten out of bed, or fallen, based on classification of a person's behavior scene determined based on the photographed data from the photographing unit 23.
[0040] The second process is a process for determining whether the care recipient 71 has stopped, similarly performed by classifying human behavior scenes. The classification of human behavior scenes in the first and second processes will be described later with reference to FIG.
[0041] The third process is performed using height information obtained by a stereo camera to determine a sudden change in the height of the care recipient based on the rate of change over time.
[0042] The fourth process is a process in which a determination is made based on a loud sound (a volume equal to or greater than a predetermined value) from the microphone 26.
[0043] (Determination process based on classification of behavioral scenes in the first and second processes) 4 is a schematic diagram showing the person / object detection process and the joint point etc. detection process performed based on the captured image data. The determination unit 311 detects the occurrence of a specific situation regarding the care recipient 71 by performing the following person / object detection process, joint point etc. detection process, and behavior scene analysis process as image analysis processes of the captured image data.
[0044] (Human and object detection processing) The determination unit 311 detects a person rectangle 710 and object rectangles 810 and 820 corresponding to people and non-human objects, respectively, from the captured data using a person / object detection process described below. The person rectangle 710 is an area within a rectangle (dashed line frame) that includes the care recipient 71 in the captured data. The object rectangles 810 and 820 are areas within a rectangle (dashed line frame) that includes specific objects other than people (e.g., a bed, wheelchair, or chair).
[0045] In the person / object detection process, areas in the captured data where objects, including people, exist are detected as object presence areas, and a reliability score (also called likelihood; hereafter, the reliability score will be referred to simply as score) is calculated for each predetermined category of objects contained in the detected object presence areas. The score is the likelihood of the target object. The person / object detection process can calculate the score using known technology using a DNN (Deep Neural Network).
[0046] The predetermined category may be, for example, a person, a chair, furniture, and bedding. In the person / object detection process, the object existence region with the highest score in the person category is detected as a person rectangle 710. That is, in the person / object detection process, the care recipient 71 is detected as an object (moving object). Similarly, the object existence region with the highest score in a predetermined object category is detected as an object rectangle 820 (e.g., an object region of a wheelchair) of the category with the highest score.
[0047] Alternatively, as another example of a method for detecting the human rectangle 710, a background subtraction method may be used, which extracts the difference between the photographic data of the detection target and a background image of the photographed area that has been extracted in advance by the fixed photographing unit 23. Also, an inter-frame (temporal) subtraction method may be used, which extracts the difference between the photographic data of the detection target and an average of past photographic data.
[0048] By such human / object detection processing, coordinates (object position coordinates) on the shooting data calculated based on the object rectangle 810 of an object other than a human are output. This output coordinate information is stored in the storage unit 12 or the storage unit 32 as time-series object information in association with each frame. The object position coordinates (object coordinates) may use the center position of the object rectangle or the positions of two or more vertices that are diagonally opposite each other.
[0049] (Joint point detection processing) In the joint point etc. detection process, the determination unit 311 detects a head rectangle 720 of the care recipient 71 and multiple feature points 730 (see FIG. 4 for both). More specifically, in the joint point etc. detection process, an area including the head of the care recipient 71 is detected (estimated) as the head rectangle 720 from the human rectangle 710. In addition, in the joint point etc. detection process, feature points 730 including joints related to the body of the care recipient 71 are detected (estimated) from the human rectangle 710. In FIG. 4, the positions of the multiple feature points 730 are indicated by open circles. Note that thick lines connecting the joint points (between the feature points 730) indicate the skeleton (bones). The feature points 730 include, for example, the head, neck, shoulders, elbows, hands, waist, thighs, knees, and feet. The feature points 730 may include feature points 730 other than those described above, or may not include any of the above. In the joint point etc. detection process, the feature points 730 of the care recipient 71 can be estimated from the human rectangle 710 using a DNN that reflects a dictionary for detecting feature points 730 from the human rectangle 710. For example, in the joint point etc. detection process, the feature points 730 can be estimated using a known technique that uses a DNN. In the joint point etc. detection process, the feature points 730 can be output as their respective coordinates on the imaging data. The coordinates of the human rectangle 710, the head rectangle 720, and the feature points 730 (joint point coordinates, head coordinates) are associated with each other for each frame of the imaging data and stored in the storage unit 32 as chronological human information.
[0050] The determination unit 311 determines the behavior of a person based on changes in the position of the person's joint points or skeleton over time. The determination unit classifies behavior scenes using a trained model. This trained model is trained through supervised learning using time-series information on people (joint point coordinates, head coordinates), object information (object position coordinates of objects such as wheelchairs and beds), and correct labels for behavior scenes. For example, the trained model is stored in the storage unit 32. This learning is performed using a recurrent neural network (RNN) method or a hidden Markov model (HMM) method. The behavior scenes to be determined include getting up, getting out of bed, falling, and tripping. The behavior scenes to be determined also include standing, sitting, walking, and stopping walking (stopping). Getting up is the behavior of getting up from bed 81, getting out of bed is the behavior of getting up and then leaving bed 81, and falling is the behavior of rolling on the floor. When a specific behavioral scene is determined from the video data from the image capture unit 23, it is notified to the staff as an event. Furthermore, among the determined behavioral scenes, specific behavioral scenes such as getting up, getting out of bed, and falling are recorded as specific situations in the care list, linked to the occurrence circumstances such as time, room number, and name of the care recipient. The room number of the room can be obtained from the location information (room number) where the image capture unit 23 is installed. The name of the care recipient can be traced from the room number and the care recipient information shown below.
[0051] (Storage unit 32) The storage unit 32 stores video data, care recipient information, staff information, care records, a care list, a trained model, etc. The care list includes a list of specific situations determined by the determination unit 311. The care list will be described later (see FIGS. 8 and 12).
[0052] The "care recipient information" is information about the care recipient 71, including the room number in which the care recipient 71 resides, family information, and the age, gender, medical history, and medical conditions of the care recipient 71. The care recipient information also includes information about the mental and physical condition and information about the care recipient's lifestyle. Examples of the mental and physical condition information include the height, weight, and level of care required of the care recipient 71. The mental and physical condition information may also include the level of independent walking and the level of transfer from the bed to a care chair. The mental and physical condition information may also include the resident's dementia level. The dementia level may be input by a facility manager or the like. Alternatively, the control device 30 may determine the level of unsteadiness of the care recipient 71 when moving from images captured by the imaging unit 23, and use this information to evaluate and determine the dementia level. Information about the resident's lifestyle may include the resident's activity level (travel time, travel distance), sleep time (life rhythm), and number of care calls. Medical history is past information about medical events and health, such as past illnesses, injuries, surgeries, treatments, and allergies.
[0053] "Staff information" includes information such as staff name, affiliated unit, and work schedule.
[0054] "Nursing care records" include nursing record sheets, nursing record sheets (medical charts, life records), and diagnostic charts that record the care and treatment provided by staff, caregivers, nurses, and doctors (hereinafter referred to as "staff, etc.") who provide care to the care recipient 71. These records may be entered as text by scanning sheets handwritten by the staff, etc., and performing OCR processing. Alternatively, they may be entered by the staff, etc., through a staff terminal such as a smartphone that the staff, etc., carries while on duty. Care records include the amount of medication administered to the care recipient 71, medication care related to medication dosage, the amount of food eaten by the care recipient 71, dietary and hydration care related to fluid intake, the amount and quality of excretion, vital values (body temperature, blood pressure, pulse rate, etc.), and the time of occurrence of these events.
[0055] The video data is acquired by the image capturing unit 23 of the detection unit 20 in each room. The video data is retained for a predetermined period of time. Furthermore, when the occurrence of a specific situation (hereinafter also referred to as an event) is determined, video data of a predetermined time range including the time of occurrence (also referred to as the event start point) linked to the specific situation is extracted, recorded, and retained for a long period of time (see the care list shown in Figures 8 and 12 described below). For example, the predetermined time range is based on the time of occurrence (occurrence time), and is a total of two minutes from one minute before to one minute after the occurrence. This predetermined time range is set in advance. This set range is not limited to this, and may be video data of any time range. For example, it may be 30 seconds before and after, or a total of one minute from one minute before to the time of occurrence.
[0056] The trained model is used by the determination unit 311 as described above, and is useful for estimating joints and bones and classifying action scenes.
[0057] (Information processing device 10) The configuration of the information processing device 10 will be described below with reference to Fig. 5. Fig. 5 is a block diagram showing the configuration of the information processing device 10.
[0058] 5, the information processing device 10 includes a control unit 11, a storage unit 12, and a communication unit 13. The control unit 11 is composed of a CPU, RAM, ROM, etc., and controls each unit of the information processing device 10 and performs calculation processing according to a program. The control unit 11 functions as an acquisition unit 111 and an output unit 113 in cooperation with the communication unit 13. The control unit 11 also functions as a caption generation unit 112.
[0059] The storage unit 12 is configured with a hard disk that stores various programs and various data, etc. The communication unit 13 is an interface circuit for communicating with other devices via a network 50.
[0060] A trained model for generating captions is stored in the storage unit 12. This trained model is a transform-type model based on DL (deep learning) that performs natural language processing. The trained model will be described later.
[0061] (Control unit 11) (Acquisition part 111) The acquisition unit 111 acquires video data obtained by imaging by the imaging unit 23 of the detection unit 20. This video data is video data for a predetermined time period linked to a specific situation whose occurrence has been determined by the determination unit 311.
[0062] (Caption generation unit 112) The caption generation unit 112 uses a machine-learned model to generate captions including the causes of specific situations when video data is input. The trained model of the caption generation unit 112 may be generated from scratch. In this case, the caption generation unit 112 is trained using commonly used learning data for generative AI (also referred to as teacher data or training data) and learning data that is a pair of correct-labeled captions for captured data that is unique human behavior monochrome top images acquired by the image capture unit 23. Alternatively, a pre-trained trained model may be used and additionally trained (fine-tuned) using learning data that is a pair of correct-labeled captions for captured data that is unique human behavior monochrome top images acquired by the image capture unit 23.
[0063] (output unit 113) The output unit 113 generates a care list. The care list is a list of information about specific situations, including captions for the specific situations. For example, by checking the care list via the control device 30, a manager or the like can immediately understand what specific situation occurred in the care recipient by checking the caption. In addition, the care list is registered with video data corresponding to a predetermined time range, including the time when the specific situation occurred, and the manager or the like can check the video data by operating an icon button indicating the video data on the display screen of the control device 30 (button b1 in FIG. 12, described below).
[0064] (Detection of specific situations) The process of detecting a specific situation will be described below with reference to Figures 6, 7A to 7D, and 8. Figure 6 is a flowchart showing the process of registering a detected specific situation in the care list.
[0065] (Step S01) The inside of the room is photographed by the photographing unit 23.
[0066] (Step S02) The determination unit 311 determines whether or not a specific situation has occurred from the video data captured by the image capture unit 23 or the audio data captured by the microphone 26.
[0067] The processing of step S02 will be described with reference to subroutine flowcharts corresponding to various processes in FIGS. 7A to 7D.
[0068] (First process) 7A is a subroutine flowchart showing the process of determining the occurrence of a specific situation by the first process. The specific situation regarding getting up, getting out of bed, and falling determined by this first process is also referred to as "specific situation 1."
[0069] (Step S211) The control unit 31 acquires the video data captured by the imaging unit 23.
[0070] (Steps S212 and S213) The determination unit 311 detects objects from video data. Objects include people and non-human objects such as beds, chairs, etc. In the case of people, the movement is monitored.
[0071] (Step S214) If getting up, getting out of bed, or falling is detected (YES), the determination unit 311 proceeds to step S215, and if not, skips step S215 and ends the subroutine flowchart (return).
[0072] (Step S215) The determination unit 311 sets a flag 1n (11 to 13) for a specific situation corresponding to getting up, getting out of bed, or falling. The above is the first process.
[0073] (Second process) 7B is a subroutine flowchart showing the process of determining the occurrence of a specific situation by the second process. The specific situation of continued standing still in a standing position, i.e., stopping, determined by the second process, is also referred to as "specific situation 2."
[0074] (Steps S221 to S223) The processing here is the same as steps S211 to S213, and the description thereof will be omitted.
[0075] (Step S224) If the care recipient 71 is in an upright position (YES), the determination unit 311 advances the process to step S225. If the care recipient 71 is not in an upright position (NO), the subroutine process ends (returns).
[0076] (Step S225) If the behavior of the care recipient 71 has started walking and then stopped (YES), the determination unit 311 proceeds to step S226. On the other hand, if the behavior of the care recipient 71 has not started walking or if the care recipient 71 is still walking (NO), the determination unit 311 continues the processing.
[0077] (Step S226) The determination unit 311 starts the timer, and if the duration of the stoppage is equal to or longer than the predetermined time (YES), the process proceeds to step S227. If the duration is less than the predetermined time (NO), the subroutine process ends (return). The predetermined time can be set for each individual care recipient 71, for example. For example, for a given care recipient 71, the duration when the care recipient 71 goes from walking to stopping over a predetermined period is monitored, and the mean + 2σ (standard deviation) is calculated from the variation in this duration, and this is used as the predetermined time.
[0078] (Step S227) The determination unit 311 sets flag 2 for the specific situation. The above is the second process. This second process is a process for recording situations where patients with advanced dementia often stop in their tracks.
[0079] (Third Processing) 7C is a subroutine flowchart showing the process of determining the occurrence of a specific situation by the third process. The third process is a process when the imaging unit 23 is a stereo camera. The specific situation regarding the time rate of change in the height of the care recipient 71 determined by this third process is also referred to as "specific situation 3."
[0080] (Step S231) The control unit 31 acquires video data captured by the first and second cameras of the imaging unit 23.
[0081] (Step S232) The determination unit 311 detects an object from video data captured by one of the pair of cameras (for example, the first camera). The process here is the same as that in step S212.
[0082] (Step S233) The determination unit 311 extracts a pair of frame images (still images) taken at the same time from the two video data, generates a distance image from the two image data based on the parallax information, which indicates the distance value from the camera position, and converts this into height information.
[0083] (Step S234) The determination unit 311 detects and monitors the height of the care recipient 71 by matching the position information of the care recipient 71 in the video data with the height information. For example, it acquires the height value of the distance image (height information) corresponding to the area of the head rectangle (see FIG. 4) of the care recipient 71.
[0084] (Step S235) If the time rate of change of height is equal to or greater than a predetermined value (YES), the determination unit 311 proceeds to step S236. If the change rate is less than the predetermined value (NO), the subroutine processing ends (returns). The time rate of change of height is equal to or greater than a predetermined value, for example, when the height suddenly decreases due to a fall. The time rate of change of height may be determined to be equal to or greater than a predetermined value when the height value is acquired every 0.5 seconds (15 frames) or every second (30 frames), for example, and when the change in height before and after is equal to or greater than a predetermined value, the time rate of change of height is determined to be equal to or greater than a predetermined value.
[0085] (Step S236) The determination unit 311 sets flag 3 for the specific situation. The above is the third process.
[0086] (Fourth process) 7D is a subroutine flowchart showing the process of determining the occurrence of a specific situation by a fourth process. The fourth process is a process using the microphone 26. The specific situation determined by this fourth process, which is a volume equal to or greater than a predetermined value, i.e., a loud volume, is also referred to as "specific situation 4." Note that the time information added to the audio data from the microphone 26 is synchronized with the time information added to the video data obtained by the imaging unit 23.
[0087] (Step S241) The control unit 31 acquires the sound data collected by the microphone 26 in the room.
[0088] (Step S242) The determination unit 311 determines whether the volume is equal to or greater than a predetermined value. If it is equal to or greater than the predetermined value (YES), the process proceeds to step S243. If it is less than the predetermined value (NO), the subroutine process ends (returns). The threshold value here is set to a level that can determine, for example, when the care recipient 71 has dementia and is yelling loudly or making strange noises, or when the care recipient 71 is banging hard on a wall. This predetermined value for volume can also be set for each individual care recipient 71. For example, the volume of the target care recipient 71 is monitored over a predetermined period, and the average + 2σ (standard deviation) is calculated from the variation in volume, and this is used as the predetermined value for volume. The sampling period for volume monitoring can be set appropriately, for example, to one minute.
[0089] (Step S246) The determination unit 311 sets flag 4 for the specific situation. The above is the fourth process.
[0090] (Step S03) Referring again to FIG. 6, the control unit 31 records the specific situation corresponding to the specific situation flag in the care list. Specifically, the control unit 31 links the extracted video data of a predetermined time and records it in the care list. Here, the video data of a predetermined time refers to the time tx (the shooting time or the audio recording time) when the occurrence of the specific situation is determined as the starting point, from a first predetermined time before the occurrence time tx to a second predetermined time after the occurrence time tx. For example, the first predetermined time is any time in the range of 30 to 120 seconds, and the second predetermined time is any time in the range of 0 to 120 seconds. In the following description, it is assumed that both the first and second predetermined times are one minute. In this case, the video data is two minutes long, one minute before and one minute after the occurrence time tx.
[0091] Figure 8 is an example of a care list. In Figure 8, captions have not yet been generated. Each row in the care list indicates a specific situation. Each specific situation is recorded with an automatically assigned unique ID, a flag (FLG) number corresponding to the specific situation, the specific situation corresponding to the flag, the room number indicating the situation in which it occurred, the name of the person receiving care, and the time of occurrence. In addition, video data for a specified period (2 minutes) including the time of occurrence tx is linked and recorded.
[0092] (Additional training of a trained model) In the following, we will explain how to use a pre-trained trained model that performs natural language processing and additionally train it using training data that is a pair of captions with correct labels for the photographic data, which is a unique human behavior monochrome top image acquired by the photographing unit 23.
[0093] Fig. 9 is a flowchart showing additional training of a trained model. Fig. 10 is a diagram showing an example of a dataset of video data and correct labels included in the training data. The video data v01 to v03 and the correct label descriptions w01 to w03 in Fig. 10 are each pairs of training data. Note that the number of datasets used for the training data is preferably several hundred to several thousand.
[0094] The video data v01 to v03 are past video data related to care that was actually captured and acquired by the imaging unit 23 at the facility. This video data is video data of one minute before and after (two minutes in total) the time when the specific situations occurred, acquired when the determination unit 311 determined the above-mentioned specific situations 1 to 4. In other words, the video data includes the specific situations. The length of this video data used for learning and its relationship to the time of occurrence are preferably the same as those of the video data extracted in step S03, but may be different. In addition, the correct answer label includes the cause of the specific situation. For example, in the explanation w01, specific situation 1 is determined, and the cause of the specific situation is described as "sitting in a chair in the room watching TV."
[0095] The explanations for the correct labels in this training data, such as w01, were created by experts in the nursing care industry. The experts here refer to medical professionals such as nursing staff, physical therapists, and occupational therapists who are normally involved in nursing care.
[0096] (Step S31) The control unit 11 functions as a learning machine. The control unit 11 acquires learning data. The learning data is a set of video data containing a specific situation generated as described above and a correct caption (explanatory text).
[0097] (Step S32) The control unit 11 performs additional learning on the pre-trained trained model using the training data acquired in step S31, thereby updating the parameters.
[0098] (Step S33) The control unit 11 updates the trained model and stores it in the storage unit 12.
[0099] (Automatic caption generation process) Next, the automatic caption generation process performed automatically by the information processing system 1000 will be described with reference to Fig. 11 and Fig. 12. The automatic generation process shown in Fig. 11 is performed at predetermined intervals, such as every day or every hour.
[0100] (Step S51) The control unit 11 acquires the care list stored in the storage unit 32.
[0101] (Step S52) The acquisition unit 111 of the control unit 11 acquires video data linked to a specific situation in the recording list for an unanalyzed specific situation. For example, in the example of Fig. 8, video data id10001 to id10006 are acquired in order. For example, video data VD0101 is acquired first.
[0102] (Step S53) The caption generation unit 112 inputs the video data into the trained model acquired in step S52 and generates captions. For example, as shown in Fig. 10, by inputting video data v01 to v04, captions such as explanatory sentences w01 to w04 are obtained, respectively. These captions are output by detecting the care recipient (person) and other objects from the video data and based on the relationship between the person and the objects over time.
[0103] (Step S54) The output unit 113 of the control unit 11 adds the caption obtained in step S53 to the care list.
[0104] (Step S55) If there is a specific situation where the caption has not been analyzed (YES), the process from step S51 onwards is repeated. On the other hand, if there is no specific situation where the caption has not been analyzed (NO), the process ends (END).
[0105] FIG. 12 is a diagram showing an example of an operation screen 341 displayed on the display unit 34. The operation screen 341 displays a care list to which captions have been added by the processing of FIG. 11. By checking the captions for specific situations that have occurred, the manager or the like can immediately understand what situation occurred in the care recipient being monitored. In addition, for each specific situation, the manager or the like can view the associated video for a predetermined period of time (e.g., two minutes) by operating button b1. For example, if a fall occurs as a specific situation, the manager or the like can understand the circumstances under which the fall occurred by viewing the video.
[0106] As described above, the care system according to this embodiment includes an acquisition unit that acquires video data obtained by capturing video of a care recipient in a room using a video capture unit. The care system also includes a determination unit that determines whether a specific situation occurs in the room for the care recipient, and a caption generation unit that generates a caption including the cause of the specific situation by inputting the video data acquired by the acquisition unit into a trained model that is trained based on past data related to care. This allows the situation of the care recipient being monitored to be easily understood simply by checking the caption.
[0107] In the past, even if the occurrence of a specific situation was determined and recorded, it was only possible to determine what type of specific situation occurred, such as waking up. In other words, since care staff did not know what caused the person to wake up, they could not determine whether or not they needed to take action, and ultimately had to visit the room of the person being cared for or rewatch the video. In response to such situations, the care system of this embodiment makes it easy to understand the situation simply by checking the captions in the care list. Furthermore, care records can be easily created by referring to and transcribing the care list.
[0108] The configuration of the information processing system 1000 described above is a main configuration for explaining the features of the above embodiment, but is not limited to the above configuration and can be modified in various ways within the scope of the claims. Furthermore, configurations that are included in general information processing systems are not excluded.
[0109] 1 and the like, the control device 30 and the information processing device 10 are described as separate entities, but they may be integrated. For example, the control device 30 functions as the information processing device 10. Alternatively, the functions of the determination unit 311 of the control device 30 may be provided on the control unit 11 side of the information processing device 10.
[0110] Furthermore, if the control device 30 has the functions of the caption generation unit 112 of the information processing device 10, the control device 30 may process the information in real time. In this case, the determination unit 311 may record a specific situation in the care list each time it occurs (each time it makes a determination), and input video data linked to this specific situation into the caption generation unit to generate a caption. The generated caption is then recorded in the care list.
[0111] In addition, in the present embodiment, the determination unit 311 determines the first to fourth processes (FIGS. 7A to 7D) as a determination of a specific situation, but it may perform only one or some of these processes. For example, the determination unit 311 may perform only the first process (FIG. 7A).
[0112] Furthermore, the means and methods for performing various processes in the information processing device or information processing system according to the above-described embodiments can be realized by either a dedicated hardware circuit or a programmed computer. The program may be provided, for example, by a computer-readable recording medium such as a USB memory or a DVD (Digital Versatile Disc)-ROM, or may be provided online via a network such as the Internet. In this case, the program recorded on the computer-readable recording medium is typically transferred and stored in a storage unit such as a hard disk. The program may also be provided as standalone application software, or may be incorporated as a function into the software of a device such as a detection unit.
[0113] While embodiments of the present invention have been described and illustrated in detail, the disclosed embodiments are made for purposes of illustration and example only and are not intended to be limiting, and the scope of the present invention should be construed by the language of the appended claims. [Explanation of symbols]
[0114] 1000 Information Processing Systems 10. Information processing equipment 11 Control section 111 Acquisition Department 112 Caption Generation Unit 113 Output section 20 Detector 21 Control section 22 Communications Department 23 Camera 24 sensors 25 Care Call Department 26. Mike 30 Control device 31 Control Unit 311 Judgment section 32 Storage section 33 Communications Department 34 Display section 35 Operation input section
Claims
1. an acquisition unit that acquires video data obtained by capturing an image of the care recipient in the room using the image capturing unit; a determination unit that determines whether a specific situation of the care recipient occurs in the room; a caption generation unit that generates a caption including a cause for the specific situation by inputting the video data acquired by the acquisition unit into a trained model that has been trained based on past data related to care; A care system comprising:
2. When the determination unit determines that the specific situation has occurred, The nursing care system according to claim 1 , wherein video data of a predetermined time range including the time when the specific situation occurred is input into the trained model.
3. The nursing care system of claim 1, wherein the trained model is trained using training data that pairs video data containing a specific situation with captions that explain the specific situation for the video data.
4. The care system according to claim 3 , wherein the captions of the correct labels used in the training data are created by a care professional.
5. The nursing care system according to claim 1 , wherein the determination unit determines that the specific situation has occurred based on the video data.
6. The care system according to claim 5 , wherein the specific situation includes at least one of the care recipient getting out of bed, getting up, and falling.
7. The care system according to claim 5 , wherein the determination unit determines that a specific situation has occurred when detecting that the care recipient has been standing still for a predetermined period of time or more.
8. the imaging unit is a stereo camera disposed above the room, The care system of claim 1, wherein the determination unit extracts height information from the parallax of two video data, and determines that the specific situation has occurred when the time rate of change in the height of the care recipient is greater than or equal to a predetermined value.
9. The nursing care system according to claim 1, wherein the determination unit determines that the specific situation has occurred when a volume greater than a predetermined value is detected from audio data obtained by a sound collection unit that collects sound within the room.
10. Further, an output unit is provided for generating a care list that lists information on a specific situation including the caption, Each time the specific situation occurs, video data is input to the trained model of the caption generation unit to generate the caption; and The care system according to claim 1 , wherein the output unit adds information about the newly occurring specific situation to the care list.
11. The nursing care system according to claim 10 , wherein information on a specific situation in the nursing care list is linked to video data of a predetermined time range including a time point at which the specific situation occurred.
12. a step (a) of acquiring video data obtained by capturing an image of the care recipient in a room using an image capturing unit; (b) determining whether a specific situation of the care recipient occurs in the room; A step (c) of generating a caption including a cause for the specific situation by inputting the video data acquired in the step (a) into a trained model trained based on past data related to caregiving; Monitoring methods, including:
13. a step (a) of acquiring video data obtained by capturing an image of the care recipient in a room using an image capturing unit; (b) determining whether a specific situation of the care recipient occurs in the room; A step (c) of generating a caption including a cause for the specific situation by inputting the video data acquired in the step (a) into a trained model trained based on past data related to caregiving; A control program for causing a computer to execute a process including the above.
Citation Information
Patent Citations
Photo book creation system and server device
JP2019160186A
System and method for generating nursing records, system and method for generating medical records, and program
JP2021162927A