Detection device and program
The detection device improves the accuracy of identifying behavior-related events by analyzing the positional relationship between a boundary and specific body parts using a machine-learned model, effectively distinguishing between normal and risky bed positions.
Patent Information
- Application Number
- JP2025124598
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2041-03-12
AI Technical Summary
Existing technologies struggle to accurately distinguish between normal lying positions and boundary positions near the edge of a bed, which can indicate a high risk of falling, leading to potential inaccuracies in detecting behavior-related events.
A detection device that utilizes a control unit to analyze the positional relationship between a predetermined boundary and specific parts of a person using a machine-learned learning model, based on captured images with depth information, to differentiate between sitting-on-the-edge and sitting-up states.
Enhances the accuracy of detecting behavior-related events, particularly distinguishing between normal lying positions and potentially dangerous boundary positions, thereby improving safety monitoring.
Smart Images

Figure 2025142263000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a detection device that detects behavior-related events related to a person and to techniques related thereto. [Background technology]
[0002] There is a technology that detects behavior-related events (such as "lying down," "standing up," "getting up," and "boundary position") related to a target person (such as a person being monitored).
[0003] For example, Patent Document 1 describes a technology for monitoring the behavior of a person (such as an inpatient) in a hospital or a nursing home, etc., and predicting the behavior of the person based on the monitoring results. In this technology, the area in which the monitored person is present and the status of the monitored person (such as "lying down," "standing up," "sitting up," or "borderline position") are detected, and events of the monitored person are determined based on the area and changes in the status. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] International Publication No. 2019 / 030880 Summary of the Invention [Problem to be solved by the invention]
[0005] However, the technology of Patent Document 1 may not be able to accurately detect behavior-related events relating to a person (for example, the status of a person being monitored).
[0006] For example, when a person is lying in bed, it may not be possible to accurately distinguish between a state in which the person is lying normally near the center of the bed (normal lying position) and a state in which the person is lying near the edge of the bed (boundary position). The boundary position is a dangerous state in which there is a high possibility of the person falling from the bed, and it is desirable to accurately detect such a state.
[0007] Therefore, an object of the present invention is to provide a technique that can more accurately detect behavior-related events related to a person. [Means for solving the problem]
[0008] In order to solve the above problem, the detection device of the present invention includes a control unit that detects behavior-related events of a target person, and the control unit detects the behavior-related events of the target person based on the positional relationship between a predetermined boundary and multiple specific parts of the target person.
[0009] The control unit may detect the behavior-related event related to the target person based on information indicating the positional relationship between the specified boundary and the multiple specific parts of the target person, using a learning model that has been machine-learned using multiple training data that takes information indicating the positional relationship between the specified boundary and the multiple specific parts of the target person as input and outputs the behavior-related event related to the target person.
[0010] The predetermined boundary may be a boundary of a bed.
[0011] The control unit may detect the behavior-related event by distinguishing between a sitting-on-the-edge state in which the target person is sitting on the edge of the bed and a sitting-up state in which the target person's upper body is raised on the bed.
[0012] The control unit may determine whether the target person is present in a human candidate area in the captured image based on the captured image and depth information of each pixel in the captured image.
[0013] The control unit may automatically set, as the predetermined boundary, the long side of the bed area set using an adjustment image obtained by photographing the target space, which is closest to the center of the adjustment image.
[0014] The control unit may detect the behavior-related event of the target person based also on a positional relationship between a bed area and the plurality of specific parts of the target person.
[0015] The control unit may set the bed area based on a predetermined point in an adjustment image obtained by capturing the target space.
[0016] The predetermined point may be designated in response to a position designation operation by a user within the adjustment image.
[0017] In order to solve the above problem, the detection method of the present invention is characterized by comprising: a) a step of acquiring a positional relationship between a predetermined boundary and multiple specific parts of a target person; and b) a step of detecting a behavior-related event related to the target person based on the positional relationship.
[0018] In order to solve the above problem, the learning model production method of the present invention is a learning model production method for producing a learning model that estimates behavior-related events related to a person, and is characterized by comprising: a) a step of machine learning the learning model using a plurality of training data in which information indicating the positional relationship between a predetermined boundary and a plurality of specific parts of the person is input and behavior-related events related to the person are output. [Effects of the Invention]
[0019] According to the present invention, it is possible to more accurately detect behavior-related events relating to a person. [Brief explanation of the drawings]
[0020] [Figure 1] FIG. 1 is a schematic diagram illustrating a detection system. [Figure 2] FIG. 2 is a functional block diagram showing a schematic configuration of a detection device. [Figure 3] FIG. 10 is a conceptual diagram showing how information on the three-dimensional position of each specific part of a photographed person is acquired (calculated) based on a photographed image and depth information. [Figure 4] This is a vertical cross-sectional view near the bed. [Figure 5] This is a view of the area around the bed from above. [Figure 6] FIG. 1 is a conceptual diagram illustrating processing at the learning stage in machine learning. [Figure 7] This is a conceptual diagram showing the processing at the inference stage using a trained model. [Figure 8] 10 is a flowchart showing the process of the learning stage. [Figure 9] 10 is a flowchart showing the processing at the inference stage. [Figure 10] FIG. 10 is a diagram showing "getting up from bed." [Figure 11] FIG. 10 is a diagram showing the "edge sitting position." [Figure 12] This is a diagram showing "one arm reaching out from the bed." [Figure 13] This is a diagram showing "lower body sliding off the bed." [Figure 14] This is a diagram showing "whole body sliding off the bed." [Figure 15] FIG. 10 is a diagram showing the positional relationship between a specific portion and a boundary surface, etc.; [Figure 16] 10 is a flowchart showing a three-dimensional object determination process. [Figure 17] FIG. 10 is a conceptual diagram illustrating a three-dimensional object determination process. [Figure 18] 10 is a flowchart showing a process for setting a bed area, etc. [Figure 19] 10A to 10C are diagrams showing adjustment images and the like obtained by photographing a target space including a bed. [Figure 20] 10A and 10B are diagrams showing that a bed upper surface area is detected in response to a bed position designation operation, etc. FIG. [Figure 21] FIG. 10 is a diagram showing how a bed boundary is automatically set based on the bed upper surface area. [Figure 22] FIG. 10 is a diagram showing a situation in which a target person has slipped off to the side (nearby) of a bed. [Figure 23] FIG. 10 is a diagram showing a situation in which a target person is lying down at a location away from a bed. [Figure 24] FIG. 1 is a diagram showing a situation in which a target person is lying in bed in a normal state. DETAILED DESCRIPTION OF THE INVENTION
[0021] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.
[0022] 1. First Embodiment <1-1. System Overview> Fig. 1 is a schematic diagram showing a detection system 1. As shown in Fig. 1, the detection system 1 includes a plurality of detection devices 10 and a plurality of terminal devices 70, 80. The terminal device 70 is also referred to as a management device 70, and the terminal device 80 is also referred to as a mobile terminal device 80. Note that Fig. 1 illustrates some of the detection devices 10 and the terminal devices 70, 80.
[0023] Here, the following mainly illustrates an example in which the detection system 1 is used in a nursing facility. However, the present invention is not limited to this, and the detection system 1 may also be used in a nursing facility (such as a hospital) or a general home.
[0024] 1, the detection devices 10 and the terminal devices 70, 80 are connected to each other via a network 108. The network 108 is configured by a LAN (Local Area Network), the Internet, etc. The connection to the network 108 may be a wired connection or a wireless connection. For example, the management device 70 is connected to the network 108 by wire, and the detection devices 10 and the mobile terminal devices 80 are connected to the network 108 wirelessly. Alternatively, all of the devices 10, 70, 80 may be connected to the network 108 wirelessly.
[0025] Each detection device 10 is arranged in each room 90 (e.g., each care recipient's private room) of each person to be observed (here, a care recipient). Each detection device 10 (and detection system 1) is a device that detects various "behavior-related events" related to the person to be observed (such as a care recipient) based on captured images of the person to be observed. "Behavior-related events" include the person's behavior (movement) itself and / or states related to the person's behavior. "Behavior-related events" are events that should be detected related to the person to be observed (events that are the target of detection processing), and are also referred to as detection target events. Note that the detection device 10 and detection system 1 are also referred to as monitoring devices and monitoring systems, as they monitor the behavior of people.
[0026] Examples of a person's "behavioral events" include "normal lying position" (event E0), "sitting up (upper body) in bed" (event E1), "boundary position" (event E2), and "slipping (out of bed)" (event E4).
[0027] Event E0 "normal lying position (normal lying position)" is an event that indicates a state in which a person is lying normally on the bed (near the center of the bed) (see FIG. 5). Note that FIG. 5 is a view of the area around the bed seen from above (vertically above). Also, in FIG. 5 and other figures, people and the like are shown in a simplified manner for convenience of illustration. For example, a person is shown as multiple connected ellipses.
[0028] Event E1 "sitting up (upper body) on bed" is a state in which the upper body is sitting up on the bed (also referred to as an upper body sitting up state) (see FIG. 10).
[0029] Event E2 "boundary position" can be further classified into event E2a "edge sitting position" and event E2b "edge lying position", etc. Event E2a "edge sitting position" is a sitting position near the edge of the bed (a posture state in which the person sits on the edge of the bed) (see Figure 11). On the other hand, event E2b "edge lying position" is a lying position near the edge of the bed. "Edge lying position" includes, for example, a state in which the person is lying very close to the edge of the bed and / or a state in which one arm is stretched out beyond the edge of the bed ("one hand outstretched from the bed") (see Figure 12).
[0030] Event E4 "sliding down" is a state in which a person slides down from the bed. Event E4 can be further classified depending on whether the part of the body that slides down from the bed to the bedside is the upper body, the lower body, or the whole body. Specifically, event E4 can be further classified into event E4a "upper body sliding down from the bed," event E4b "lower body sliding down from the bed" (see FIG. 13), event E4c "whole body sliding down from the bed," etc.
[0031] 5 and 10 to 13 are diagrams showing examples of each behavior-related event (events E0, E1, E2a, E2b, E4b).
[0032] <1-2. Overview of the detection device 10> As shown in FIG. 1, the detection device 10 includes a camera unit 20 and a processing unit 30. The camera unit 20 is installed on the ceiling of a room 90 or the like and is capable of capturing images of the interior of the room 90. The processing unit 30 is capable of acquiring three-dimensional positions (particularly time-series information) of multiple specific body parts (eyes, ears, nose, chest, waist, shoulders, elbows, wrists, knees, ankles, etc.) of the person being observed based on images (particularly moving images) captured by the camera unit 20. Note that, although the camera unit 20 and the processing unit 30 are provided separately here, this is not limiting and the camera unit 20 and the processing unit 30 may be provided as an integrated unit.
[0033] The camera unit 20 is a three-dimensional camera. The camera unit 20 acquires a captured image with depth information. Specifically, the camera unit 20 captures a captured image (such as an infrared image) 110 (see FIG. 3) of a subject object (such as a person, a wall, a floor 91, or a bed 92), and acquires depth information (depth distance information) 120 for each pixel in the captured image 110. The depth information 120 for each pixel in the captured image is information on the distance (distance from the camera unit 20) to the subject object (such as a person, a wall, a floor, or a bed) corresponding to each pixel in the captured image 110, and is information on the distance in a direction perpendicular to the captured image 110. In other words, the depth information is distance information (depth distance information) in the normal direction of the plane of the captured image.
[0034] For example, camera unit 20 includes sensor unit 23 (FIG. 2), which includes an infrared camera (e.g., an infrared image sensor) and an infrared projector. The infrared camera captures a captured image 110 (e.g., an infrared image) of a subject object. The infrared projector projects a predetermined infrared dot pattern onto the subject object, and the infrared camera captures an infrared image of the dot pattern projected onto the subject object. The distance (depth distance information) 120 to the subject object is calculated (acquired) by utilizing the fact that the geometric shape of the captured dot pattern changes depending on the distance to the subject object. The calculation process of the depth distance information is executed by controller 21 or the like incorporated in camera unit 20. Controller 21 has the same hardware configuration as controller 31 (described later) or the like. Note that sensor unit 23 may be provided with an RGB image sensor or the like that captures visible light images (e.g., color images), and a color image may be captured.
[0035] In this way, the camera unit 20 acquires a captured image (captured image information) 110 (see FIG. 3) of the subject person and depth information 120 for each pixel in the captured image (information on the distance to the object corresponding to each pixel (depth distance information)).
[0036] FIG. 3 is a conceptual diagram showing how information on the three-dimensional position of each specific part of a subject person is acquired (calculated) based on a captured image 110 and depth information 120. Note that FIG. 3 shows the captured image 110 and the depth information 120 in a schematic manner. For example, in the actual captured image 110, the person is captured as is, whereas in the captured image 110 of FIG. 3, the person is represented (simplified) by a figure made up of connected ellipses. The same is true for other depth information 120. Furthermore, while the actual depth information 120 includes information on the distance to an object corresponding to each pixel, in the depth information 120 of FIG. 3, the distance to the object corresponding to each pixel is expressed by converting it into the density of each pixel. In detail, the magnitude of the distance is expressed by the magnitude of density (shade).
[0037] First, the processing unit 30 (particularly the controller 31) analyzes the captured image 110 and acquires skeletal information (skeletal model information) 140 (see the lower left portion of FIG. 3 ) of the captured person based on the texture information, etc., of the captured image 110. The skeletal information 140 is information that (simplified) represents the skeleton of the captured person using a plurality of specific parts (specifically, the chest, nose, shoulders, elbows, wrists, waist, knees, ankles, eyes, ears, etc.) (mainly joints) of the captured person and skeletal lines (links) connecting the plurality of specific parts. In the skeletal information 140 of FIG. 3 , each specific part is represented by a "point," and each skeletal line (connecting line) is represented by a "line segment." The skeletal information 140 based on the captured image includes information about a plurality of specific parts (here, B0 to B17) (such as planar position information of each part within the image).
[0038] In addition, the processing unit 30 acquires planar position information (two-dimensional position information within the captured image in the camera coordinate system) of each specific part (for example, each representative position) of the captured person within the captured image 110 based on the skeletal information 140 of the captured person.
[0039] Furthermore, the processing unit 30 also acquires distance information (depth position information in the camera coordinate system) from the camera unit 20 to each specific part based on depth information 120 of one or more pixels at a plane position (in the photographed image 110) corresponding to the specific part. The distance information can also be expressed as depth information in the normal direction (camera line of sight direction) of the photographed image plane.
[0040] In this way, the processing unit 30 acquires information regarding the planar position of each specific part of the subject person within the captured image (two-dimensional position information within the captured image) and distance information (depth information) to each specific part of the subject person.
[0041] Then, the processing unit 30 acquires three-dimensional position information 150 (three-dimensional position information within the living space) of each specific part of the subject person, based on information about the planar position of each specific part within the captured image and distance information (depth information) to each specific part, by coordinate transformation or the like. Specifically, the processing unit 30 converts the position information in the camera coordinate system Σ1 (planar position within the captured image and depth position in the normal direction of the captured image) into three-dimensional position information 150 in a coordinate system Σ2 fixed with respect to the living room 90 (more specifically, the bed 92). The camera coordinate system Σ1 is, for example, a three-dimensional Cartesian coordinate system based on three orthogonal axes, namely, two orthogonal axes parallel to the plane of the captured image and one axis extending in a direction perpendicular to the plane of the captured image. Furthermore, the coordinate system Σ2 after the conversion is, for example, a three-dimensional Cartesian coordinate system based on three orthogonal axes, namely, two orthogonal axes parallel to the horizontal plane and one axis extending in the vertical direction (height direction). In this detection system 1 (controller 31, etc.), the position of camera unit 20 in real space is acquired in advance, and adjustments (calibration related to camera positions, etc.) are performed in advance based on the positions of the cameras in camera unit 20. In other words, the positional relationship between the two coordinate systems Σ1 and Σ2 is adjusted in advance.
[0042] Furthermore, as will be described later, the transformed coordinate system Σ2 is a coordinate system in which, for example, the bed upper surface position is the reference position (Z=0) in the vertical direction (Z direction) and the bed boundary surface 210 (see Figure 4, etc.) is the reference position (X=0) in one horizontal direction (X direction).
[0043] In this way, the processing unit 30 acquires (calculates) information on the three-dimensional position of each specific part of the subject person. Here, the subject person in the captured image is generally considered to be the person to be observed. If there are multiple subject people within the camera's field of view, all of the subject people need only be identified as the person to be observed. However, this is not limited to this, and a desired observation target object may be identified from among the multiple subject people using information such as height (or face recognition technology, etc.).
[0044] Here, the controller 31 of the processing unit 30 generates the skeletal information 140 and the three-dimensional position information 150 of each specific part, but this is not limiting.
[0045] For example, the controller 21 of the camera unit 20 may generate the skeletal information 140. Alternatively, the controller 21 of the camera unit 20 may generate three-dimensional position information 150 of each specific part. In particular, the controller 21 may acquire information about the position of each specific part of the subject person in the captured image and distance information (depth information) to each specific part, and acquire the three-dimensional position information 150 of each specific part by coordinate transformation or the like. The controller 31 of the processing unit 30 may then acquire the three-dimensional position information 150 of each specific part from the camera unit 20. Alternatively, the controllers 21, 31 of the camera unit 20 and the processing unit 30 may cooperate to execute these various processes. In other words, the skeletal information 140 and the three-dimensional position information 150 may be generated by both or one of the controllers 21, 31.
[0046] Also, here, a 3D camera using a pattern illumination method that acquires depth information for each pixel based on the illumination of a dot pattern is exemplified, but is not limited to this. The 3D camera may be a stereoscopic camera or a TOF (Time of Flight) camera. Any type of 3D camera may be used as long as it can acquire a captured image of a person being observed and depth information (such as depth information for each pixel) that is information on the distance to a subject object in the captured image.
[0047] Furthermore, the detection device 10 is also referred to as an image processing device because it is a device that executes image processing on captured images, etc. The detection system 1 is also referred to as an image processing system, etc.
[0048] <1-3. Detailed configuration of the detection device 10> FIG. 2 is a functional block diagram showing a schematic configuration of the detection device 10. As shown in FIG.
[0049] As shown in the functional block diagram of FIG. 2, the detection device 10 includes a camera unit 20 and a processing unit 30.
[0050] As described above, the camera unit 20 includes an infrared camera, an infrared projector, and the like.
[0051] The processing unit 30 includes a controller (also referred to as a control unit) 31, a storage unit 32, a communication unit 34, and an operation unit 35.
[0052] The controller 31 is a control device that is built into the processing unit 30 and controls the detection device 10 .
[0053] The controller 31 is configured as a computer system including one or more hardware processors (for example, a central processing unit (CPU) and a graphics processing unit (GPU)). The controller 31 performs various processes by executing, in the CPU or the like, a predetermined software program (hereinafter also simply referred to as a program) stored in a storage unit (a non-volatile storage unit such as a ROM and / or a hard disk) 32. The program (more specifically, a group of program modules) may be recorded on a portable recording medium such as a USB memory, read from the recording medium, and installed in the detection device 10. Alternatively, the program may be downloaded via a communication network or the like and installed in the detection device 10.
[0054] The controller 31 acquires captured image information about a subject person (such as a person to be observed) from the camera unit 20, and also acquires distance information (depth information) to each specific part of the subject from the camera unit 20. Then, the controller 31 acquires three-dimensional position information of each specific part based on the captured image information (particularly, planar position information of each specific part of the subject in the captured image) and the distance information (depth information of the subject).
[0055] Furthermore, the controller 31 detects behavior-related events of the target person based on the positional relationship (described later) between a predetermined boundary or the like and a plurality of specific parts of the target person. An example of the predetermined boundary is the boundary of the bed 92 (more specifically, the boundary surface 210 between the space 211 above the bed and the external space 212 beside the bed (see FIGS. 1 and 4, etc.)).
[0056] Here, the controller 31 uses the learning model 410 to detect behavior-related events of the target person.
[0057] As the learning model 410, for example, a neural network model composed of multiple layers is used. Then, weighting coefficients (learning parameters) between multiple layers (input layer, (one or more) intermediate layers, and output layer) in the neural network model are adjusted using a predetermined machine learning method. The learning model 410 after being trained by machine learning is also referred to as a trained model 420. Specifically, the learning parameters of the learning model 410 (learner) are adjusted using a predetermined machine learning method, and a trained learning model 410 (trained model 420) is generated (see FIG. 6).
[0058] Therefore, first, the controller 31 executes the process of the learning stage (see FIG. 6) in machine learning. Specifically, the controller 31 performs machine learning in advance on the learning model 410 using a plurality of training data. As each training data, training data is used that receives information indicating the positional relationship between a predetermined boundary or the like and a plurality of specific parts of a person as input and that outputs behavior-related events related to the person. Then, a predetermined machine learning method is used to adjust weighting coefficients (learning parameters) between multiple layers (input layer, (one or more) intermediate layers, and output layer) in the neural network model. As a result, a trained learning model 410 (trained model 420) is generated (produced).
[0059] Thereafter, the controller 31 executes the processing of the inference stage (see FIG. 7 ) using the trained learning model 410 (420). Specifically, the controller 31 uses the learning model 410 (trained model 420) that has been machine-learned using the plurality of training data to detect behavior-related events related to the target person (also referred to as a person to be detected or a person to be determined) based on information indicating the positional relationships of the target person's multiple specific parts with respect to a predetermined boundary or the like. More specifically, the controller 31 inputs the positional relationships of the target person's multiple specific parts with respect to a predetermined boundary or the like into the trained model 420, and obtains an output (detection result of a behavior-related event of the target person) from the trained model 420. The positional relationships between the predetermined boundary or the like and the target person's multiple specific parts will be described in detail later.
[0060] Here, the detection device 10 (controller 31) performs both the learning stage process and the inference stage process in machine learning. However, this is not limited to this, and for example, the learning stage process and the inference stage process may be performed by different devices.
[0061] Furthermore, the controller 31 cooperates with the communication unit 34 and the like to transmit the detection results to the terminal devices 70, 80, and the terminal devices 70, 80 output the detection results (such as display output and / or audio output).
[0062] The storage unit 32 is configured with a storage device such as a hard disk drive (HDD) or a solid state drive (SSD). The storage unit 32 stores (memorizes) the above-mentioned programs and various data. For example, the storage unit 32 stores (memorizes) time-series data of the three-dimensional positions of multiple specific parts of the observed person, as well as various data and programs used for learning and using the learning model 410.
[0063] The communication unit 34 is capable of performing network communication via the network 108. This network communication utilizes various protocols such as TCP / IP (Transmission Control Protocol / Internet Protocol). By utilizing this network communication, the detection device 10 can exchange various data with desired destinations (e.g., terminal devices 70 and 80).
[0064] The operation unit 35 includes an operation input unit 35a that receives operation input to the detection device 10, and a display unit 35b that displays and outputs various types of information.
[0065] <1-4. Learning stage processing of learning model 410> First, the learning stage processing will be described.
[0066] 8 is a flowchart showing the process of the learning stage. The process shown in FIG. 8 is executed by the controller 31 and the like.
[0067] 8 is also a diagram showing a method for generating a trained model. In this application, generating a trained model 420 means producing (manufacturing) the trained model 420, and the "method for generating a trained model" means the "method for producing a trained model."
[0068] First, in step S11, the controller 31 acquires information 310 indicating the positional relationship between a predetermined boundary or the like and a plurality of specific parts (for example, B0 to B17) of the target person.
[0069] Specifically, first, relative position information of a plurality of specific parts B0 to B17 with respect to a predetermined boundary (for example, a boundary surface 210 virtually separating two horizontally adjacent spaces 211 and 212) is acquired as the information 310. Here, as shown in FIGS. 4 and 5, the boundary surface 210 is the boundary surface between a space 211 above the bed 92 and an external space 212 to the side of the bed 92. In other words, the boundary surface 210 is a vertical plane including one side surface 92b of the bed 92. Note that FIG. 4 is a vertical cross-sectional view of the vicinity of the bed 92 (a cross-sectional view of the bed 92, etc., cut along a vertical cross-section extending both vertically and laterally across the bed 92), and is a view from the toes to the head of a person lying on the bed 92. Also, FIG. 5 is a view (top view) of the vicinity of the bed 92 as seen from above.
[0070] The positional relationship between boundary surface 210 and each specific portion (B0 to B17) is expressed by a signed normal distance X to boundary surface 210 (signed shortest distance (differential position) X to boundary surface 210). The sign is positive (+) at a horizontal position closer to the bed (inside (right side in FIG. 4)) than boundary surface 210, and is negative (-) at a horizontal position closer to the bed outside (left side in FIG. 4) than boundary surface 210. Note that signed normal distance X to boundary surface 210 can also be expressed as the relative position in the X direction (of each specific portion) with respect to boundary surface 210.
[0071] Furthermore, in step S11, the controller 31 also acquires, as information 310, relative position information between a predetermined reference horizontal plane 230 (a horizontal plane 230 virtually separating two adjacent spaces 231, 232 adjacent in the vertical direction (height direction)) and a plurality of specific parts B0 to B17. Here, the reference horizontal plane 230 is a horizontal plane that extends two-dimensionally in the horizontal direction, and is a plane 230 that includes the upper surface 92a of the bed (the upper surface and its extended plane). The reference horizontal plane 230 is a horizontal plane that virtually separates a space above the bed (upper space) 231 from a space below the bed (lower space) 232.
[0072] The positional relationship between the reference horizontal plane 230 and each specific portion (B0 to B17) is expressed by a signed normal distance Z with respect to the reference horizontal plane 230 (signed shortest distance (differential position) Z with respect to the reference horizontal plane 230). The sign is positive (+) at a vertical (height) position above (on the upper side) the reference horizontal plane 230, and is negative (-) at a vertical (height) position below (on the lower side) the reference horizontal plane 230. The signed normal distance Z with respect to the reference horizontal plane 230 can also be expressed as the relative position in the Z direction (of each specific portion) with respect to the reference horizontal plane 230.
[0073] The above positional relationship can be easily expressed by adopting the following coordinate system Σ2.
[0074] The coordinate system Σ2 is an XYZ Cartesian coordinate system with an origin O2 (see FIGS. 4 and 5) at a point on the intersection 215 between the upper surface 92a of the bed 92 and one of the side surfaces 92b of the bed 92 (see FIG. 4) (here, the end position of the person's feet in the bed 92). In the coordinate system Σ2, the vertical direction is the Z direction, the left-right direction (lateral direction) of the bed is the X direction, and the longitudinal direction of the bed is the Y direction. The coordinate system Σ2 is a coordinate system in which the Z direction position of the bed upper surface (reference horizontal plane 230) is the reference position (Z=0) in the vertical direction (Z direction) and the bed boundary surface 210 is the reference position (X=0) in one horizontal direction (left-right direction (X direction) of the bed). In FIGS. 4 and 5, coordinate axes in the XYZ coordinate system are appropriately shown. FIG. 4 can also be expressed as a diagram showing the XZ plane, and FIG. 5 as a diagram showing the XY plane. The boundary surface 210 is also expressed as a plane where X=0, and the reference horizontal plane 230 is also expressed as a plane where Z=0.
[0075] Then, the X coordinate value Xi and Z coordinate value Zi of a plurality (here, 18) of specific parts Bi (i = 0 to 17) are acquired as information 310 indicating the positional relationship between a predetermined boundary or the like and a plurality of specific parts (for example, B0 to B17) of the target person. In detail, a total of 36 values are acquired as information 310: value X0 and value Z0 of specific part B0, value X1 and value Z1 of specific part B1, value X2 and value Z2 of specific part B2, value X3 and value Z3 of specific part B3, ..., and value X17 and value Z17 of specific part B17. Note that the Y coordinate value of each specific part Bi is not used as information 310 here.
[0076] In this way, the relative positions of the multiple specific parts Bi are calculated using the boundary surface 210 of the bed (specifically, the boundary surface 210 (plane where X=0) between the space 211 above the bed and the external space 212 beside the bed) as a reference. The relative positions of the multiple specific parts Bi are also calculated using the reference horizontal plane 230 (specifically, the horizontal plane (plane where Z=0) including the upper surface 92a of the bed) as a reference. That is, the relative positions of the multiple specific parts Bi are found as relative positions with respect to the boundary surface 210 and the reference horizontal plane 230. In other words, the relative positions of the multiple specific parts Bi are found as relative positions with respect to a reference axis (specifically, the intersection line 215 (Y-axis) between the boundary surface 210 and the reference horizontal plane 230) included in the boundary surface 210 (and the reference horizontal plane 230). The relative positions of the multiple specific parts Bi can also be expressed as two-dimensional positions obtained by projecting the three-dimensional position of each specific part Bi onto a plane (projection plane) (such as the plane of Y=0) that is perpendicular to both the boundary surface 210 and the reference horizontal plane 230. The intersection line 215 can also be expressed as a line (boundary line) that is included in both the boundary surface 210 and the reference horizontal plane 230.
[0077] Next, in step S12, the controller 31 assigns a corresponding "behavior-related event" (for example, "normal lying position" or "edge-sitting position") as a label (correct answer data) to the information 310 indicating the positional relationship between a predetermined boundary or the like and a plurality of specific parts of the person. More specifically, for example, from among the above-mentioned events E0, E1, E2a, E2b, E4a, E4b, and E4c, the corresponding "behavior-related event" is identified and assigned as a label.
[0078] The association (processing for identifying the corresponding behavior-related event) may be performed by an operator (operator) or the like. In detail, the operator determines the "behavior-related event" corresponding to the information 310, and the "behavior-related event" related to the determination result is assigned to the detection device 10 in response to an operation input by the operator. That is, the controller 31 assigns a label (correct answer data) to the information 310 in response to the operation input. In this way, training data is generated that uses the information 310 as an input and the behavior-related event as an output.
[0079] By repeatedly executing steps S11 and S12, a plurality of pieces of training data are generated. Note that the repeated steps are not shown in FIG.
[0080] Next, in step S13, these multiple pieces of training data are used to perform machine learning on the learning model 410. When information 310 indicating the positional relationship between a predetermined boundary or the like and multiple specific parts of a person is input, the learning model 410 is machine-learned to output behavior-related events related to the person. As a result, a trained model 420 that estimates behavior-related events related to the person is generated (produced) (step S14).
[0081] For example, in the event E0 "normal lying position" (see FIG. 5), the X-coordinate value Xi and the Z-coordinate value Zi of the multiple specific parts Bi (i=0 to 17) are all positive values. Note that FIG. 5 is a diagram illustrating the event E0 "normal lying position" and is a diagram (top view) of a person in a normal lying position viewed from above (vertically above).
[0082] In the event E0, the values Xi and Zi of each specific part Bi are both positive (plus). Also, since the person is located near the bed upper surface 92a (see FIG. 4) (e.g., inside the elliptical region indicated by the two-dot chain line in FIG. 4), the value Zi is relatively small (e.g., 100 mm (millimeters)). By the above-described machine learning, a trained model 420 is generated that outputs the event E0 "normal lying position" as a behavior-related event in response to input of information 310 having such characteristics.
[0083] On the other hand, suppose a situation occurs in which a person's posture changes, causing a transition from event E0 "normal lying position" to event E1 "sitting up in bed" (see FIG. 10). In this case, for example, specific part B1 ("chest") moves to position P1 in FIG. 4. As a result, for specific part B1, value Z1 changes to a relatively large value (e.g., 400 mm) (while value X1 remains approximately the same). Similarly, for the other specific parts B0, B2, B5, B14 to B17, value Zi changes to relatively large values (while value Xi remains approximately the same). By the above-described machine learning, a trained model 420 is generated that outputs event E1 "sitting up in bed" as a behavior-related event in response to input of information 310 having such characteristics.
[0084] For each of the other behavior-related events, the information 310 at the time of occurrence of the behavior-related event has its own unique properties.
[0085] For example, consider a situation in which event E1 "sitting up in bed" transitions to event E2a "sitting on the edge of bed" (see Figure 11). In this case, specific parts B10 (right ankle) and B13 (left ankle) move to the vicinity of point P3 in Figure 4. As a result, the values Xi and Zi of specific parts B10 and B13 both change to negative (-) values. In addition, the values Xi (absolute values) of specific parts B9 (right knee) and B12 (left knee) and specific parts B0 to B7, B14 to B17, etc. of the upper body each change to relatively small values. In particular, the values Xi of the specific parts of the left and right upper body (left and right shoulders, left and right elbows, etc.) become relatively small on both sides (to the same extent on both sides).
[0086] Also, consider a situation in which a person reaches outward from the edge of the bed to grab a plastic bottle or the like, resulting in a transition from event E0 (Fig. 5) to event E2b (see Fig. 12). In this case, specific parts B6 (left elbow) and B7 (left wrist) (or specific parts B3 (right elbow) and B4 (right wrist)) move to the vicinity of point P2 in Fig. 4. As a result, although the values Zi of specific parts B6 and B7 (or specific parts B3 and B4) remain roughly the same, the values Xi of those specific parts B6, B7, etc. change to negative (-) values. Furthermore, with regard to other specific parts of the upper body, such as B5 (left shoulder), the absolute values of the values Xi each change to a relatively small value (while the values Zi remain roughly the same).
[0087] Event E2b also includes a state in which a person moves to the edge of the bed and lies down (boundary position). In this case, some parts of the body, such as the elbows, wrists, knees, and ankles, move to the edge of the bed. As a result, the values Xi (absolute values) of at least some of the specific parts B0 to B7, B14 to B17, etc. (for example, multiple parts on the right side and / or multiple parts on the left side) change to values smaller than a certain level.
[0088] As described above, the coordinate values change in response to changes in each event. In other words, each event has its own unique characteristics with respect to the combination of the coordinate values Xi, Zi of multiple specific body parts. These unique characteristics are then learned by the learning model 410. Specifically, the learning model 410 appropriately learns the correspondence between the coordinate values Xi, Zi of multiple specific body parts and behavior-related events. That is, according to the above-described machine learning, a trained model 420 is generated that outputs a corresponding behavior-related event in response to input of information 310 having characteristics indicated by each coordinate value.
[0089] <1-5. Inference stage processing using trained model 420> Next, the processing at the inference stage (stage of estimating behavior-related events) using the trained model 420 will be described.
[0090] Fig. 9 is a flowchart showing the processing of the inference stage. The processing shown in Fig. 9 is executed by the controller 31, etc. In the processing of the inference stage, the detection device 10 (controller 31) uses the trained model 420 to detect a behavior-related event of the target person based on information 310 indicating the positional relationship between a predetermined boundary (boundary surface 210) and a plurality of specific parts B0 to B17 of the target person.
[0091] Therefore, first, in step S31, a process similar to that of step S11 (FIG. 8) is executed for the person to be determined. As a result, information 310 indicating the positional relationship between the multiple specific parts Bi of the person to be determined and a predetermined boundary or the like is acquired. For example, the information 310 includes the positional relationship between the multiple specific parts Bi of the person to be determined and the bed boundary surface 210, as well as the positional relationship between the multiple specific parts Bi of the person to be determined and the reference horizontal plane 230. More specifically, the X-coordinate value Xi and Z-coordinate value Zi of the specific parts Bi (i=0 to 17) are acquired as the information 310.
[0092] In the next step S32, the controller 31 inputs the information 310 into the learning model 410 (trained model 420) and acquires the output (behavior-related event) from the learning model 410 as an estimation result (inference result). Specifically, any one of the above-mentioned events E0, E1, E2a, E2b, E4a, E4b, and E4c is output from the learning model 410 and acquired as the estimation result.
[0093] In step S33, the estimation result is output. Specifically, the controller 31 displays the name of the estimated behavior-related event (such as "normal lying position" or "edge sitting position") on the display unit 35b. Furthermore, the controller 31 transmits the estimation result to the terminal devices 70 and 80 by communication or the like, and causes the terminal devices 70 and 80 to display the result on their respective displays. In particular, when an event other than event E0 occurs, it is preferable that the occurrence of the event is notified to the terminal user on the terminal devices 70 and 80 with a warning display or a warning sound.
[0094] <1-6. Effects of the embodiment> According to the above process, the learning model 410 is machine-learned using training data in which information 310 indicating the positional relationship between a predetermined boundary or the like and a plurality of specific parts Bi of a person is input and behavior-related events related to the person are output. Therefore, the learning model 410 that properly reflects the positional relationship between the predetermined boundary or the like and a plurality of specific parts Bi of a person is generated.
[0095] In particular, the three-dimensional position information (such as three-dimensional position information in the coordinate system Σ1) of the multiple specific parts Bi is not used as an input as it is, but rather a value (Xi) converted into information (relative position information) indicating the relative positional relationship with the boundary surface 210 is used as an input to the learning model 410. This makes it possible to appropriately reflect in the learned model 420 the positional relationship between each specific part Bi and the boundary surface 210 (the distance between each specific part and the boundary surface 210, and whether each specific part Bi is located inside or outside the boundary surface 210).
[0096] Similarly, the value (Zi) converted into information (relative position information) indicating the relative positional relationship with the reference horizontal plane 230 (bed upper surface 92a, etc.) is used as an input to the learning model 410. This makes it possible to appropriately reflect in the learned model 420 the positional relationship between each specific part Bi and the reference horizontal plane 230 (the distance between each specific part Bi and the reference horizontal plane 230, and whether each specific part Bi is located above or below the reference horizontal plane 230).
[0097] In particular, by not using the Y-direction information (Yi) among the three-dimensional information and using only the information on the other two dimensions (Xi, Zi), it is possible to appropriately reduce the amount of information and achieve efficient learning.
[0098] Such machine learning allows appropriate adjustment of weighting parameters and the like in the learning model 410. As a result, a trained model 420 is generated that has been trained to output appropriate behavior-related events even in response to new (unknown) input indicating the positional relationship between a predetermined boundary or the like and multiple specific parts Bi.
[0099] Furthermore, in the inference stage of the above embodiment, the detection device 10 (controller 31) detects behavior-related events of the target person based on the positional relationship between a predetermined boundary or the like and multiple specific parts Bi of the target person. Specifically, using the trained model 420 described above, behavior-related events related to the target person are detected based on information 310 indicating the positional relationship of the multiple specific parts Bi with respect to the boundary surface 210. In other words, behavior-related events related to the target person are detected using the trained model 420 trained to output appropriate estimation results. Therefore, behavior-related events related to the target person can be detected more accurately.
[0100] In particular, it is possible to obtain estimation results (behavior-related information) that appropriately reflect the positional relationship between each specific part and the boundary surface 210 (the distance between each specific part and the boundary surface 210, and whether each specific part is located inside or outside the boundary surface), etc.
[0101] Therefore, for example, it is possible to accurately distinguish between a state in which the target person is lying normally near the center of the bed (normal lying position) and a state in which the target person is lying near the edge of the bed (boundary position), and detect both states. Alternatively, it is possible to accurately distinguish between a state in which the target person is sitting on the edge of the bed ("edge sitting position") and a state in which the target person's upper body is raised on the bed ("sitting up on the bed" state), and detect both states (behavior-related events). In other words, it is possible to accurately detect behavior-related events in which the positional relationship with the boundary surface 210 is an important discrimination factor.
[0102] Furthermore, in the inference stage of the above embodiment, behavior-related events of the target person are detected based particularly on the positional relationship between the reference horizontal plane 230 and the target person's multiple specific body parts (B0 to B17). Specifically, using the trained model 420 described above, behavior-related events of the target person are detected based on information 310 indicating the positional relationship of the multiple specific body parts Bi with respect to the reference horizontal plane 230. This makes it possible to more accurately detect behavior-related events of the target person. In particular, it is possible to accurately detect behavior-related events in which not only the positional relationship with the boundary plane 210 but also the positional relationship with the reference horizontal plane 230 are important discriminators. For example, it is possible to accurately distinguish and detect event E2a "sitting on the edge of bed" (see FIG. 11) and event E4b "slipping of the lower body off the bed" (see FIG. 13).
[0103] Note that, in each stage (especially the estimation stage) of the above embodiment, it is not necessary to input relative position information, etc., for all of the multiple specific parts to the learning model 410. Specifically, the relative positional relationship with the boundary surface 210, etc., may be acquired for only the specific parts Bi, among the multiple specific parts Bi, that have a large weighting (weighting in the learning model 410) corresponding to a specific behavior-related event. Even in this case, the specific behavior-related event can be suitably determined. For example, when determining "sitting up in bed" (see FIG. 10), the relative positional relationship of the specific parts Bi of the upper body with the boundary surface 210, etc., is an important factor (a factor with a large weighting). Therefore, when positional information for several major specific parts of the upper body is obtained, it is possible to suitably determine "sitting up in bed" (see FIG. 10) even when positional information for specific parts of the lower body (e.g., "knees" and "ankles") is not available.
[0104] Furthermore, when recognizing (detecting) two similar behavior-related events in a distinguishable manner, it is possible to obtain favorable recognition results by acquiring position information of some specific parts Bi among multiple specific parts that are heavily weighted to correspond to the two behavior-related events. For example, to distinguish between "sitting up on the bed" (see FIG. 10) and "sitting on the edge of the bed" (see FIG. 11), the relative positional relationship of the specific parts Bi of the upper body with the boundary surface 210, etc., is an important factor (a factor with a heavy weighting). More specifically, in "sitting on the edge of the bed," the relative positions of specific parts such as "shoulders" and "elbows" on both the left and right sides and the boundary surface 210 are close. In "sitting up on the bed," the distance Xi between one of the left and right (e.g., right) specific parts such as "shoulders" and "elbows" (right shoulder, right elbow, etc.) and the boundary surface 210 is small, while the distance Xi between the other (e.g., left) specific part such as "shoulder" and "elbow" (left shoulder, left elbow, etc.) and the boundary surface 210 is large. By (substantially) learning such a difference between the left and right sides in the relative position between a specific part of the upper body and boundary surface 210, the two types of behavior-related events can be distinguished from each other and appropriately detected.
[0105] <1-7. Modifications, etc.> In the above embodiment, the coordinate values Xi, Zi of each specific portion Bi in the XYZ orthogonal coordinate system are used, but the present invention is not limited to this.
[0106] For example, coordinate values Ri, θi of each specific portion Bi in a cylindrical coordinate system (R, θ, Y) with the Y axis of the XYZ Cartesian coordinate system as the reference axis (rotation center axis) may be used (see FIG. 15). That is, coordinate values θi, Ri may be used instead of coordinate values Xi, Zi. Here, coordinate value θi is the rotation angle θi around origin O2 in the XZ plane of FIG. 4. Specifically, rotation angle θi may be determined such that a clockwise rotation angle is a positive (+) rotation angle (a counterclockwise rotation angle is a negative (-) rotation angle) with the vertical upward direction (+Z direction) as the reference (θi=0). Furthermore, coordinate value Ri is the distance (difference distance) from the reference axis (intersection line 215).
[0107] In this way, the differential position information Ri of each specific part relative to the intersection line 215 (Y axis) between the boundary surface 210 and the reference horizontal plane 230 and the angle information θi of each specific part relative to the intersection line may be used as relative position information of each specific part relative to the boundary surface 210 and the reference horizontal plane 230.
[0108] Furthermore, in the above embodiment, both the value Xi and the value Zi of each specific portion Bi are used as inputs to the learning model 410, but this is not limiting. For example, only the value Xi may be used as an input to the learning model 410. This makes it possible to obtain a learning result that reflects at least the relative positional relationship with the boundary surface 210. Similarly, of the values Ri and θi in the cylindrical coordinate system, only the value θi may be used as an input to the learning model 410.
[0109] However, in order to obtain learning results (and inference results) that also reflect the relative positional relationship with the reference horizontal plane 230, it is preferable that both the values Xi and Zi of each specific part Bi (or both the values Ri and θi, etc.) be used as inputs to the learning model 410.
[0110] In the above embodiment, the events E0, E1, E2a, E2b, E4a, E4b, and E4c are mainly exemplified as behavior-related events, but the present invention is not limited to these. For example, other events (such as "standing out of bed") may also be included as behavior-related events. Furthermore, the events may be more finely categorized, or conversely, may be more broadly categorized events. Alternatively, the behavior-related events may be any combination of all or part of these events.
[0111] 2. Second Embodiment The second embodiment is a modification of the first embodiment, and the following description will focus on the differences from the first embodiment.
[0112] In the second embodiment, a technique that can accurately extract a person on a bed or a person near a wall in the first embodiment and the like will be described.
[0113] In the first embodiment and the like, the human body (human body texture) is recognized based on texture information and the like of the photographed image 110, and skeletal information (skeletal model information) 140 of the person is acquired.
[0114] However, when recognizing a human body based only on the texture information (two-dimensional information) of the captured image 110, a texture that is actually flat and resembles the texture of a human body may be mistakenly recognized as the texture of a real (three-dimensional) human body.
[0115] For example, the above-mentioned misrecognition can occur when detecting the texture of a human body in a poster (poster stuck on a wall, etc.). Alternatively, depending on the accuracy of human body texture recognition, a flat texture on the surface of a bed (or near a wall) may be mistakenly recognized as a real human body if it happens to resemble the texture of a real (three-dimensional) human body. The above-mentioned misrecognition can also occur when the shadow of a person in a photographed image (such as an image taken with an RGB camera) is detected as the texture of a human body.
[0116] Therefore, in this second embodiment, a technique that can avoid misidentifying a planar texture as a texture of a person will be described. In detail, a technique for avoiding misidentification that utilizes the three-dimensionality of a person will be described.
[0117] Specifically, the controller 31 determines whether or not a person exists in a candidate person area in the photographed image 110, based on the photographed image 110 and the depth information 120 of each pixel in the photographed image 110. More specifically, as shown in Figures 16 and 17, the controller 31 obtains a reference plane 620 in the candidate person area in the photographed image 110, based on the depth information 120 of each pixel in the photographed image 110. Then, the controller 31 determines whether or not a target person exists in the candidate person area, based on the protrusion amount from the reference plane 620 of the point cloud that protrudes from the reference plane 620.
[0118] Such processing is executed when confirming the presence of a person in the above-mentioned step S11 (FIG. 8) and / or step S31 (FIG. 9) (or immediately before each step), etc. FIG. 16 is a flowchart showing such processing, and FIG. 17 is a conceptual diagram showing the processing. The left column of FIG. 17 shows a cross section of real space when a three-dimensional object (real) human body 610 is present, and the right column of FIG. 17 shows a cross section of real space when a three-dimensional object (real human body 610) is not present.
[0119] More specifically, as shown in FIG. 16, first, in step S51, the controller 31 extracts a judgment target area (a candidate person area and its surrounding area) including a target object (object to be judged) to be judged as to whether it is a person or not, based on the texture information of the captured image 110.
[0120] In detail, first, the controller 31 performs image processing (feature analysis processing) on the captured image 110 to extract (detect) an area (also referred to as a candidate person area) that is determined to be highly likely to be a person (having a certain degree or more of the possibility) based on the texture, etc., of the captured image 110. The candidate person area is, for example, an area from which skeletal information of a person is extracted (an area having a human shape, etc.). Such extraction processing (detection processing) may utilize image recognition processing using a neural network, etc. Then, the controller 31 sets a circumscribing rectangle that surrounds the candidate person area within the captured image 110, and sets the area inside the circumscribing rectangle (the candidate person area and its surrounding area) as a determination target area (see also the top row of FIG. 17). Note that in FIG. 17, the determination target area, which originally has a two-dimensional extent, is shown one-dimensionally (as a one-dimensional range in one cross section).
[0121] Next, in step S52, the controller 31 obtains a reference plane 620 (reference plane in real space) in the determination target area based on the depth information 120 of the point group (pixel group) in the determination target area (see also the second row from the top in FIG. 17). For example, a bed surface, floor surface, wall surface, slope, or the like is extracted as the reference plane 620.
[0122] The reference plane may be determined, for example, using a RANSAC (RANdom Sample Consensus) method. Specifically, first, any three points are extracted from a point cloud (a set of points having three-dimensional positions determined for each pixel in the region) within the determination target region (preferably, a region surrounding a candidate human region therein), and a plane (temporary plane) passing through the three points is provisionally set. Then, if the number of points in the point cloud that exist within a predetermined allowable range (for example, several mm to several tens of mm (millimeters)) from the reference plane is the largest so far, the plane formed by the three points is updated as a solution plane (optimal plane). In other words, if the number of points existing near the temporary plane is greater than the number of points existing near the optimal plane at that time, the temporary plane is determined as a new optimal plane. Thereafter, the same process (such as setting a temporary plane and comparing the temporary plane with the optimal plane) is repeated a number of times (a predetermined number of times) for the new any three points. As a result, the finally obtained solution plane is determined as the reference plane 620.
[0123] In step S53, a group of points within the determination target area, excluding the group of points that constitute the reference plane 620 (points that exist near the reference plane), is determined as a group of points that protrude from the reference plane 620. In the third row from the top of Fig. 17, the group of points that constitute the reference plane 620 is represented by white circles, and the remaining group of points represented by black circles is a group of points that protrude from the reference plane 620 (also referred to as a protruding point group). The group of points that protrude from the reference plane 620 is a group of points that do not constitute the reference plane 620, and is therefore also referred to as a non-plane point group.
[0124] Then, in step S54, the amount of protrusion (the amount of protrusion in the normal direction of the reference plane) Hi from the reference plane 620 of each point Pi in the point group (non-plane point group) protruding from the reference plane 620 is calculated. Then, the average value Hv of the protrusion amounts Hi of the multiple points Pi is calculated.
[0125] If the average value Hv is greater than a threshold value TH1 (for example, 80 mm), it is determined that a three-dimensional object (and therefore a human body) is present. On the other hand, if the average value Hv is smaller than the threshold value TH1, it is determined that a three-dimensional object (and therefore a human body) is not present. In this way, the controller 31 determines whether or not a person (three-dimensional object) is present within the determination area based on the protrusion amount Hi.
[0126] As described above, the controller 31 determines the reference plane 620 in the determination target area (the person candidate area and its surrounding area) based on the depth information 120 of the point group (pixel group) in the determination target area. The controller 31 then determines the presence or absence of a person in the determination target area based on the protrusion amount H (Hi, etc.) (from the reference plane) of the point group protruding from the reference plane 620. This makes it possible to accurately determine (confirm) whether the texture is three-dimensional, even if a planar texture in the determination target area is mistakenly detected as a human body, based on the protrusion amount Hi from the reference plane. This makes it possible to avoid misidentifying the planar texture as the texture of a real person (three-dimensional object). In this way, it is possible to accurately determine the presence or absence of a target person as a three-dimensional object.
[0127] In particular, the reference plane 620 is once determined, and a relative displacement with respect to the reference plane 620 (such as a relative distance (difference) from the reference plane 620) is used to appropriately determine the presence or absence of a protrusion (three-dimensional object) based on the reference plane 620. In other words, the presence or absence of a three-dimensional object on the reference plane 620 can be appropriately determined without necessarily having to determine the absolute value of the height of the reference plane 620.
[0128] The concept of the second embodiment has been described as a modified example of the first embodiment, but is not limited to this. For example, as in the first embodiment, a situation is assumed in which the position of camera unit 20 in real space is acquired in advance, adjustments based on the position of camera unit 20 (calibration related to the camera position, etc.) are performed, and then the positions of each specific part in coordinate system Σ2, etc. are measured. However, the concept of the second embodiment can be applied not only to this situation but also to other situations.
[0129] For example, this concept can also be applied to the case where it is determined simply whether or not a three-dimensional human body actually exists in a human body candidate region on a floor or wall surface. This concept itself does not necessarily require calculating the absolute value of the height of the reference plane 620. Therefore, in such a case, it is possible to determine the presence or absence of a three-dimensional object using the relative displacement of each point cloud with respect to the reference plane 620, without making adjustments based on the position of the camera unit 20 in real space (such as setting camera parameters (camera installation position information)).
[0130] 3. Third Embodiment In the above embodiments, the coordinate system Σ2 is exemplified as an XYZ Cartesian coordinate system having the origin O2 (see FIGS. 4 and 5) as a point on the intersection line 215 between the upper surface 92a of the bed 92 and the side surface 92b of one of the beds 92 (see FIG. 4). This intersection line 215 may be set, for example, as shown below. In the third embodiment, details of setting the coordinate system Σ2 (particularly the intersection line 215) will be described. Here, the setting process of the intersection line 215 is executed in conjunction with the setting process of the bed area 323 (see FIG. 21, etc.).
[0131] FIG. 18 is a flowchart showing the setting process of the bed area 323 etc. (the process of the controller 31).
[0132] As shown in FIG. 18, in step S71, a captured image 110C and depth information 120C relating to a target space (target space of the detection process) including the bed 92 are acquired (see also FIG. 19).
[0133] The left side of FIG. 19 shows an adjustment image (a photographed image 110C for calibration (pre-adjustment)) of a target space including a bed 92, and the right side of FIG. 19 shows visualized depth information 120C of each pixel in the photographed image 110C. Note that the depth information 120C is expressed by converting the distance to an object corresponding to each pixel into the density of each pixel. Here, an image photographed from the ceiling downward (an image photographed by the camera unit 10 (infrared image)) is shown as an example of the photographed image 110C. Note that the photographed image 110C is not limited to this, and may be one photographed of a target space such as a living room (particularly a target space including the space where the bed is placed) from various angles.
[0134] In steps S72 to S74, the bed area 323 etc. is set using an adjustment image (photographed image) 110C obtained by photographing the target space. In detail, the bed area 323 etc. is set based on a predetermined point 315 (see FIG. 20) in the photographed image 110C.
[0135] Specifically, first, in step S72, the controller 31 accepts a user's operation to specify a predetermined point 315 (a bed position specification operation). Specifically, on a display screen (display screen displayed on the display unit 35b or the like) displaying the photographed image 110C, the user moves a cursor 311 for specifying a bed position to a position near the center of the bed 92 in the photographed image 110C by using a mouse and double-clicks the cursor. In short, the user specifies the bed position by pointing (pointing) at it using the cursor 311. In response to this operation, the controller 31 acquires the cursor position (specifically, the center position of the circular cursor 311) within the display screen (within the photographed image 110C) as the predetermined point 315. In this way, the predetermined point 315 is specified in response to a position specification operation within the photographed image 110C by the user. The predetermined point 315 is not limited to a position near the center of the bed, and may be, for example, another position (a position other than near the center) on the upper surface of the bed.
[0136] The controller 31 then acquires the three-dimensional position (X, Y, Z) of the object surface position (position on the upper surface of the bed) corresponding to the two-dimensional position of the predetermined point 315. Specifically, the three-dimensional position (X, Y, Z) of a point on the upper surface of the bed 92 in the coordinate system Σ3 is acquired. In other words, the position information in the camera coordinate system Σ1 regarding the predetermined point 315 (the plane position in the captured image and the depth position in the normal direction of the captured image) is converted into coordinate values (X, Y, Z) in the coordinate system Σ3 and acquired. The coordinate system Σ3 is, for example, a three-dimensional Cartesian coordinate system based on three orthogonal axes, namely, two orthogonal axes parallel to the horizontal plane and one axis extending in the vertical direction (height direction). The origin of the coordinate system Σ3 may be set at an appropriate position (such as an appropriate position on the floor). The Z coordinate in the coordinate system Σ3 is acquired as a value representing the height from the floor (for example, "56 cm"). It is also assumed that the positional relationship between the coordinate systems Σ1 and Σ3 has been adjusted in advance. Note that the above-mentioned coordinate system Σ2 is a coordinate system obtained by translating the coordinate system Σ3 (in three-dimensional space).
[0137] In the next step S73, the controller 31 extracts a region (continuous region) having a height equivalent to the height (e.g., 56 cm) of the object surface corresponding to a predetermined point 315 (a position designated by the user (e.g., near the center of the bed)) in real space. In detail, the controller 31 extracts a continuous region (continuous planar region) having a height (e.g., 56 cm ± 5 cm) within an allowable error range (e.g., ± 5 cm) for the height of the object surface corresponding to the predetermined point 315.
[0138] Then, in step S74, the controller 31 extracts the continuous region (more specifically, the circumscribing quadrilateral (circumscribing rectangle, etc.) of the continuous region) as the bed upper surface region 322 (see the bottom part of FIG. 20).
[0139] The bottom row of Fig. 20 shows a bed top area 322 surrounded by four sides 322a, 322b, 32c, and 322d. Here, the bed top area 322 has a rectangular shape and is surrounded by four sides: long sides 322a and 322b and short sides 322c and 322d. The bed top area 322 is a planar area located at a predetermined height from the floor. Note that the height of the bed top area 322 from the floor may be the height (e.g., 56 cm) of the object surface (bed surface) corresponding to the predetermined point 315 (a position designated by the user), or may be recalculated as the average value of multiple points included in the continuous area.
[0140] Furthermore, the controller 31 extracts, as the bed planar area 321, a projected planar area obtained by projecting the bed upper surface area 322 onto the floor surface (a plane equivalent to the floor surface).
[0141] In step S74, an extended bed area 325 is further set. The extended bed area 325 is a planar area obtained by expanding the bed planar area 321 by a predetermined range. Specifically, the extended bed area 325 is a planar area obtained by expanding the bed planar area 321 by a predetermined length (for example, 1 m to 2 m) in the left-right and up-down directions (directions perpendicular to each of the four sides 322a, 322b, 32c, and 322d) in a planar view. In other words, the extended bed area 325 is a planar area including both the bed planar area 321 itself and the area surrounding the bed planar area 321 (also referred to as the bed peripheral area). The bed peripheral area is the area of the extended bed area 325 excluding the bed planar area 321. The extended bed area 325 is mainly used in the fourth embodiment.
[0142] In this way, the bed area 323 (bed upper surface area 322 and bed plane area 321) and the extended bed area 325 are set. Here, both the bed upper surface area 322 and the bed plane area 321 are set as the bed area 323, but this is not limitative, and only one of the bed upper surface area 322 and the bed plane area 321 may be set.
[0143] Thereafter, the controller 31 sets (automatically sets) the long side 322a, which is closest to the center of the photographed image 110C among the long sides 322a and 322b surrounding the bed upper surface area 322, as the boundary of the bed 92. In other words, the long side 322a is set as the intersection line 215 (see FIGS. 21 and 5). Furthermore, one of the two endpoints of the long side (line segment) 322a is determined as the origin O2 of the coordinate system Σ2. For example, the one endpoint may be selected so that the Y axis extending from the one endpoint to the other endpoint (the Y axis extending along the intersection line 215 (long side 322a)), the Z axis extending vertically upward from the one endpoint, and the X axis extending from the one endpoint toward the other long side 322b (the side opposite to long side 322a) constitute the right-handed coordinate system Σ2. In other words, one of the end points may be selected so that the bed upper surface area 322 is positioned on the +X side of the intersection line 215.
[0144] In this way, the coordinate system Σ2 is set. Then, the relationship between the coordinate system Σ1 and the coordinate system Σ2 is determined based on the relationship between the coordinate system Σ3 and the coordinate system Σ2 (such as the direction and amount of translation between them) and the relationship between the adjusted coordinate system Σ1 and the coordinate system Σ3.
[0145] Here, it is possible to manually set (set on the screen) the positions of the endpoints of the four sides surrounding the bed upper surface area 322. However, such a setting method is cumbersome because it requires specifying the positions of four points. In contrast, by using the method described above, it is sufficient to specify only one point near the center of the bed, making the setting operation relatively easy. Furthermore, an appropriate line segment from among the four sides is automatically set as the intersection line 215, making the setting operation even easier.
[0146] Furthermore, if such automatic settings are not necessarily accurate, adjustment processing (manual setting) based on user operation may be further performed after the above-described automatic setting has been performed. For example, the position of each side may be fine-tuned by moving the positions of the endpoints of the four sides surrounding the bed upper surface area 322 as needed using a mouse or the like. Furthermore, if the long side 322b (or 322a) opposite to the long side 322a (or 322b) automatically set in response to the operation of specifying the predetermined point 315 is suitable as the intersection line 215, an appropriate long side 322b or the like may be specified in response to a user operation. In particular, in a mode for changing (specifying) the "bed boundary," an appropriate long side (322b or the like) may be updated (specified) as the intersection line 215 by a mouse click or the like.
[0147] 4. Fourth Embodiment In the fourth embodiment, a mode will be described in which a behavior-related event of a target person is detected based on the positional relationship (particularly the planar positional relationship (positional relationship within a projection plane (floor surface, etc.))) between the bed area 323 (particularly the bed plane area 321) and multiple specific parts Bi of the target person. Note that the setting process of the bed area 323, etc. (see FIG. 18) may be executed in the same manner as in the third embodiment.
[0148] In this fourth embodiment, two learning models 410 are utilized.
[0149] One learning model 410 (also referred to as 410A) is a learning model that mainly detects behavior-related events of a person near a bed (behavior-related events related to the bed), and is the learning model (learning model for bed-related event detection) described in each of the above-mentioned embodiments. The learning model 410A is machine-trained to output behavior-related events related to a person when information 310 indicating the positional relationship between a predetermined boundary (specifically, the boundary of the bed) or the like and multiple specific parts Bi of the person is input. More specifically, the learning model 410A appropriately learns the correspondence between the coordinate values Xi, Zi of the multiple specific parts Bi and the behavior-related events.
[0150] The other learning model 410 (also referred to as 410B) is a learning model (learning model for detecting bed-unrelated events) that mainly detects behavior-related events of a person at a position away from the bed (behavior-related events not related to the bed). Learning model 410B is similar to learning model 410A. However, learning model 410B is machine-trained to output behavior-related events related to a person when information indicating the heights of multiple specific parts Bi of the person is input. More specifically, the correspondence between the coordinate values Zi of the multiple specific parts Bi and behavior-related events is appropriately learned by learning model 410B.
[0151] Here, in order to accurately detect a variety of behavior-related events, multiple (two in detail) learning models are used instead of a single learning model. Specifically, learning model 410A, which is good at detecting behavior-related events of people mainly near the bed, and learning model 410B, which is good at detecting behavior-related events of people mainly away from the bed, are used.
[0152] Specifically, the learning model 410A is machine-trained to detect various behavior-related events described in the first embodiment, etc. For example, bed-related actions such as "normal lying position (on the bed)," "sitting up (upper body) on the bed," "boundary position (sitting on the edge of the bed, etc.)," and "slipping off (the bed)" can be detected.
[0153] On the other hand, learning model 410B is machine-trained to detect two behavior-related events, "standing" and "lying down." However, without being limited to this, learning model 410B may detect behavior-related events including "falling" and / or "sitting down" in addition to "standing" and "lying down."
[0154] These two learning models 410A (420A) and 410B (420B) are also used during inference. Specifically, inference results (output results) are obtained by the two learning models 410A and 410B based on images, etc., of the same target person taken at the same time.
[0155] In the fourth embodiment, when selecting one of these two inference results, the positional relationship between the bed area 323 and the target person (the specific part Bi of the target person) is used. Specifically, one of the two inference results is selected depending on the distance of the target person from the bed area 323.
[0156] In detail, when it is determined that the target person is located inside (inside) the extended bed area 325 (a planar area obtained by extending the bed planar area 321 by a predetermined range), a behavior-related event of the target person is detected based on the inference result (output result) by the learning model 410A out of the two inference results. That is, the inference result by the learning model 410A (trained model 420A) is selected as the final inference result. Conversely, when it is determined that the target person is located outside (outside) the extended bed area 325, a behavior-related event of the target person is detected based on the inference result (output result) by the learning model 410B out of the two inference results. That is, the inference result by the learning model 410B is selected as the final inference result.
[0157] Fig. 22 shows a situation in which a target person has slipped down beside a bed 92 in a living room. The top part of Fig. 22 shows the interior of the living room as seen from vertically above.
[0158] In the situation shown in the upper part of FIG. 22, the learning model 410A obtains an inference result of "whole body sliding off (from the bed)," and the learning model 410B obtains an inference result of "lying down." In this case, it is determined that the target person is present within the extended bed area 325 (more specifically, outside the bed plane area 321 and inside the extended bed area 325) based on the captured image 110, etc., and based on this determination result, the inference result "whole body sliding off from the bed" from the learning model 410A is selected (determined) as the final inference result. Note that whether the target person is present within a predetermined plane area (e.g., the extended bed area 325) may be determined based on whether the average position of multiple specific parts Bi of the target person is within the predetermined plane area. Alternatively, it may be determined based on whether a predetermined number or more (e.g., a majority) of the multiple specific parts Bi of the target person are present within the predetermined plane area.
[0159] Also, for example, in a situation such as that shown in the upper part of Fig. 23, learning model 410A may obtain an inference result of "whole body sliding (off the bed)" (or "normal lying position"), and learning model 410B may obtain an inference result of "lying down." Fig. 23 shows a situation (abnormal state) in which a target person is lying down in a place away from bed 92 (due to a fall, etc.) in a living room, and the upper part of Fig. 23 shows the room as viewed vertically from above.
[0160] In the case of Figure 23, it is determined based on the captured image 110, etc. that the target person (more specifically, the majority of the multiple specific parts Bi or the average position of the multiple specific parts Bi) is outside the bed area 323 (particularly outside the extended bed area 325). Then, based on such a determination result, the inference result "lying down" from the learning model 410B is selected (determined) as the final inference result.
[0161] Furthermore, FIG. 24 shows a situation in which a target person is lying in a normal state on a bed 92 in a living room, and the upper part of FIG. 24 shows the state of the living room as viewed vertically from above.
[0162] For example, in the situation shown in the upper part of Fig. 24, the learning model 410A obtains an inference result of "normal lying position," and the learning model 410B obtains an inference result of "lying down." In this case, based on the captured image 110, etc., it is determined that the target person (more specifically, a large number of their specific parts Bi) is present within the bed area 323 (and thus within the extended bed area 325), and based on this determination result, the inference result "normal lying position" from the learning model 410A is selected (determined) as the final inference result.
[0163] As described above, in the fourth embodiment, a learning model 410A suitable for detecting behavior-related events related to a bed and a learning model 410B suitable for detecting behavior-related events not related to a bed are individually generated. Then, the two learning models 410 are used to detect the behavior-related events of a target person. This makes it possible to obtain highly accurate detection results by using a learning model optimized (specialized) for each scene (for each scene where a bed is present or absent). Note that in a room without a bed (such as a rehabilitation room), only the learning model 410B may be used (the learning model 410A is not used).
[0164] In the fourth embodiment, the bed plane area 321 and the like are used when selecting an appropriate output result from the output results of the two learning models 410. In particular, the behavior-related events of the target person are detected based on the positional relationship (positional relationship (planar positional relationship) in the projection plane (floor surface, etc.)) between the bed area 323 (particularly the bed plane area 321) and the target person's multiple specific parts Bi. In particular, depending on whether the target person is located inside or outside the extended bed area 325, the output result (inference result of the behavior-related event) from the corresponding learning model 410 of the corresponding learning models 410A and 410B is selected (determined) as the final inference result. This makes it possible to appropriately select an appropriate inference result from the inference results of the two learning models 410 depending on the target person's location.
[0165] In the fourth embodiment, the bed plane area 321 and the like are used when selecting an appropriate output result from the output results of the two learning models 410. However, without being limited to this, the bed plane area 321 and the like may be used when learning and inferring one learning model 410A. For example, in the first embodiment (or the fourth embodiment, etc.), information on whether or not the target person is present in the bed plane area 321 may also be added as an input to the learning model 410A. More specifically, not only height information on the part Bi of the target person but also information on the presence or absence of the target person in the bed plane area 321 may be used to perform learning and inferencing for the learning model 410A.
[0166] <5. Modifications, etc.> Although the embodiment of the present invention has been described above, the present invention is not limited to the above-described contents.
[0167] For example, the specific body parts of a person are not limited to the specific body parts described above, but may be the head, neck, collarbone, etc. of the person.
[0168] Furthermore, in each of the above embodiments, it is assumed that a person moves from outside the bed to on the bed (or from on the bed to outside the bed) only from one side surface 92b of the bed, out of the two left and right sides of the bed 92. Nursing care beds often have a fall prevention fence or the like provided on the other side surface 92c (see FIG. 15), and the above-mentioned configuration (a configuration that takes into account the relative positional relationship with one of the boundary surfaces 210) can accommodate many situations.
[0169] However, the present invention is not limited to this. For example, the above-described concept may be applied to a situation in which a person moves from outside the bed to on the bed (or from on the bed to outside the bed) from either side of the bed. In this case, it is preferable to consider the positional relationship of each specific part Bi with respect to not only the boundary surface 210 on one side but also the boundary surface 220 on the opposite side (see FIG. 15). For example, not only the values Xi and Zi of each specific part Bi but also the value Wi of each specific part Bi may be used as input to the learning model 410. Here, the value Wi is the signed normal distance W (signed shortest distance (differential position) W with respect to the boundary surface 220) of each specific part Bi based on the boundary surface 220 (described next). The boundary surface 220 is a vertical plane including the other side surface 92c of the bed 92 (see FIG. 15). The sign may be defined to be positive (+) at a horizontal position closer to the bed (inside) than the boundary surface 220 (left side in Figure 15), and negative (-) at a horizontal position closer to the bed (right side in Figure 15) than the boundary surface 220.
[0170] Furthermore, in the third and fourth embodiments, several application examples of the setting process of the bed area 323 (see FIG. 18, etc.) have been described, but the present invention is not limited thereto, and the setting process of the bed area 323 (such as a setting process based on a designation operation of a predetermined point) can be applied to various other uses. According to such a setting process, a designation operation of only one point on the bed is sufficient, and the setting operation is much easier than when the setting process is performed with a designation operation of the four sides of the bed area 323. Therefore, the setting operation does not need to be performed by a specialized worker of the device manufacturer, and a general user (such as a care worker) can perform the setting operation. [Explanation of symbols]
[0171] 1. Detection System 10. Detection Device 20 Camera Unit 30 processing units 70,80 Terminal equipment 90 Room 91 beds 92 beds 210,220 Interface (bed interface) 230 Reference horizontal plane 610 Human body (three-dimensional object) 620 Reference plane Bi specific part Ei event Hi Protrusion from the reference plane
Claims
1. a control unit that detects behavior-related events of a target person; Equipped with The control unit detects the behavior-related event of the target person based on a positional relationship between a predetermined boundary and a plurality of specific parts of the target person.
2. The detection device of claim 1, wherein the control unit detects the behavior-related event related to the target person based on information indicating the positional relationship between the predetermined boundary and the multiple specific parts of the target person, using a learning model that has been machine-learned using multiple training data that takes information indicating the positional relationship between the predetermined boundary and the multiple specific parts of the target person as input and outputs the behavior-related event related to the target person.
3. 3. The detection device according to claim 1, wherein the predetermined boundary is a boundary of a bed.
4. The detection device according to any one of claims 1 to 3, characterized in that the control unit detects the behavior-related event by distinguishing between an edge-sitting state in which the target person is sitting on the edge of the bed and an upper-body sitting state in which the target person's upper body is raised on the bed.
5. The detection device according to any one of claims 1 to 4, characterized in that the control unit determines whether the target person is present within a candidate person area in the captured image based on the captured image and depth information of each pixel in the captured image.
6. The detection device according to any one of claims 1 to 5, characterized in that the control unit automatically sets the long side of the bed area set using an adjustment image taken of the target space that is closest to the center of the adjustment image as the predetermined boundary.
7. 6. The detection device according to claim 1, wherein the control unit detects the behavior-related event of the target person based also on a positional relationship between a bed area and the plurality of specific parts of the target person.
8. 8. The detection device according to claim 6, wherein the control unit sets the bed area based on a predetermined point in an adjustment image obtained by capturing the target space.
9. 9. The detection device according to claim 8, wherein the predetermined point is designated in response to a position designation operation by a user within the adjustment image.
10. a) acquiring a positional relationship between a predetermined boundary and a plurality of specific parts of a target person; b) detecting a behavior-related event related to the target person based on the positional relationship; A detection method comprising:
11. A learning model production method for producing a learning model that estimates behavior-related events related to a person, comprising: a) machine-learning a learning model using a plurality of training data sets, the training data sets having information indicating the positional relationship between a predetermined boundary and a plurality of specific parts of the person as input and behavior-related events related to the person as output; A learning model production method comprising:
Citation Information
Patent Citations
Object detection device, object detection method and program
JP2014035302A
Information processing device, information processing method, and program
JP2014174627A
Operation recognition device
JP2017041079A
Behavior detection device, method and program, and monitored person monitoring device
JP2017168105A
Watch support system and control method thereof
JP2018147089A