Detection device and program

JP7902325B2Active Publication Date: 2026-08-07GLORY LTD
View PDF 9 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
GLORY LTD
Filing Date
2025-07-25
Publication Date
2026-08-07

AI Technical Summary

Benefits of technology

【0019】 本発明によれば、人物に関する行動関連事象をより正確に検知することが可能である。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007902325000001
    Figure 0007902325000001
  • Figure 0007902325000002
    Figure 0007902325000002
  • Figure 0007902325000003
    Figure 0007902325000003
Patent Text Reader

Abstract

To provide a technique capable of more accurately detecting a behavior-related event related to a person.SOLUTION: On the basis of a positional relationship between a predetermined boundary and a plurality of specific portions Bi of an object person, a detection device detects a behavior-related event of the object person (for example, [usual lying position (normal lying position)], [(upper body) getting up on the bed], [boundary position], etc.). For example, the detection device uses a learning model, which has been machine-trained using a plurality of teacher data with information indicative of a positional relationship between a boundary surface 210 and the plurality of specific portions Bi of the person as input and the behavior-related event related to the person as output, to detect the behavior-related event related to the object person on the basis of the information indicative of the positional relationship of the plurality of specific portions Bi of the object person with respect to the boundary surface 210.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0004] , , , , , , ,

[0006] , , , ,

[0005] , , , , ,

[0001] The present invention relates to a detection device for detecting action-related events related to a person and technologies related thereto.

Background Art

[0002] There is a technology for detecting action-related events (such as "lying position", "standing position", "getting up", "boundary position", etc.) related to a target person (such as a person to be watched).

[0003] For example, Patent Document 1 describes a technology for monitoring the actions of a person (such as an inpatient) in a hospital or a nursing facility and predicting the actions of the person based on the monitoring results. In this technology, the area where the monitoring target person is present and the status of the monitoring target person (such as "lying position", "standing position", "getting up", "boundary position", etc.) are detected, and an event of the monitoring target person is determined based on changes in the area and status.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, in the technology of Patent Document 1, it may not be possible to accurately detect action-related events related to a person (for example, the status of the monitoring target person).

[0006] For example, when a person is lying in bed, it may not be possible to accurately distinguish between a state where the person is lying normally near the center of the bed (normal lying position) and a state where the person is lying near the end of the bed (boundary position). The boundary position is a dangerous state with a high possibility of falling from the bed, and it is desirable to accurately detect such a state.

[0007] Therefore, the objective of this invention is to provide a technology that can more accurately detect behavioral events related to people. [Means for solving the problem]

[0008] To solve the above problems, the detection device according to the present invention comprises a control unit for detecting behavior-related events of a target person, and the control unit detects the behavior-related events of the target person based on the positional relationship between a predetermined boundary and a plurality of specific parts of the target person. The predetermined boundary comprises a first boundary surface which is a horizontal plane including the bed surface and a second boundary surface which is a vertical plane including the bed side surface. The control unit, in detecting the action-related events of the target person based on the positional relationship between the first boundary surface and a plurality of specific parts of the target person, and the positional relationship between the second boundary surface and a plurality of specific parts of the target person, accepts a user operation to specify a predetermined point indicating the position of the bed in an adjustment image captured of the target space. In response to the user operation, it acquires the three-dimensional position of the object surface corresponding to the two-dimensional position of the predetermined point, sets a plane having a height equivalent to the height of the object surface as the first boundary surface having the height of the bed surface, and, in response to the user operation, acquires a continuous region having a height equivalent to the height of the object surface and a rectangular region in a top view as the bed region, and sets a vertical plane containing the longest side of the bed region closest to the center in the adjustment image as the second boundary surface.

[0009] The control unit may selectively use a first learning model and a second learning model to detect the behavior-related events relating to the target person, wherein the first learning model is a machine learning model that outputs the behavior-related events relating to a person when information indicating the positional relationship between a predetermined boundary and a plurality of specific parts of the person is input, and the second learning model is a learning model different from the first learning model, and the control unit may selectively use the first learning model and the second learning model based on the positional relationship between the target person and the bed area to detect the behavior-related events relating to the target person.

[0010] The control unit may, when the target person is in the bed area, use the first learning model to detect the behavior-related events of the target person based on the positional relationship between the predetermined boundary and a plurality of specific body parts of the target person, and may, when the target person is located at a position horizontally at a distance of a predetermined amount or more from the bed area, use the second learning model to detect the behavior-related events of the target person.

[0011] The second learning model may be a machine learning model that takes information indicating the height of multiple specific body parts of the person as input and outputs the behavior-related events relating to that person.

[0012] To solve the above problems, the present invention provides a program for a computer to perform a process to detect behavior-related events of a target person based on the positional relationship between a predetermined boundary and a plurality of specific parts of the target person, wherein the predetermined boundary has a first boundary surface which is a horizontal plane including the top surface of the bed and a second boundary surface which is a vertical plane including the side surface of the bed, and the program a) detects the behavior-related events of the target person based on the positional relationship between the first boundary surface and a plurality of specific parts of the target person and the positional relationship between the second boundary surface and a plurality of specific parts of the target person, by providing a predetermined point indicating the position of the bed in an adjustment image captured of the target space. The program is characterized by causing a computer to perform the following steps: a) receiving a user operation to specify a point; b) in response to the user operation, obtaining the three-dimensional position of the object surface corresponding to the two-dimensional position of the predetermined point, and setting a plane having a height equivalent to the height of the object surface as the first boundary surface having the height of the bed surface; and c) in response to the user operation, obtaining a continuous region having a height equivalent to the height of the object surface and a rectangular region in a top view as the bed region, and setting a vertical plane containing the longest side of the bed region that is closest to the center in the adjustment image as the second boundary surface.

[0013] To solve the above problems, the detection device according to the present invention comprises a control unit that selectively uses a first learning model and a second learning model to detect behavior-related events of a target person, wherein the first learning model is a machine learning model that outputs behavior-related events concerning a person when information indicating the positional relationship between a predetermined boundary and a plurality of specific parts of the person is input, the predetermined boundary has a first boundary surface which is a horizontal plane including the top surface of the bed and a second boundary surface which is a vertical plane including the side surface of the bed, the second learning model is a learning model different from the first learning model, and the control unit is characterized in that it selectively uses the first learning model and the second learning model based on the positional relationship between the target person and the bed area to detect the behavior-related events concerning the target person.

[0014] The control unit may, when the target person is in the bed area, use the first learning model to detect the behavior-related events of the target person based on the positional relationship between the predetermined boundary and a plurality of specific body parts of the target person, and may, when the target person is located at a position horizontally at a distance of a predetermined amount or more from the bed area, use the second learning model to detect the behavior-related events of the target person.

[0015] The second learning model may be a machine learning model that takes information indicating the height of multiple specific body parts of a person as input and outputs the aforementioned behavior-related events concerning that person.

[0016] To solve the above problems, the present invention provides a program that causes a computer to perform the step of a) selectively using a first learning model and a second learning model to detect behavior-related events of a target person, wherein the first learning model is a learning model that has been trained to output behavior-related events concerning a person when information indicating the positional relationship between a predetermined boundary and a plurality of specific parts of a person is input, the predetermined boundary has a first boundary surface which is a horizontal plane including the top surface of the bed and a second boundary surface which is a vertical plane including the side surface of the bed, the second learning model is a learning model different from the first learning model, and in step a), the behavior-related events concerning the target person are detected by selectively using the first learning model and the second learning model based on the positional relationship between the target person and the bed area. [Effects of the Invention]

[0019] According to the present invention, it is possible to detect behavior-related events concerning individuals with greater accuracy. [Brief explanation of the drawing]

[0020] [Figure 1] This is a schematic diagram showing the detection system. [Figure 2] This is a functional block diagram showing the general configuration of the detection device. [Figure 3]It is a conceptual diagram showing that information on the three-dimensional positions of specific parts of a subject person is acquired (calculated) based on a photographed image and depth information. [Figure 4] It is a vertical sectional view near the bed. [Figure 5] It is a view of the area near the bed seen from above. [Figure 6] It is a conceptual diagram showing the processing in the learning stage in machine learning. [Figure 7] It is a conceptual diagram showing the processing in the inference stage using a learned model. [Figure 8] It is a flowchart showing the processing in the learning stage. [Figure 9] It is a flowchart showing the processing in the inference stage. [Figure 10] It is a diagram showing "getting up from the bed". [Figure 11] It is a diagram showing "sitting upright". [Figure 12] It is a diagram showing "stretching one hand from the bed". [Figure 13] It is a diagram showing "sliding down the lower body from the bed". [Figure 14] It is a diagram showing "sliding down the whole body from the bed". [Figure 15] It is a diagram showing the positional relationship between a specific part and a boundary surface, etc. <所 [Figure 16] It is a flowchart showing the three-dimensional object determination process. [Figure 17] It is a conceptual diagram showing the three-dimensional object determination process. [Figure 18] It is a flowchart showing the setting process of the bed area, etc. [Figure 19] It is a diagram showing an adjustment image obtained by photographing a target space including a bed, etc. [Figure 20] It is a diagram showing that the upper surface area of the bed is detected in response to a bed position designation operation, etc. [Figure 21] ]END]]It is a diagram showing that the bed boundary is automatically set based on the upper surface area of the bed, etc. [Figure 22] It is a diagram showing a situation where the target person has slipped down to the side (perimeter) of the bed. [Figure 23] This diagram shows the subject lying down in a location away from the bed. [Figure 24] This diagram shows the subject in a normal sleeping position in a bed. [Modes for carrying out the invention]

[0021] Embodiments of the present invention will be described below with reference to the drawings.

[0022] <1. First Embodiment> <1-1. System Overview> Figure 1 is a schematic diagram of the detection system 1. As shown in Figure 1, the detection system 1 comprises a plurality of detection devices 10 and a plurality of terminal devices 70, 80. Terminal device 70 is also called a management device 70, and terminal device 80 is also called a portable terminal device 80. Note that in Figure 1, only some of the plurality of detection devices 10 and plurality of terminal devices 70, 80 are shown.

[0023] This section primarily provides examples of how detection system 1 is used in nursing care facilities. However, it is not limited to this, and detection system 1 may also be used in nursing facilities (hospitals, etc.) or private homes.

[0024] As shown in Figure 1, each detection device 10 and each terminal device 70, 80 are connected to each other via a network 108. The network 108 consists of a LAN (Local Area Network) and the Internet, etc. The connection to the network 108 may be wired or wireless. For example, the management device 70 may be wired to the network 108, and each detection device 10 and each mobile terminal device 80 may be wirelessly connected to the network 108. Alternatively, all devices 10, 70, and 80 may be wirelessly connected to the network 108.

[0025] Each detection device 10 is placed in the living room 90 of each person being observed (in this case, a person receiving care) (for example, in each person's individual room). Each detection device 10 (and detection system 1) is a device that detects various "behavior-related events" related to the person being observed (person receiving care, etc.) based on captured images, etc. "Behavior-related events" include the person's actions (movements) themselves and / or states related to the person's actions. "Behavior-related events" are events that should be detected with respect to the person being observed (events that are the target of detection processing), and are also called detection target events. Since the detection device 10 and detection system 1 monitor the behavior of people, etc., they are also called monitoring devices and monitoring systems, etc.

[0026] Examples of "behavioral events" for a person include "normal supine position" (event E0), "sitting up (upper body) in bed" (event E1), "borderline position" (event E2), and "sliding off (the bed)" (event E4).

[0027] Event E0, "Normal Supine Position," represents a state where a person is lying normally in bed (near the center of the bed) (see Figure 5). Figure 5 is a view of the area around the bed from above (vertically above). Also, in Figure 5 and other figures, people and other elements are simplified for illustrative purposes. For example, people are represented by multiple connected ellipses.

[0028] Event E1, "(Upper body) sitting up in bed," is a state in which the upper body is raised while in bed (also referred to as the upper body sitting up state) (see Figure 10).

[0029] Event E2, "Borderline Position," can be further classified into Event E2a, "Sitting Position," and Event E2b, "Lying Position," etc. Event E2a, "Sitting Position," is a sitting position near the edge of the bed (a posture in which the person is sitting on the edge of the bed) (see Figure 11). On the other hand, Event E2b, "Lying Position," is a lying position near the edge of the bed. "Lying Position" includes, for example, a state in which the person is lying very close to the edge of the bed, and / or a state in which the person is lying on the bed with one arm extended outside the bed (from the edge of the bed) ("One-arm extension from the bed") (see Figure 12), etc.

[0030] Event E4, "sliding off," describes a situation where a person is sliding off a bed. Event E4 can be further classified depending on whether the part of the body that is sliding off the bed to the bedside is the upper body, the lower body, or the entire body. Specifically, Event E4 can be further classified into Event E4a, "upper body sliding off the bed," Event E4b, "lower body sliding off the bed" (see Figure 13), Event E4c, "entire body sliding off the bed," etc.

[0031] Figures 5 and 10-13 show examples of each behavior-related event (events E0, E1, E2a, E2b, E4b).

[0032] <1-2. Overview of detection device 10> As shown in Figure 1, the detection device 10 comprises a camera unit 20 and a processing unit 30. The camera unit 20 is installed on the ceiling of the living room 90 and is capable of capturing images of the inside of the living room 90. The processing unit 30 is capable of acquiring the three-dimensional position (especially time-series information) of multiple specific body parts (eyes, ears, nose, chest, waist, shoulders, elbows, wrists, knees, ankles, etc.) of the person being observed, based on images (especially moving images) captured by the camera unit 20. In this example, the camera unit 20 and the processing unit 30 are provided separately, but the device is not limited to this arrangement, and the camera unit 20 and the processing unit 30 may be provided as an integrated unit.

[0033] The camera unit 20 is a three-dimensional camera. The camera unit 20 acquires captured images with depth information. Specifically, the camera unit 20 captures an image (infrared image, etc.) 110 (see Figure 3) of a subject object (person, wall, floor 91, bed 92, etc.) and acquires depth information (depth distance information) 120 for each pixel in the captured image 110. The depth information 120 for each pixel in the captured image is information about the distance (distance from the camera unit 20) to the subject object (person, wall, floor, bed, etc.) corresponding to each pixel in the captured image 110, and is distance information in the direction perpendicular to the captured image 110. In other words, this depth information is distance information (depth distance information) in the direction normal to the plane of the captured image.

[0034] For example, the camera unit 20 includes a sensor unit 23 (Figure 2), which comprises an infrared camera (infrared image sensor, etc.) and an infrared projector. The infrared camera acquires a captured image 110 (for example, an infrared image) of the subject object. The infrared projector projects a predetermined dot pattern using infrared light onto the target object, and the infrared camera also acquires an infrared image of the dot pattern projected onto the target object. The distance to the target object (depth distance information) 120 is calculated (acquired) by utilizing the fact that the geometric shape of the captured dot pattern changes depending on the distance to the target object. The calculation process of this depth distance information is performed by a controller 21, etc., built into the camera unit 20. The controller 21 has a hardware configuration similar to that of the controller 31 (described later), etc. Note that the sensor unit 23 may also be provided with an RGB image sensor, etc., that captures a visible light image (color image, etc.), and a color image may be captured.

[0035] In this way, the camera unit 20 acquires a captured image of the subject (captured image information) 110 (see Figure 3) and depth information (information on the distance to the object corresponding to each pixel (depth distance information)) 120 for each pixel in the captured image.

[0036] Figure 3 is a conceptual diagram showing how 3D position information of specific parts of a subject person is acquired (calculated) based on the captured image 110 and depth information 120. Note that in Figure 3, the captured image 110 and depth information 120 are shown schematically. For example, in the actual captured image 110, the person is captured as it was, whereas in the captured image 110 in Figure 3, the person is represented (simplified) as a figure of connected ellipses. The same applies to the other depth information 120. Also, while the actual depth information 120 contains information on the distance to the object corresponding to each pixel, in the depth information 120 in Figure 3, the distance to the object corresponding to each pixel is expressed by converting it into the density of each pixel. In detail, the magnitude of the distance is expressed by the magnitude of the density (shade).

[0037] First, the processing unit 30 (especially the controller 31) analyzes the captured image 110 and obtains skeletal information (skeletal model information) 140 (see the lower left part of Figure 3) of the subject person based on the texture information of the captured image 110. This skeletal information 140 is information that (simplifies) represents the skeleton of the person using multiple specific parts of the subject person (in detail, the chest, nose, shoulders, elbows, wrists, waist, knees, ankles, eyes, ears, etc.) (mainly joints) and skeletal lines (links) connecting these multiple specific parts. In the skeletal information 140 in Figure 3, each specific part is shown as a "point," and each skeletal line (connecting line) is shown as a "line segment." The skeletal information 140 based on the captured image contains information about multiple specific parts (here, B0 to B17) (such as the planar position information of each part in the image).

[0038] Furthermore, the processing unit 30 acquires planar position information (two-dimensional position information within the captured image in the camera coordinate system) of each specific part of the subject (for example, their respective representative positions) within the captured image 110, based on the subject's skeletal information 140.

[0039] Furthermore, the processing unit 30 also acquires distance information (depth position information in the camera coordinate system) from the camera unit 20 to each specific part based on the depth information 120 of one or more pixels at a planar position (within the captured image 110) corresponding to each specific part. This distance information can also be expressed as depth information in the direction of the normal of the captured image plane (camera line of sight direction).

[0040] In this way, the processing unit 30 acquires information regarding the planar position of each specific part of the subject person within the captured image (two-dimensional position information within the captured image) and distance information to each specific part of the subject person (depth information).

[0041] The processing unit 30 then acquires 3D position information 150 (3D position information within the living space) for each specific part of the subject, based on information regarding the planar position of each specific part of the subject within the captured image and the distance information (depth information) to each specific part, through coordinate transformation, etc. Specifically, the processing unit 30 transforms the position information in the camera coordinate system Σ1 (planar position within the captured image and depth position in the direction of the normal of the captured image) into 3D position information 150 in a coordinate system Σ2 fixed to the living space 90 (more specifically, the bed 92). The camera coordinate system Σ1 is, for example, a 3D orthogonal coordinate system based on three orthogonal axes: two orthogonal axes parallel to the plane of the captured image and one axis extending perpendicular to the plane of the captured image. The transformed coordinate system Σ2 is, for example, a 3D orthogonal coordinate system based on three orthogonal axes: two orthogonal axes parallel to the horizontal plane and one axis extending vertically (height direction). In this detection system 1 (controller 31, etc.), the position of the camera unit 20 in real space is acquired in advance, and adjustments (calibration regarding camera positions, etc.) based on the positions of each camera within the camera unit 20 are performed in advance. In other words, the positional relationship between the two coordinate systems Σ1 and Σ2 is adjusted in advance.

[0042] Furthermore, as will be described later, the transformed coordinate system Σ2 is a coordinate system in which, for example, the bed surface position is the reference position (Z=0) in the vertical direction (Z direction) and the bed boundary surface 210 (see Figure 4, etc.) is the reference position (X=0) in the horizontal direction (X direction).

[0043] In this way, the processing unit 30 acquires (calculates) information on the three-dimensional position of each specific part of the subject person. Here, the subject person in the captured image is, in principle, considered to be the person to be observed. If there are multiple subject people in the camera's field of view, it is sufficient if all of them are identified as the person to be observed. However, this is not limited to this, and the desired object to be observed may be identified from among multiple subject people using information such as height (or using facial recognition technology, etc.).

[0044] In this example, the controller 31 of the processing unit 30 generates skeletal information 140 and 3D position information 150 for each specific part, but this is not limited to this.

[0045] For example, the controller 21 of the camera unit 20 may generate skeletal information 140. Alternatively, the controller 21 of the camera unit 20 may generate 3D position information 150 for each specific body part. More specifically, the controller 21 may acquire information regarding the position of each specific body part of the subject in the captured image and distance information (depth information) to each specific body part, and acquire 3D position information 150 for each specific body part through coordinate transformation or the like. The controller 31 of the processing unit 30 may then acquire the 3D position information 150 for each specific body part from the camera unit 20. Alternatively, both controllers 21 and 31 of the camera unit 20 and the processing unit 30 may cooperate to perform these various processes. In other words, the skeletal information 140 and the 3D position information 150 can be generated by either or both controllers 21 and 31.

[0046] Furthermore, while a 3D camera using a pattern illumination method that acquires depth information for each pixel based on the illumination of a dot pattern is given as an example here, it is not limited to this. The 3D camera may also be a stereoscopic camera or a Time of Flight (TOF) camera, etc. Any type of 3D camera is required to acquire an image of the person being observed and depth information (depth information for each pixel, etc.), which is information about the distance to the subject object in the image.

[0047] Furthermore, since the detection device 10 is a device that performs image processing on captured images, etc., it can also be described as an image processing device. Similarly, the detection system 1 can also be described as an image processing system.

[0048] <1-3. Detailed configuration of detection device 10> Figure 2 is a functional block diagram showing the schematic configuration of the detection device 10.

[0049] As shown in the functional block diagram of Figure 2, the detection device 10 comprises a camera unit 20 and a processing unit 30.

[0050] As described above, the camera unit 20 includes an infrared camera and an infrared projector, etc.

[0051] Furthermore, the processing unit 30 includes a controller (also called a control unit) 31, a storage unit 32, a communication unit 34, and an operation unit 35.

[0052] The controller 31 is a control device built into the processing unit 30 that controls the detection device 10.

[0053] The controller 31 is configured as a computer system equipped with one or more hardware processors (for example, a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit)). The controller 31 performs various processes by executing a predetermined software program (hereinafter also simply referred to as a program) stored in a storage unit (ROM and / or a non-volatile storage unit such as a hard disk) 32 using the CPU, etc. The program (more specifically, a group of program modules) may be recorded on a portable recording medium such as a USB memory stick and read from the recording medium to be installed on the detection device 10. Alternatively, the program may be downloaded via a communication network or the like and installed on the detection device 10.

[0054] The controller 31 acquires captured image information about the subject (person being observed, etc.) from the camera unit 20, and also acquires distance information (depth information) to each specific part of the person from the camera unit 20. Then, the controller 31 acquires 3D position information for each specific part based on the captured image information (particularly the planar position information of each specific part of the subject within the captured image) and the distance information (depth information of the subject).

[0055] Furthermore, the controller 31 detects behavior-related events of the target person based on the positional relationship (described later) between predetermined boundaries, etc., and multiple specific body parts of the target person. Examples of predetermined boundaries include the boundary of the bed 92 (in detail, the interface 210 between the space on the bed 211 and the external space 212 to the side of the bed (see Figures 1 and 4, etc.)).

[0056] Here, the controller 31 uses the learning model 410 to detect behavior-related events of the target person.

[0057] As the learning model 410, for example, a neural network model consisting of multiple layers is used. Then, the weighting coefficients, etc. (learning parameters) between the layers (input layer, (one or more) hidden layers, output layer) of the neural network model are adjusted using a predetermined machine learning method. The learning model 410 after being trained by machine learning is also called the trained model 420. Specifically, the learning parameters of the learning model 410 (learner) are adjusted using a predetermined machine learning method, and a trained learning model 410 (trained model 420) is generated (see Figure 6).

[0058] Therefore, first, the controller 31 performs the learning phase (see Figure 6) in machine learning. Specifically, the controller 31 pre-trains the learning model 410 using multiple training data sets. Each training data set uses information indicating the positional relationship between a predetermined boundary and multiple specific body parts of a person as input, and behavior-related events concerning the person as output. Then, using a predetermined machine learning method, the weighting coefficients (learning parameters) between the layers (input layer, (one or more) hidden layers, output layer) of the neural network model are adjusted. As a result, the trained learning model 410 (trained model 420) is generated (produced).

[0059] Subsequently, the controller 31 performs the inference stage (see Figure 7) using the trained learning model 410 (420). Specifically, the controller 31 uses the learning model 410 (trained model 420), which has been machine-learned using the multiple training data, to detect behavior-related events concerning the target person (also referred to as the person to be detected or the person to be judged) based on information indicating the positional relationships of multiple specific body parts of the target person with respect to a predetermined boundary, etc. More specifically, the controller 31 inputs the positional relationships of multiple specific body parts of the target person with respect to a predetermined boundary, etc., into the trained model 420 and obtains the output from the trained model 420 (the detection result of behavior-related events of the target person). The positional relationships between the predetermined boundary, etc., and multiple specific body parts of the target person will be described in detail later.

[0060] Here, the detection device 10 (controller 31) is assumed to perform both the learning phase processing and the inference phase processing in machine learning. However, it is not limited to this, and for example, the learning phase processing and the inference phase processing may be performed by separate devices.

[0061] Furthermore, the controller 31, in cooperation with the communication unit 34, transmits the detection result to the terminal devices 70 and 80, and the terminal devices 70 and 80 output the detection result (display output and / or audio output, etc.).

[0062] The memory unit 32 is composed of a storage device such as a hard disk drive (HDD) or a solid-state drive (SSD). The memory unit 32 stores the aforementioned programs and various data. For example, the memory unit 32 stores time-series data of the three-dimensional positions of multiple specific body parts of the observed person, as well as various data and programs used for training and utilizing the learning model 410.

[0063] The communication unit 34 is capable of performing network communication via the network 108. Various protocols, such as TCP / IP (Transmission Control Protocol / Internet Protocol), are used in this network communication. By using this network communication, the detection device 10 can exchange various types of data with a desired partner (for example, terminal devices 70, 80).

[0064] The operation unit 35 includes an operation input unit 35a that receives operation input to the detection device 10, and a display unit 35b that outputs various information.

[0065] <1-4. Processing of the learning phase of learning model 410> First, let's explain the process during the learning phase.

[0066] Figure 8 is a flowchart showing the processing during the learning phase. The processing shown in Figure 8 is executed by the controller 31, etc.

[0067] Furthermore, Figure 8 also illustrates a method for generating a pre-trained model. In this application, generating a pre-trained model 420 means producing (manufacturing) a pre-trained model 420, and "method for generating a pre-trained model" means "method for producing a pre-trained model."

[0068] First, in step S11, the controller 31 acquires information 310 that shows the positional relationship between a predetermined boundary, etc., and multiple specific parts of the target person (for example, B0 to B17).

[0069] Specifically, first, relative position information of multiple specific parts B0 to B17 with respect to a predetermined boundary (for example, a boundary surface 210 that virtually separates two adjacent spaces 211 and 212 in the horizontal direction) is acquired as the information 310. Here, as shown in Figures 4 and 5, the boundary surface 210 is the boundary surface between the space 211 on the bed 92 and the lateral external space 212 of the bed 92. In other words, the boundary surface 210 is a vertical plane that includes one side surface 92b of the bed 92. Figure 4 is a vertical cross-sectional view of the vicinity of the bed 92 (a cross-sectional view obtained by cutting the bed 92 etc. with a vertical cross-section that extends in both the vertical direction and the left-right direction of the bed 92), and is a view from the foot side to the head side of a person lying on the bed 92. Figure 5 is a view of the vicinity of the bed 92 from above (top view).

[0070] The positional relationship between the boundary surface 210 and each specific part (B0 to B17) is expressed by the signed normal distance X with respect to the boundary surface 210 (signed shortest distance (difference position) X with respect to the boundary surface 210). The sign is positive (+) at horizontal positions on the bed side (inside (right side in Figure 4)) of the boundary surface 210, and negative (-) at horizontal positions on the bed side (left side in Figure 4) of the boundary surface 210. Note that the signed normal distance X with respect to the boundary surface 210 can also be expressed as the relative position in the X direction with respect to the boundary surface 210 (of each specific part).

[0071] Furthermore, in step S11, the controller 31 also acquires relative position information between a predetermined reference horizontal plane 230 (a horizontal plane 230 that virtually separates two adjacent spaces 231 and 232 adjacent in the vertical direction (height direction)) and a plurality of specific parts B0 to B17 as information 310. Here, the reference horizontal plane 230 is a horizontal plane that extends two-dimensionally in the horizontal direction and is a plane 230 (the upper surface and its extension plane) that includes the upper surface 92a of the bed. The reference horizontal plane 230 is a horizontal plane that virtually separates the space above the bed (upper space) 231 and the space below the bed (lower space) 232.

[0072] The positional relationship between the reference horizontal plane 230 and each specific location (B0 to B17) is expressed by the signed normal distance Z with respect to the reference horizontal plane 230 (the signed shortest distance (difference position) Z with respect to the reference horizontal plane 230). The sign is positive (+) for vertical (height) positions above (upper side) the reference horizontal plane 230, and negative (-) for vertical (height) positions below (lower side) the reference horizontal plane 230. The signed normal distance Z with respect to the reference horizontal plane 230 can also be expressed as the relative position (of each specific location) in the Z direction with respect to the reference horizontal plane 230.

[0073] The positional relationships described above can be easily represented by adopting the following coordinate system Σ2.

[0074] The coordinate system Σ2 is an XYZ Cartesian coordinate system with its origin O2 (see Figures 4 and 5) set at a point on the intersection line 215 between the top surface 92a of bed 92 and the side surface 92b of bed 92 (see Figure 4) (in this case, the position of the foot end of the bed 92). In coordinate system Σ2, the vertical direction is the Z direction, the left-right direction (horizontal direction) of the bed is the X direction, and the long-right direction of the bed is the Y direction. Furthermore, coordinate system Σ2 is a coordinate system in which the Z-direction position of the top surface of the bed (reference horizontal plane 230) is the reference position (Z=0) in the vertical direction (Z direction), and the bed boundary surface 210 is the reference position (X=0) in the horizontal direction (left-right direction (X direction) of the bed). In Figures 4 and 5, the coordinate axes in the XYZ coordinate system are shown as appropriate. Figure 4 can also be described as a diagram showing the XZ plane, and Figure 5 as a diagram showing the XY plane. Furthermore, the interface 210 can also be described as the plane where X=0, and the reference horizontal plane 230 can also be described as the plane where Z=0.

[0075] Then, the X-coordinate values ​​Xi and Z-coordinate values ​​Zi of multiple (in this case, 18) specific body parts Bi (i=0 to 17) are acquired as information 310 indicating the positional relationship between a predetermined boundary, etc., and multiple specific body parts of the subject (for example, B0 to B17). Specifically, a total of 36 values ​​are acquired as information 310: X0 and Z0 for specific body part B0, X1 and Z1 for specific body part B1, X2 and Z2 for specific body part B2, X3 and Z3 for specific body part B3, ..., and X17 and Z17 for specific body part B17. Note that the Y-coordinate values ​​of each specific body part Bi are not used as information 310 here.

[0076] Thus, the relative positions of multiple specific parts Bi are calculated with reference to the bed interface 210 (specifically, the interface 210 between the space on the bed 211 and the external space 212 to the side of the bed (the plane where X=0)). Furthermore, the relative positions of multiple specific parts Bi are also calculated with reference to the reference horizontal plane 230 (specifically, the horizontal plane including the top surface 92a of the bed (the plane where Z=0)). In other words, the relative positions of multiple specific parts Bi are determined as relative positions with respect to the interface 210 and the reference horizontal plane 230. To put it another way, the relative positions of multiple specific parts Bi are determined as relative positions with respect to the reference axis included in the interface 210 (and the reference horizontal plane 230) (specifically, the intersection line 215 (Y-axis) of the interface 210 and the reference horizontal plane 230). Furthermore, the relative positions of multiple specific parts Bi can also be expressed as two-dimensional positions obtained by projecting the three-dimensional position of each specific part Bi onto a plane (projection plane) (such as the plane of Y=0) that is orthogonal to both the interface 210 and the reference horizontal plane 230. Additionally, the intersection line 215 can be expressed as a line (boundary line) that is included in both the interface 210 and the reference horizontal plane 230.

[0077] Next, in step S12, the controller 31 assigns corresponding "behavior-related events" (for example, "normal supine position" or "sitting position") as labels (ground truth data) to the information 310 that shows the positional relationship between a predetermined boundary, etc., and multiple specific body parts of a person. More specifically, for example, among the events E0, E1, E2a, E2b, E4a, E4b, and E4c mentioned above, the corresponding "behavior-related events" are identified and assigned as labels.

[0078] This matching (the process of identifying the corresponding behavior-related event) can be performed by an operator (user). Specifically, the operator determines the "behavior-related event" corresponding to the information 310, and the "behavior-related event" related to the determination result is assigned to the detection device 10 in accordance with the operator's input. That is, the controller 31 assigns a label (correct data) to the information 310 in accordance with the input. In this way, training data is generated that takes the information 310 as input and the behavior-related event as output.

[0079] By repeatedly executing steps S11 and S12, multiple training data sets are generated. Note that the repeated portion is not shown in Figure 8.

[0080] Next, in step S13, machine learning is performed on the learning model 410 using these multiple training data. When information 310 indicating the positional relationship between a predetermined boundary, etc., and multiple specific body parts of a person is input, the learning model 410 is trained to output behavior-related events concerning the person. As a result, a trained model 420 that estimates behavior-related events concerning a person is generated (produced) (step S14).

[0081] In this case, for example, in event E0 "normal supine position" (see Figure 5), the X coordinate value Xi and Z coordinate value Zi of multiple specific body parts Bi (i=0~17) are all positive values. Figure 5 is a diagram showing event E0 "normal supine position," etc., and is a top view (view from above) of a person in a normal supine position.

[0082] In event E0, the values ​​Xi and Zi for each specific part Bi are both positive. Also, since the person is in the vicinity of the bed surface 92a (see Figure 4) (e.g., inside the elliptical region shown by the dashed line in Figure 4), the value Zi is relatively small (for example, 100 mm). Through machine learning as described above, a trained model 420 is generated that outputs event E0 "normal supine position" as an action-related event in response to input information 310 with these characteristics.

[0083] On the other hand, consider a situation where a person transitions from event E0 "normal lying position" to event E1 "getting up in bed" (see Figure 10) due to a change in their posture. In this case, for example, a specific body part B1 ("chest") moves to position P1 in Figure 4. As a result, with respect to the specific body part B1, the value Z1 changes to a relatively large value (for example, "400" mm), while the value X1 remains almost the same. Similarly, for other specific body parts B0, B2, B5, B14~B17, the value Zi changes to a relatively large value, while the value Xi remains almost the same. Through machine learning as described above, a trained model 420 is generated that outputs event E1 "getting up in bed" as an action-related event in response to input information 310 having these characteristics.

[0084] For each other behavior-related event, the information 310 at the time of occurrence of that event also has its own unique characteristics.

[0085] For example, consider a situation where the patient transitions from event E1, "getting up in bed," to event E2a, "sitting on the edge of a bed" (see Figure 11). In this case, specific body parts B10 (right ankle) and B13 (left ankle) move to the vicinity of point P3 in Figure 4. As a result, the values ​​Xi and Zi for specific body parts B10 and B13 both change to negative (-) values. Also, the values ​​Xi (absolute values) for specific body parts B9 (right knee) and B12 (left knee), as well as specific upper body parts B0-B7, B14-B17, etc., change to relatively small values. In particular, the values ​​Xi for specific left and right upper body parts (left shoulder and right shoulder, and left elbow and right elbow, etc.) become relatively small on both sides (to roughly the same extent).

[0086] Furthermore, consider a scenario where a person reaches out from the edge of the bed to grab a plastic bottle or similar item, transitioning from event E0 (Figure 5) to event E2b (see Figure 12). In this case, specific body parts B6 (left elbow) and B7 (left wrist) (or specific body parts B3 (right elbow) and B4 (right wrist)) move to the vicinity of point P2 in Figure 4. As a result, although the values ​​Zi for specific body parts B6 and B7 (or specific body parts B3 and B4) remain almost the same, the values ​​Xi for these specific body parts B6 and B7 change to negative (-) values. Additionally, for other specific body parts of the upper body, such as B5 (left shoulder), the absolute values ​​of their values ​​Xi change to relatively smaller values ​​(while their values ​​Zi remain almost the same).

[0087] A state in which a person moves to the edge of the bed while sleeping (boundary position) is also one of the events E2b. In this case, some parts of the body, such as the elbows, wrists, knees, and ankles, move to the edge of the bed. As a result, the absolute value of the value Xi of at least some of the specific body parts B0-B7, B14-B17, etc. (for example, several parts on the right side and / or several parts on the left side) changes to a value smaller than a certain extent.

[0088] As described above, the coordinate values ​​change in accordance with the changes in each event. In other words, each event has its own unique characteristics with respect to the combination of coordinate values ​​Xi and Zi of multiple specific body parts. These unique characteristics are then learned by the learning model 410. Specifically, the correspondence between the coordinate values ​​Xi and Zi of multiple specific body parts and the action-related events is appropriately learned by the learning model 410. That is, according to the machine learning described above, a trained model 420 is generated that outputs the corresponding action-related event in response to the input of information 310 having the characteristics indicated by each coordinate value.

[0089] <1-5. Inference stage processing using the 420 pre-trained models> Next, we will explain the inference process (the stage of estimating behavior-related events) using the pre-trained model 420.

[0090] Figure 9 is a flowchart of the processing in the inference stage. The processing shown in Figure 9 is performed by the controller 31, etc. In the processing of the inference stage, the detection device 10 (controller 31) uses the trained model 420 to detect behavior-related events of the target person based on information 310 that shows the positional relationship between a predetermined boundary (boundary surface 210) and multiple specific body parts B0 to B17 of the target person.

[0091] Therefore, in step S31, the same processing as in step S11 (Figure 8) is performed with respect to the person to be judged. As a result, information 310 is obtained that shows the positional relationship between multiple specific body parts Bi of the person to be judged and predetermined boundaries, etc. For example, this information 310 includes the positional relationship between multiple specific body parts Bi of the person to be judged and the bed boundary surface 210, as well as the positional relationship between multiple specific body parts Bi of the person to be judged and the reference horizontal plane 230. More specifically, the X coordinate value Xi and Z coordinate value Zi of the specific body parts Bi (i=0~17) are obtained as this information 310.

[0092] In the next step S32, the controller 31 inputs the information 310 into the learning model 410 (trained model 420) and obtains the output (behavior-related event) from the learning model 410 as the estimation result (inference result). Specifically, one of the events E0, E1, E2a, E2b, E4a, E4b, or E4c described above is output from the learning model 410 and obtained as the estimation result.

[0093] In step S33, the estimation result is output. Specifically, the controller 31 displays the name of the estimated behavior-related event (e.g., "normal supine position," "sitting position") on the display unit 35b. Furthermore, the controller 31 transmits the estimation result to the terminal devices 70 and 80 via communication or other means, and displays it on the respective displays of the terminal devices 70 and 80. In particular, if an event other than event E0 occurs, it is preferable that the terminal devices 70 and 80 notify the terminal user of the occurrence of the event with a warning display or warning sound.

[0094] <1-6. Effects of the Embodiments> Through the above process, the learning model 410 is trained using training data that takes information 310 indicating the positional relationship between a predetermined boundary, etc., and multiple specific body parts Bi of a person as input, and outputs behavior-related events concerning the person. Therefore, a learning model 410 is generated that accurately reflects the positional relationship between the predetermined boundary, etc., and multiple specific body parts Bi of a person.

[0095] In particular, the 3D position information of multiple specific parts Bi (such as 3D position information in coordinate system Σ1) is not used directly as input, but rather the value (Xi) converted into information indicating the relative positional relationship with the boundary surface 210 (relative positional information) is used as input to the learning model 410. This makes it possible to appropriately reflect the positional relationship between each specific part Bi and the boundary surface 210 (the distance between each specific part and the boundary surface 210, and whether each specific part Bi is inside or outside the boundary surface 210) in the learning model 420.

[0096] Similarly, a value (Zi) converted into information indicating the relative positional relationship with the reference horizontal plane 230 (bed surface 92a, etc.) (relative positional information) is used as input to the trained model 410. This makes it possible to appropriately reflect the positional relationship between each specific part Bi and the reference horizontal plane 230 (the distance between each specific part Bi and the reference horizontal plane 230, and whether each specific part Bi is located above or below the reference horizontal plane 230) in the trained model 420.

[0097] Furthermore, by using only the information from the other two dimensions (Xi, Zi) and not the information in the Y direction (Yi) of the three-dimensional information, it is possible to appropriately reduce the amount of information and achieve efficient learning.

[0098] Through this machine learning process, the weighting parameters and other settings in the learning model 410 are appropriately adjusted. As a result, a trained model 420 is generated that is trained to output appropriate behavior-related events even for new (unknown) inputs that indicate the positional relationship between a predetermined boundary and multiple specific parts Bi.

[0099] Furthermore, in the inference stage of the above embodiment, the detection device 10 (controller 31) detects behavior-related events of the target person based on the positional relationship between a predetermined boundary, etc., and multiple specific body parts Bi of the target person. Specifically, using the trained model 420 described above, behavior-related events concerning the target person are detected based on information 310 indicating the positional relationship of the multiple specific body parts Bi with respect to the boundary surface 210. In other words, behavior-related events concerning the target person are detected using the trained model 420, which has been trained to output appropriate estimation results. Therefore, it is possible to detect behavior-related events concerning the target person more accurately.

[0100] In particular, it is possible to obtain estimation results (behavior-related information) that appropriately reflect the positional relationship between each specific part and the interface 210 (the distance between each specific part and the interface 210, and whether each specific part is located inside or outside the interface).

[0101] Therefore, for example, it is possible to accurately distinguish between a state in which the subject is lying normally near the center of the bed (normal supine position) and a state in which the subject is lying near the edge of the bed (boundary position), and to detect both states. Alternatively, it is possible to accurately distinguish between a state in which the subject is sitting on the edge of the bed ("edge sitting position") and a state in which the subject's upper body is raised on the bed ("raising on the bed" state), and to detect both states (behavior-related events). In other words, it is possible to accurately detect behavior-related events in which the positional relationship with the boundary surface 210 is an important identifying element.

[0102] Furthermore, in the inference stage of the above embodiment, behavior-related events of the target person are detected, in particular, based on the positional relationship between the reference horizontal plane 230 and multiple specific body parts (B0 to B17) of the target person. Specifically, using the trained model 420 described above, behavior-related events concerning the target person are detected based on information 310 that also shows the positional relationship of multiple specific body parts Bi with respect to the reference horizontal plane 230. Therefore, it is possible to detect behavior-related events concerning the target person with greater accuracy. More specifically, it is possible to accurately detect behavior-related events in which not only the positional relationship with the interface 210 but also the positional relationship with the reference horizontal plane 230 is an important identification element. For example, it is possible to accurately distinguish between events E2a "sitting on the edge" (see Figure 11) and E4b "lower body sliding off the bed" (see Figure 13) and detect both events.

[0103] In the above embodiment, at each stage (especially the estimation stage), it is not necessary for relative position information, etc., for all of the multiple specific body parts to be input to the learning model 410. Specifically, relative positional relationships with the interface 210, etc., may be obtained only for the specific body parts Bi that have a large weight (weight in the learning model 410) corresponding to a particular behavior-related event. Even in this case, it is possible to suitably determine the specific behavior-related event. For example, when determining "getting up from bed" (see Figure 10), the relative positional relationship of the upper body specific body parts Bi with the interface 210, etc., is a particularly important element (an element with a large weight). Therefore, if positional information for several major specific body parts of the upper body is obtained, it is possible to appropriately determine "getting up from bed" (see Figure 10) even if positional information for specific body parts of the lower body (e.g., "knees" and "ankles") is not available.

[0104] Furthermore, when distinguishing and recognizing (detecting) two similar behavioral events, it is possible to obtain a favorable recognition result if the positional information of some of the specific body parts Bi that have a large weight corresponding to the two behavioral events is acquired. For example, in order to distinguish and detect "getting up in bed" (see Figure 10) and "sitting on the edge of a bed" (see Figure 11), the relative positional relationship between the specific body parts Bi of the upper body and the interface 210 is a particularly important factor (a factor with a large weight). More specifically, in "sitting on the edge of a bed," the relative positions of the specific body parts such as the left and right "shoulders" and "elbows" and the interface 210 are close. In "getting up in bed," the distance Xi between one of the left and right specific body parts such as the "shoulder" and "elbow" (right shoulder, right elbow, etc.) and the interface 210 is small, while the distance Xi between the other (left) specific body part such as the "shoulder" and "elbow" (left shoulder, left elbow, etc.) and the interface 210 is large. By (substantially) learning the left-right differences in the relative position between specific parts of the upper body and the interface 210, the two types of behavior-related events can be distinguished from each other and appropriately detected.

[0105] <1-7. Variations, etc.> In the above embodiment, the coordinate values ​​Xi and Zi of each specific part Bi in the XYZ Cartesian coordinate system are used, but the embodiment is not limited to this.

[0106] For example, the coordinate values ​​Ri and θi of each specific part Bi in a cylindrical coordinate system (R, θ, Y) with the Y axis of the above XYZ Cartesian coordinate system as the reference axis (rotation center axis) may be used (see Figure 15). That is, the coordinate values ​​θi and Ri may be used instead of the coordinate values ​​Xi and Zi. Here, the coordinate value θi is the rotation angle θi around the origin O2 in the XZ plane in Figure 4. More specifically, the rotation angle θi should be defined such that a clockwise rotation angle is considered a positive (+) rotation angle (a counterclockwise rotation angle is considered a negative (-) rotation angle) with the vertical upward direction (+Z direction) as the reference (θi=0). Also, the coordinate value Ri is the distance (difference distance) from the reference axis (intersection line 215).

[0107] Thus, the difference position information Ri of each specific part with respect to the intersection line 215 (Y-axis) of the interface surface 210 and the reference horizontal plane 230, and the angle information θi of each specific part with respect to the said intersection line, may be used as relative position information of each specific part with respect to the interface surface 210 and the reference horizontal plane 230.

[0108] Furthermore, in the above embodiment, both the value Xi and the value Zi of each specific part Bi are used as input to the learning model 410, but this is not limited to this. For example, only the value Xi may be used as input to the learning model 410. This makes it possible to obtain learning results that reflect at least the relative positional relationship with the interface surface 210. Similarly, of the values ​​Ri and θi in the cylindrical coordinate system, only the value θi may be used as input to the learning model 410.

[0109] However, in order to obtain learning results (and inference results) that also reflect the relative positional relationship with the reference horizontal plane 230, it is preferable that both the value Xi and the value Zi of Bi for each specific part (or both the values ​​Ri and θi, etc.) are used as input to the learning model 410.

[0110] Furthermore, in the above embodiment, events E0, E1, E2a, E2b, E4a, E4b, and E4c are mainly exemplified as behavior-related events, but the embodiment is not limited to these. For example, other events (e.g., "standing outside of bed") may be included as behavior-related events. Also, the events may be more finely classified, or conversely, more broadly classified. Alternatively, the behavior-related events may be any combination of all or some of these events.

[0111] <2. Second Embodiment> The second embodiment is a modification of the first embodiment. The differences from the first embodiment will be explained below.

[0112] In the second embodiment, a technique will be described that enables accurate extraction of a person on a bed or a person near a wall, etc., in the first embodiment and the like.

[0113] In the first embodiment, etc., a human body (human body texture) is recognized based on the texture information of the captured image 110, and the skeletal information (skeletal model information) 140 of the person is acquired.

[0114] However, when recognizing a human body based solely on the texture information (2D information) of the captured image 110, it is possible that a flat texture similar to the texture of a human body may be mistakenly identified as the texture of a real (3D) human body.

[0115] For example, the above-mentioned misrecognition can occur when the texture of a human body is detected within a poster (a poster pasted on a wall, etc.). Alternatively, depending on the accuracy of the human body texture recognition, a flat texture on the bed surface (or near a wall, etc.) that coincidentally resembles the texture of a real (three-dimensional) human body may be mistakenly recognized as a real human body. Furthermore, the above-mentioned misrecognition can occur when the shadow of a person in a captured image (such as an image taken with an RGB camera) is detected as a human body texture.

[0116] Therefore, in this second embodiment, a technique is described that can avoid misidentifying a flat texture as a human texture. In detail, a technique for avoiding misidentification that utilizes the three-dimensionality of a human will be described.

[0117] Specifically, the controller 31 determines the presence or absence of a person within the candidate person region in the captured image 110 based on the captured image 110 and the depth information 120 of each pixel in the captured image 110. More specifically, as shown in Figures 16 and 17, the controller 31 determines a reference plane 620 in the candidate person region of the captured image 110 based on the depth information 120 of each pixel in the captured image 110. Then, the controller 31 determines the presence or absence of the target person within the candidate person region based on the amount of projection of the point cloud protruding from the reference plane 620.

[0118] Such processing is performed when confirming the presence of a person in step S11 (Figure 8) and / or step S31 (Figure 9) (or immediately before each of these), etc. Figure 16 is a flowchart illustrating such processing, and Figure 17 is a conceptual diagram illustrating the processing. The left column of Figure 17 shows a cross-section of real space when a three-dimensional object (real) human body 610 is present, and the right column of Figure 17 shows a cross-section of real space when a three-dimensional object (human body 610) is not present.

[0119] More specifically, as shown in Figure 16, in step S51, the controller 31 first extracts a determination target region (a candidate person region and its surrounding region) that includes an object to be determined as to whether or not it is a person (an object to be determined) based on the texture information of the captured image 110.

[0120] In detail, first, the controller 31 performs image processing (feature analysis) on the captured image 110 to extract (detect) areas that are highly likely to be human (having a certain degree of such possibility) based on the texture of the captured image 110 (also referred to as human candidate areas). These human candidate areas are, for example, areas from which human skeletal information is extracted (areas with a human shape, etc.). Image recognition processing using neural networks, etc., can be used for such extraction (detection) processing. The controller 31 then sets a bounding rectangle surrounding the human candidate area within the captured image 110, and sets the area inside this bounding rectangle (the human candidate area and its surrounding area) as the area to be judged (see also the top row of Figure 17). Note that in Figure 17, the area to be judged, which originally has a two-dimensional extent, is shown one-dimensionally (as a one-dimensional range in one cross-section).

[0121] Next, in step S52, the controller 31 determines a reference plane 620 (reference plane in real space) in the target area based on the depth information 120 of the point cloud (pixel group) within the target area (see also the second row from the top in Figure 17). For example, the bed surface, floor surface, wall surface, or slope may be extracted as the reference plane 620.

[0122] The reference plane can be determined, for example, using the RANSAC (RANDOM SAmple Consensus) method. Specifically, first, three arbitrary points are extracted from the point cloud (a set of points with 3D positions determined for each pixel in the region) within the target region (preferably the region surrounding the candidate person region), and a plane (provisional plane) passing through these three points is provisionally set. Then, if the number of points in the point cloud that are within a predetermined allowable range (for example, a few millimeters to tens of millimeters) from the plane is the largest number so far, the plane composed of these three points is updated as the solution plane (optimal plane). In other words, if the number of points in the vicinity of the provisional plane is greater than the number of points in the vicinity of the current optimal plane, the provisional plane is determined as the new optimal plane. After that, the same process (setting the provisional plane, and comparing the provisional plane with the optimal plane, etc.) is repeated many times (a predetermined number of times) for the new arbitrary three points. As a result, the solution plane finally obtained is determined as the reference plane 620.

[0123] In step S53, the point cloud within the region to be judged, excluding the point cloud that constitutes the reference plane 620 (points located in the vicinity of the reference plane), is determined as the point cloud that protrudes from the reference plane 620. In the third row from the top of Figure 17, the point cloud that constitutes the reference plane 620 is represented by white circles, and the remaining point cloud, represented by black circles, is the point cloud that protrudes from the reference plane 620 (also called the protruding point cloud). Since the point cloud that protrudes from the reference plane 620 does not constitute the reference plane 620, it is also called the non-planar point cloud.

[0124] Then, in step S54, the amount of projection Hi of each point Pi from the reference plane 620 (the amount of projection in the direction normal to the reference plane) is calculated for each point Pi in the point cloud (non-planar point cloud) that protrudes from the reference plane 620. Then, the average value Hv of the projection amounts Hi of multiple points Pi is calculated.

[0125] If the average value Hv is greater than the threshold TH1 (for example, 80 mm), it is determined that a three-dimensional object (and therefore a human body) is present. On the other hand, if the average value Hv is less than the threshold TH1, it is determined that a three-dimensional object (and therefore a human body) is not present. In this way, the controller 31 determines the presence or absence of a person (three-dimensional object) within the determination area based on the protrusion amount Hi.

[0126] As described above, the controller 31 determines a reference plane 620 in the target area (the candidate person area and its surrounding area) based on the depth information 120 of the point cloud (pixel group) within the target area. Then, the controller 31 determines the presence or absence of a person within the target area based on the amount of protrusion H (Hi, etc.) of the point cloud that protrudes from the reference plane 620. With this, even if a planar texture is mistakenly detected as a human body within the target area, it is possible to accurately determine (confirm) whether the texture is three-dimensional or not based on the amount of protrusion Hi from the reference plane.Therefore, it is possible to avoid misidentifying a planar texture as the texture of a real (three-dimensional) person.In this way, it is possible to accurately determine the presence or absence of a target person as a three-dimensional object.

[0127] In particular, a reference plane 620 is determined, and the relative displacement with respect to the reference plane 620 (such as the relative distance (difference) from the reference plane 620) is used to appropriately determine whether or not there is a protrusion (three-dimensional object) relative to the reference plane 620. In other words, it is not always necessary to determine the absolute value of the height of the reference plane 620, and the presence or absence of a three-dimensional object on the reference plane 620 can be appropriately determined.

[0128] It should be noted that the concept of the second embodiment is described as a modification of the first embodiment, but is not limited thereto. For example, here, similar to the first embodiment, it is assumed that the position of the camera unit 20 in real space is acquired in advance, and after adjustments (calibration related to the camera position, etc.) based on the position of the camera unit 20 are performed, the positions of each specific part in the coordinate system Σ2, etc., are measured. However, the concept of the second embodiment can be applied not only to this situation but also to other situations.

[0129] For example, this concept can be applied even when simply determining whether or not a three-dimensional human body actually exists in a candidate area for a human body on a floor or wall. This concept itself does not necessarily require determining the absolute value of the height of the reference plane 620. Therefore, in such cases, it is possible to determine the presence or absence of a three-dimensional object using the relative displacement of each point cloud with respect to the reference plane 620, without performing adjustments based on the position of the camera unit 20 in real space (such as setting camera parameters (camera installation position information)).

[0130] <3. Third Embodiment> In each of the above embodiments, the coordinate system Σ2 is exemplified as an XYZ Cartesian coordinate system with the origin O2 (see Figures 4 and 5) being a point on the intersection line 215 between the upper surface 92a of the bed 92 and the side surface 92b of one bed 92 (see Figure 4). This intersection line 215 may be set, for example, as shown below. In the third embodiment, the details of setting the coordinate system Σ2 (particularly the intersection line 215) will be described. Here, the setting process for the intersection line 215 is performed along with the setting process for the bed area 323 (see Figure 21, etc.).

[0131] Figure 18 is a flowchart showing the setting process for the bed area 323, etc. (processing by the controller 31).

[0132] As shown in Figure 18, in step S71, captured images 110C and depth information 120C relating to the target space (the target space for detection processing) including the bed 92 are acquired (see also Figure 19).

[0133] On the left side of Figure 19, an adjustment image (calibration image 110C for pre-adjustment) of the target space including the bed 92 is shown, and on the right side of Figure 19, the depth information 120C of each pixel in the captured image 110C is visualized and displayed. The depth information 120C is expressed by converting the distance to the object corresponding to each pixel into the density of each pixel. Here, an image taken from the ceiling downwards (image taken by camera unit 10 (infrared image)) is used as an example of captured image 110C. However, it is not limited to this, and captured image 110C may be an image of the target space such as a living room (especially the target space including the bed placement) taken from various angles.

[0134] In steps S72 to S74, the bed area 323 and other elements are set using an adjustment image (captured image) 110C of the target space. Specifically, the bed area 323 and other elements are set based on a predetermined point 315 (see Figure 20) within the captured image 110C.

[0135] Specifically, in step S72, the controller 31 first receives a user's operation to specify a predetermined point 315 (bed position specification operation). More specifically, on the display screen showing the captured image 110C (the display screen displayed on the display unit 35b, etc.), the user moves the bed position specification cursor 311 to a position near the center of the bed 92 in the captured image 110C using a mouse and double-clicks. In short, the user specifies the bed position by pointing (indicating) using the cursor 311. In response to this operation, the controller 31 acquires the cursor position (more specifically, the center position of the circular cursor 311) within the display screen (within the captured image 110C) as the predetermined point 315. In this way, the predetermined point 315 is specified in accordance with the user's position specification operation within the captured image 110C. Note that the predetermined point 315 is not limited to a position near the center of the bed, but may be, for example, another position on the upper surface of the bed (a position other than near the center).

[0136] The controller 31 then acquires the three-dimensional position (X, Y, Z) of the object surface position (position on the bed surface) corresponding to the two-dimensional position of a predetermined point 315. More specifically, the three-dimensional position (X, Y, Z) of a point on the bed surface 92 in coordinate system Σ3 is acquired. In other words, the position information of the predetermined point 315 in camera coordinate system Σ1 (planar position in the captured image and depth position in the normal direction of the captured image) is converted into coordinate values ​​(X, Y, Z) in coordinate system Σ3 and acquired. Coordinate system Σ3 is, for example, a three-dimensional orthogonal coordinate system based on three orthogonal axes: two orthogonal axes parallel to the horizontal plane and one axis extending vertically (height direction). The origin of coordinate system Σ3 can be set at an appropriate position (such as an appropriate position on the floor surface). The Z coordinate in coordinate system Σ3 is acquired as a value representing, for example, the height from the floor surface (for example, "56 cm"). Furthermore, the relative positions of coordinate systems Σ1 and Σ3 are assumed to be pre-adjusted. Note that the aforementioned coordinate system Σ2 is a coordinate system obtained by translating coordinate system Σ3 (in 3D space).

[0137] In the next step S73, the controller 31 extracts a region (continuous region) having a height equivalent to the height of the object surface corresponding to a predetermined point 315 (a user-specified position (e.g., near the center of the bed)) in real space (e.g., 56 cm ± 5 cm). More specifically, the controller 31 extracts a continuous region (continuous planar region) having a height (e.g., 56 cm ± 5 cm) within a tolerance range (e.g., ± 5 cm) of the height of the object surface corresponding to the predetermined point 315.

[0138] Then, in step S74, the controller 31 extracts the continuous region (more specifically, the circumscribed quadrilateral (circumscribed rectangle, etc.) of the continuous region) as the bed surface region 322 (see bottom row of Figure 20).

[0139] In the bottom row of Figure 20, a bed surface area 322 enclosed by four sides 322a, 322b, 32c, and 322d is shown. Here, the bed surface area 322 has a rectangular shape and is enclosed by four sides: the long sides 322a and 322b and the short sides 322c and 322d. The bed surface area 322 is a planar area positioned at a predetermined height from the floor. The height of the bed surface area 322 from the floor may be the height of the object surface (bed surface) corresponding to a predetermined point 315 (a position specified by the user) (for example, 56 cm), or it may be recalculated as the average value of multiple points included in the continuous area.

[0140] Furthermore, the controller 31 extracts a projected plane region as the bed plane region 321 by projecting the bed top surface region 322 onto the floor surface (or a plane corresponding to it).

[0141] In step S74, an extended bed area 325 is further set. The extended bed area 325 is a planar area obtained by extending the bed planar area 321 by a predetermined range. Specifically, the extended bed area 325 is a planar area obtained by extending the bed planar area 321 by a predetermined length (for example, 1m to 2m) in the left, right, up, and down directions (directions perpendicular to each of the four sides 322a, 322b, 32c, and 322d) when viewed from above. In other words, the extended bed area 325 is a planar area that includes both the bed planar area 321 itself and the area surrounding the bed planar area 321 (also referred to as the bed periphery area). The bed periphery area is the area of ​​the extended bed area 325 excluding the bed planar area 321. The extended bed area 325 is mainly used in the fourth embodiment.

[0142] In this way, the bed area 323 (bed top area 322 and bed plane area 321) and the extended bed area 325 are set. Here, both the bed top area 322 and the bed plane area 321 are set as the bed area 323, but this is not limited to this, and only one of the bed top area 322 or the bed plane area 321 may be set.

[0143] Subsequently, the controller 31 sets (automatically sets) the longest side 322a, which is closest to the center in the captured image 110C, as the boundary of the bed 92. In other words, the longest side 322a is set as the intersection line 215 (see Figures 21 and 5). Furthermore, one of the two endpoints of the longest side (line segment) 322a is determined to be the origin O2 of the coordinate system Σ2. For example, the one endpoint should be selected such that the Y-axis (the Y-axis along the intersection line 215 (longest side 322a)) extending from that endpoint to the other endpoint, the Z-axis extending vertically upward from that endpoint, and the X-axis extending from that endpoint to the other longest side 322b (the opposite side of longest side 322a) constitute a right-handed coordinate system Σ2. In other words, one of the endpoints should be selected such that the bed surface area 322 is located on the +X side of the intersection line 215.

[0144] In this way, coordinate system Σ2 is established. Then, based on the relationship between coordinate system Σ3 and coordinate system Σ2 (such as the direction and amount of translation between them), and the adjusted relationship between coordinate system Σ1 and coordinate system Σ3, the relationship between coordinate system Σ1 and coordinate system Σ2 is determined.

[0145] Here, it is possible to manually set (on the screen) the endpoint positions of each of the four sides surrounding the bed surface area 322. However, such a setting method is cumbersome because it requires specifying the positions of four points. In contrast, by using the method described above, it is only necessary to specify one point near the center of the bed, making the setting operation relatively easy. Furthermore, since an appropriate line segment is automatically set as the intersection line 215 from among the four sides, the setting operation is made even easier.

[0146] Furthermore, if such automatic settings are not necessarily accurate, further adjustment processing (manual settings) based on user operation may be performed after the automatic settings described above have been executed. For example, the position of each side may be finely adjusted by moving the endpoint positions of each of the four sides surrounding the bed surface area 322 using mouse operations or the like as needed. Also, if the opposite long side 322b (or 322a) to the long side 322a (or 322b) that was automatically set in response to the operation of specifying a predetermined point 315 is suitable as the intersection line 215, then an appropriate long side 322b, etc. may be specified in response to user operation. Specifically, in the mode for changing (specifying) the "bed boundary", an appropriate long side (322b, etc.) may be updated (specified) as the intersection line 215 by mouse click operations or the like.

[0147] <4. Fourth Embodiment> In the fourth embodiment, a method for detecting behavioral events of a target person will be described based on the positional relationship between the bed area 323 (particularly the bed plane area 321) and multiple specific body parts Bi of the target person (particularly the planar positional relationship (positional relationship within the projection plane (floor surface, etc.))). The setting process for the bed area 323, etc. (see Figure 18) can be performed in the same manner as in the third embodiment.

[0148] In this fourth embodiment, two learning models 410 are used.

[0149] One learning model 410 (also referred to as 410A) is a learning model that primarily detects human behavior-related events (bed-related behavior-related events) near the bed, and is the learning model (learning model for detecting bed-related events) described in each of the embodiments above. Learning model 410A is machine-learned to output behavior-related events concerning a person when it receives information 310 indicating the positional relationship between a predetermined boundary (specifically the bed boundary) and multiple specific body parts Bi of a person. More specifically, the correspondence between the coordinate values ​​Xi,Zi of multiple specific body parts Bi and behavior-related events is appropriately learned by learning model 410A.

[0150] Another learning model, 410 (also referred to as 410B), is a learning model that primarily detects behavior-related events of a person at locations away from the bed (behavior-related events unrelated to the bed) (a learning model for detecting bed-unrelated events). Learning model 410B is similar to learning model 410A. However, learning model 410B is machine-trained to output behavior-related events related to a person when it is input with information indicating the height of multiple specific body parts Bi of the person. More specifically, the correspondence between the coordinate values ​​Zi of multiple specific body parts Bi and behavior-related events is appropriately learned by learning model 410B.

[0151] Here, instead of a single learning model, multiple (specifically two) learning models are used to accurately detect a variety of behavior-related events. Specifically, learning model 410A, which excels at detecting behavior-related events of people mainly near the bed, and learning model 410B, which excels at detecting behavior-related events of people mainly at locations away from the bed, are used.

[0152] Specifically, the learning model 410A is trained using machine learning to detect various behavior-related events as described in the first embodiment, etc. For example, actions related to the bed, such as "normal supine position (on the bed)", "sitting up (upper body) on the bed", "boundary position (e.g., sitting on the edge of the bed)", and "sliding off (the bed)", can be detected.

[0153] On the other hand, the learning model 410B is machine-trained to detect two behavioral events: "standing" and "lying down." However, it is not limited to these two; the learning model 410B may also detect behavioral events including "falling" and / or "sitting" in addition to "standing" and "lying down."

[0154] These two learning models, 410A (420A) and 410B (420B), are also used during inference. Specifically, based on images of the same subject taken at the same time, the inference results (output results) from the two learning models 410A and 410B are obtained, respectively.

[0155] In this fourth embodiment, the positional relationship between the bed area 323 and the target person (specific body part Bi) is used when selecting one of the two inference results. Specifically, one of the two inference results is selected depending on the degree of separation of the target person from the bed area 323.

[0156] In detail, if it is determined that the target person is inside the extended bed area 325 (a planar area obtained by extending the bed planar area 321 by a predetermined range), the target person's behavior-related events are detected based on the inference result (output result) of the learning model 410A among the two inference results. That is, the inference result of the learning model 410A (trained model 420A) is selected as the final inference result. Conversely, if it is determined that the target person is outside the extended bed area 325, the target person's behavior-related events are detected based on the inference result (output result) of the learning model 410B (trained model 420B) among the two inference results. That is, the inference result of the learning model 410B is selected as the final inference result.

[0157] Figure 22 shows the situation in the living room where the subject has slid down beside bed 92. The upper part of Figure 22 shows the living room viewed from vertically above.

[0158] In the situation shown in the upper part of Figure 22, the learning model 410A provides the inference result "(whole body sliding off the bed)", and the learning model 410B provides the inference result "lying down". In this case, based on the captured image 110, etc., it is determined that the subject person is within the extended bed area 325 (more specifically, outside the bed planar area 321 and inside the extended bed area 325), and based on this determination, the inference result from the learning model 410A, "whole body sliding off the bed", is selected (determined) as the final inference result. Whether or not the subject person is within a predetermined planar area (for example, the extended bed area 325) can be determined based on whether or not the average position of multiple specific body parts Bi of the subject person is within that predetermined planar area. Alternatively, it may be determined based on whether or not a predetermined number or more (e.g., a majority) of the multiple specific body parts Bi of the subject person are within the predetermined planar area.

[0159] Furthermore, in situations such as those shown in the upper part of Figure 23, the learning model 410A may infer that the person is "sliding off the bed" (or "normal supine position"), while the learning model 410B may infer that the person is "lying down". Figure 23 shows a situation (abnormal state) in which the subject is lying down (due to a fall, etc.) away from the bed 92 in the living room, and the upper part of Figure 23 shows the living room as viewed from vertically above.

[0160] In the case of Figure 23, based on the captured image 110, etc., it is determined that the subject person (more specifically, a number of their specific body parts Bi or the average position of those specific body parts Bi) is outside the bed area 323 (particularly outside the extended bed area 325). Based on this determination, the inference result "lying down" from the learning model 410B is selected (determined) as the final inference result.

[0161] Furthermore, Figure 24 shows the subject sleeping in bed 92 in a normal position within the room, and the upper part of Figure 24 shows the room viewed from vertically above.

[0162] For example, in the situation shown in the upper part of Figure 24, the learning model 410A provides the inference result "normal supine position," and the learning model 410B provides the inference result "lying down." In this case, based on the captured image 110, etc., it is determined that the subject person (more specifically, many of their specific body parts Bi) is within the bed area 323 (and consequently within the extended bed area 325), and based on this determination, the inference result "normal supine position" from the learning model 410A is selected (determined) as the final inference result.

[0163] As described above, in the fourth embodiment, a learning model 410A suitable for detecting behavior-related events associated with beds and a learning model 410B suitable for detecting behavior-related events not associated with beds are generated separately. Then, the two learning models 410 are used to detect behavior-related events of the target person. With this, it is possible to obtain highly accurate detection results by using learning models that are optimized (specialized) for each situation (for each situation where there is a bed or not). In addition, in rooms where there is no bed (such as a rehabilitation room), only learning model 410B may be used (learning model 410A may not be used).

[0164] Furthermore, in the fourth embodiment, the bed plane region 321 and the like are used when selecting an appropriate output result from the output results of the two learning models 410. Specifically, the action-related events of the target person are detected based on the positional relationship (positional relationship within the projection plane (floor, etc.) (planar positional relationship)) between the bed region 323 (especially the bed plane region 321) and multiple specific body parts Bi of the target person. Specifically, depending on whether the target person is inside or outside the extended bed region 325, the output result (inference result of action-related events) from the corresponding learning model 410 (from the corresponding learning models 410A and 410B) is selected (determined) as the final inference result. This makes it possible to appropriately select an appropriate inference result from the inference results of the two learning models 410 according to the location of the target person.

[0165] In the fourth embodiment, the bed plane region 321 is used when selecting the appropriate output result from the output results of the two learning models 410. However, the invention is not limited to this, and the bed plane region 321 may be used during the learning and inference of a single learning model 410A. For example, in the first embodiment (or the fourth embodiment, etc.), information regarding whether or not the target person is located within the bed plane region 321 may be added as input to the learning model 410A. More specifically, not only the height information of the target person's body part Bi but also the presence or absence information of the target person within the bed plane region 321 may be used to perform learning and inference on the learning model 410A.

[0166] <5. Variations, etc.> The embodiments of this invention have been described above, but this invention is not limited to those described above.

[0167] For example, the specific body part of a person is not limited to the specific body parts mentioned above, but may also be the person's head, neck, collarbone, etc.

[0168] Furthermore, in each of the above embodiments, it is assumed that a person moves from outside the bed to on the bed (or from the bed to outside the bed) only from one of the two sides 92b of the bed 92. In nursing care beds, the other side 92c (see Figure 15) is often equipped with a fall prevention rail, and the above-described embodiment (an embodiment that considers the relative positional relationship with one of the boundary surfaces 210) can accommodate many situations.

[0169] However, the present invention is not limited thereto. For example, the above concept may be applied in situations where a person moves from outside the bed to on the bed (or from the bed to outside the bed) from either side of the bed. In this case, it is preferable to consider not only the boundary surface 210 on one side but also the positional relationship between each specific part Bi and the opposite boundary surface 220 (see Figure 15). For example, not only the values ​​Xi and Zi of each specific part Bi but also the value Wi of each specific part Bi may be used as input to the learning model 410. Here, the value Wi is the signed normal distance W (signed shortest distance (difference position) W with respect to the boundary surface 220) of each specific part Bi. The boundary surface 220 is a vertical plane that includes the other side 92c of the bed 92 (see Figure 15). The sign should be defined as positive (+) at a horizontal position on the bed side (inside) (left side in Figure 15) of the boundary surface 220, and negative (-) at a horizontal position on the bed side (right side in Figure 15) of the boundary surface 220.

[0170] Furthermore, while several application examples of the setting process for the bed area 323 (see Figure 18, etc.) have been described in the third and fourth embodiments above, the setting process for the bed area 323 (setting process based on specifying a predetermined point, etc.) can be applied to a variety of other uses. With such a setting process, only the operation of specifying one point on the bed is required, making the setting operation significantly easier compared to cases where the setting process involves specifying all four sides of the bed area 323. Therefore, it is not necessary for specialized workers of the equipment manufacturer to perform the setting operation, and it is possible for general users (caregivers, etc.) to perform the setting operation. [Explanation of symbols]

[0171] 1. Detection System 10 Detection device 20 Camera Units 30 processing units 70,80 Terminal devices 90 Room 91 beds 92 beds 210,220 Boundary surface (bed boundary surface) 230 Reference horizontal plane 610 Human body (three-dimensional object) 620 Reference plane Bi specific part Ei Event Hi: Projection from the reference plane

Claims

1. A control unit that detects behavioral events of the target person. Equipped with, The control unit detects the behavior-related events of the target person based on the positional relationship between a predetermined boundary and multiple specific body parts of the target person, The predetermined boundary has a first boundary surface which is a horizontal plane including the top surface of the bed, and a second boundary surface which is a vertical plane including the side surface of the bed. The control unit, In detecting the behavior-related events of the target person based on the positional relationship between the first boundary surface and multiple specific parts of the target person, and the positional relationship between the second boundary surface and multiple specific parts of the target person, the system accepts a user operation to specify a predetermined point indicating the position of the bed in an adjustment image taken of the target space. In response to the user operation, the three-dimensional position of the object surface corresponding to the two-dimensional position of the predetermined point is obtained, and a plane having a height equal to the height of the object surface is set as the first boundary surface having the height of the bed surface, and, A detection device characterized in that, in response to the user operation, it acquires a continuous region having a height equal to the height of the object surface and a rectangular region in a top view as a bed region, and sets a vertical plane containing the long side of the bed region that is closest to the center in the adjustment image as the second boundary surface.

2. The control unit selectively uses the first learning model and the second learning model to detect the behavior-related events relating to the target person, The first learning model is a machine learning model that, upon input of information indicating the positional relationship between a predetermined boundary and multiple specific body parts of a person, outputs the behavior-related events concerning that person. The second learning model described above is a different learning model from the first learning model described above. The detection device according to claim 1, characterized in that the control unit selectively uses the first learning model and the second learning model based on the positional relationship between the target person and the bed area to detect the behavior-related events relating to the target person.

3. The control unit is If the target person is present in the bed area, the first learning model is used to detect the behavior-related events of the target person based on the positional relationship between the predetermined boundary and multiple specific body parts of the target person. The detection device according to claim 2, characterized in that, when the target person is located at a position horizontally at a predetermined distance or more from the bed area, the second learning model is used to detect the behavior-related events of the target person.

4. The detection device according to claim 2 or 3, characterized in that the second learning model is a learning model that has been trained to output the behavior-related events concerning a person when it is input information indicating the height of multiple specific parts of a person.

5. A program for causing a computer to perform a process to detect behavior-related events of a target person based on the positional relationship between a predetermined boundary and a plurality of specific body parts of the target person, The predetermined boundary has a first boundary surface which is a horizontal plane including the top surface of the bed, and a second boundary surface which is a vertical plane including the side surface of the bed. The aforementioned program, a) In detecting the behavior-related events of the target person based on the positional relationship between the first boundary surface and a plurality of specific parts of the target person, and the positional relationship between the second boundary surface and a plurality of specific parts of the target person, the step of receiving a user operation to specify a predetermined point indicating the position of the bed in an adjustment image captured of the target space, b) In response to the user operation, the three-dimensional position of the object surface corresponding to the two-dimensional position of the predetermined point is obtained, and a plane having a height equivalent to the height of the object surface is set as the first boundary surface having the height of the bed surface, c) In response to the user operation, acquire a continuous region having a height equivalent to the height of the object surface and a rectangular region in a top view as the bed region, and set the vertical plane containing the long side of the bed region that is closest to the center in the adjustment image as the second boundary surface, A program that causes a computer to execute something.

6. A control unit that detects behavior-related events of a target person by selectively using a first learning model and a second learning model, Equipped with, The first learning model is a machine learning model that takes information indicating the positional relationship between a predetermined boundary and multiple specific body parts of a person as input and outputs behavior-related events concerning that person. The predetermined boundary has a first boundary surface which is a horizontal plane including the top surface of the bed, and a second boundary surface which is a vertical plane including the side surface of the bed. The second learning model described above is a different learning model from the first learning model described above. The control unit is characterized by selectively using the first learning model and the second learning model based on the positional relationship between the target person and the bed area to detect the behavior-related events concerning the target person.

7. The control unit is If the target person is present in the bed area, the first learning model is used to detect the behavior-related events of the target person based on the positional relationship between the predetermined boundary and multiple specific body parts of the target person. The detection device according to claim 6, characterized in that, when the target person is located at a position horizontally at a predetermined distance or more from the bed area, the second learning model is used to detect the behavior-related events of the target person.

8. The detection device according to claim 6 or 7, characterized in that the second learning model is a learning model that has been trained to output the behavior-related events relating to a person when it is input information indicating the height of multiple specific parts of a person.

9. A computer, a) A step of detecting behavior-related events of a target person by selectively using the first learning model and the second learning model, A program to execute, The first learning model is a machine learning model that takes information indicating the positional relationship between a predetermined boundary and multiple specific body parts of a person as input and outputs behavior-related events concerning that person. The predetermined boundary has a first boundary surface which is a horizontal plane including the top surface of the bed, and a second boundary surface which is a vertical plane including the side surface of the bed. The second learning model described above is a different learning model from the first learning model described above. In step a), the program is characterized in that it selectively uses the first learning model and the second learning model based on the positional relationship between the target person and the bed area to detect the behavior-related events concerning the target person.

Citation Information

Patent Citations

  • Object detection device, object detection method and program

    JP2014035302A

  • Information processing device, information processing method, and program

    JP2014174627A

  • Operation recognition device

    JP2017041079A

  • Behavior detection device, method and program, and monitored person monitoring device

    JP2017168105A

  • Watch support system and control method thereof

    JP2018147089A