Information processing device, information processing program, machine learning device, and machine learning program

JP7916972B2Active Publication Date: 2026-09-08KONICA MINOLTA INC
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2024514200
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-04-08
Filing Date
2023-03-13
Publication Date
2026-09-08
Estimated Expiration
2043-03-13

AI Technical Summary

Benefits of technology

【0027】 本発明においては、検出対象の複数フレーム分のキーポイント検出結果を取得し、当該複数フレーム分のキーポイント検出結果を使用して、キーポイント検出結果の未検出のキーポイントを補完する。したがって、検出対象の撮影時に検出対象の全体または一部のキーポイントが他の物体等の陰に隠れて検出できない場合でも、未検出のキーポイントを補完できる。これにより、例えば介護の現場において、検出対象としてのケア対象者を撮影中にケア対象者の全体または一部が他の人や物体の陰に隠れて検出できない場合でも、補完されたキーポイント検出結果を使用してケア対象者の姿勢や行動を推定できる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007916972000001
    Figure 0007916972000001
  • Figure 0007916972000002
    Figure 0007916972000002
  • Figure 0007916972000003
    Figure 0007916972000003
Patent Text Reader

Abstract

Provided are an information processing device, an information processing program, a machine-learning device, and a machine-learning program which can estimate the posture or behavior of a subject to be provided care even when a portion or the entirety of the subject to be provided care is hidden by the shadow of another person or an object and cannot be detected while the subject to be provided care is imaged. The information processing device 200 includes a key point acquisition unit 112 and a complementary unit 113. The key point acquisition unit 112 acquires a key point detection result for a plurality of frames of the subject to be detected. The complementary unit 113 uses the key point detection result for the plurality of frames and complements undetected key points of the key point detection result.
Need to check novelty before this filing date? Find Prior Art

Description

[[Technical Field]]

[0001] The present invention relates to an information processing apparatus, an information processing program, a machine learning apparatus, and a machine learning program. [[Background Art]]

[0002] In Japan, due to the improvement of living standards accompanying post-war high economic growth, the improvement of sanitary environments, and the improvement of medical standards, the increase in life expectancy has become remarkable. For this reason, combined with the declining birth rate, Japan has become an aging society with a high aging rate. In such an aging society, it is expected that the number of care-requiring persons (hereinafter referred to as "care subjects") who need nursing care and other responses due to illness, injury, aging, and the like will increase. In facilities such as hospitals and welfare facilities for the elderly (hereinafter simply referred to as "facilities"), care and other responses to care subjects are provided by caregivers, nurses, and the like (hereinafter referred to as "care staff").

[0003] Further, along with the increase in the number of care subjects, the burden on care staff is increasing, and technological development to reduce the burden is being promoted. For example, a technique is known in which a camera (e.g., a near-infrared camera, etc.) for photographing the state of a care subject is installed in the care subject's room, and the posture (standing, lying, etc.) and behavior (getting up, leaving bed, etc.) of the care subject are estimated from the captured image (e.g., Patent Document 1).

[0004] However, when photographing a care subject with a camera, depending on the position of the care subject in the room, the care subject may be hidden behind other people in the room such as care staff or objects such as installed beds, chairs, etc., which may make it impossible to detect all or part of the care subject (occlusion).

[0005] In relation to this, a technique for recognizing an object by complementing the occluded part from an image captured in a state where a part of the object is occluded is disclosed (Patent Document 2). Further, a technique for interpolating a missing part in an image is disclosed (Patent Document 3). [[Prior Art Documents]] [Patent Documents]

[0006] [Patent Document 1] Japanese Patent Publication No. 2020-86819 [Patent Document 2] Japanese Patent Publication No. 2020-135551 [Patent Document 3] International Publication No. 2019 / 186833 [Overview of the Initiative] [Problems that the invention aims to solve]

[0007] However, in the technology described in Patent Document 2, it is not possible to recognize an object from an image in which the entire object is obscured from the camera's field of view. Similarly, in the technology described in Patent Document 3, it is not possible to interpolate missing parts from an image in which the entire person is missing.

[0008] The present invention was made to solve the above-mentioned problems. That is, the main object of the present invention is to provide an information processing device, an information processing program, a machine learning device, and a machine learning program that can estimate the posture and behavior of a care recipient even when the care recipient is hidden in the shadow of another person or object and cannot be detected while the care recipient is being photographed. [Means for solving the problem]

[0009] The above-mentioned problems of the present invention are solved by the following means.

[0010] (1) A keypoint acquisition unit that acquires keypoint detection results for multiple frames to be detected, and a completion unit that uses the keypoint detection results for multiple frames to complete the keypoints that have not been detected. The keypoint acquisition unit changes the number of frames of the keypoint detection results to be acquired, according to the processing method of the interpolation unit. Information processing device. (2) An information processing device comprising: a keypoint acquisition unit that acquires keypoint detection results for multiple frames to be detected; and a completion unit that uses the keypoint detection results for multiple frames to complete the keypoints that have not been detected in the keypoint detection results, wherein the completion unit inputs the keypoint detection results for multiple frames, including missing frames in which at least some keypoints are missing, and the ground truth data corresponding to the keypoint detection results into a learning model, and uses a trained model that has been machine-learned with the ground truth data as its target to complete the keypoints that have not been detected in the keypoint detection results.

[0011] ( 3The keypoint acquisition unit acquires keypoint detection results in a video containing multiple temporally consecutive frames of images, The interpolation unit uses the keypoint detection results in the video to interpolate keypoints that were not detected in the video, as described in (1) above. or (2) The information processing device described above.

[0012] ( 4 The information processing apparatus according to (1) or (2) above, wherein the key point acquisition unit acquires estimated key point detection results for multiple frames of still images captured by at least one imaging device.

[0013] ( 5 ) The key point detection result consists of two-dimensional key points, as described above ( 4 Information processing device as described above.

[0014] ( 6 The keypoint acquisition unit further acquires a rectangle containing the keypoint in addition to the keypoint detection result, as described above. 5 Information processing device as described above.

[0015] ( 7 The information processing device according to (1) or (2) above, wherein the key point detection result is the detection result of multiple joint points, or the detection result of skeletal information including joint points and nodes connecting the joint points.

[0018] ( 8 The learning model is a generative model that extracts features from the keypoint detection results and reconstructs undetected keypoints based on the extracted features. 2 Information processing device as described above.

[0019] ( 9 The learning model is a transformer model that takes the keypoint detection results for multiple frames as an input sequence and the reconstructed keypoint detection results for multiple frames as the inference result. 2The information processing apparatus according to

[0020] ( 10 )The information processing apparatus according to (1) or (2) above, further comprising an action estimation unit that estimates an action using the keypoints complemented by the complementing unit.

[0021] ( 11 )The information processing apparatus according to (1) or (2) above, further comprising a people counting estimation unit that estimates the number of people using the keypoints complemented by the complementing unit.

[0022] ( 12 )The information processing apparatus according to (1) or (2) above, further comprising a posture estimation unit that estimates a posture using the keypoints complemented by the complementing unit.

[0023] ( 13 )The information processing apparatus according to (1) or (2) above, wherein the keypoint acquisition unit detects keypoints in images of a plurality of frames of the detection target.

[0024] ( 14 )A process comprising: step (a) of acquiring keypoint detection results for a plurality of frames of a detection target; and step (b) of complementing undetected keypoints in the keypoint detection results using the keypoint detection results for the plurality of frames In the above procedure (a), a process is performed to change the number of frames of the keypoint detection result to be acquired, depending on the processing method of the above procedure (b). An information processing program for causing a computer to execute . (15) An information processing program that causes a computer to perform a process including (a) a step of obtaining keypoint detection results for multiple frames to be detected, and (b) a step of using the keypoint detection results for multiple frames to complete the keypoints that have not been detected in the keypoint detection results, wherein in step (b), the program inputs the keypoint detection results for multiple frames, including missing frames in which at least some keypoints are missing, and the ground truth data corresponding to the keypoint detection results into a learning model, and uses a trained model that has been machine-trained with the ground truth data as its target to complete the process of completing the keypoints that have not been detected in the keypoint detection results.

[0025] (16) A machine learning apparatus comprising: a reception unit that receives keypoint detection results for a plurality of frames including missing frames in which at least some keypoints are missing, and correct answer data corresponding to the keypoint detection results; and a learning unit that inputs the keypoint detection results for the plurality of frames and the correct answer data into a learning model, and generates a trained model by causing the learning model to perform machine learning with the correct answer data as a target.

[0026] (17) A machine learning program that causes a computer to perform the following steps: (a) receiving keypoint detection results for multiple frames, including missing frames in which at least some keypoints are missing, and ground truth data corresponding to the keypoint detection results; and (b) inputting the keypoint detection results for multiple frames and the ground truth data into a learning model, and generating a trained model by having the learning model machine learn using the ground truth data as the target. [Effects of the Invention]

[0027] In this invention, keypoint detection results for multiple frames of the target to be detected are acquired, and the undetected keypoints in the keypoint detection results are used to complete the process. Therefore, even if all or part of the keypoints of the target to be detected are hidden by other objects or other objects during the time of shooting, the undetected keypoints can be completed. This allows, for example, in a care setting, even if all or part of the care recipient is hidden by other people or objects during shooting and cannot be detected, the posture and behavior of the care recipient can be estimated using the completed keypoint detection results. [Brief explanation of the drawing]

[0028] [Figure 1] This figure illustrates a schematic configuration of an information processing system according to one embodiment of the present invention. [Figure 2] Figure 1 is a block diagram illustrating the schematic configuration of the imaging device shown. [Figure 3] Figure 1 is a block diagram illustrating the schematic configuration of the server shown. [Figure 4] Figure 1 is a block diagram illustrating the schematic configuration of a mobile device. [Figure 5] Figure 1 is a functional block diagram illustrating the functions of the control unit when the server functions as an information processing device. [Figure 6]Figure 5 is an example of an image containing multiple frames (A) to (F) acquired by the image acquisition unit shown. [Figure 7] This is a schematic diagram illustrating the keypoint detection results obtained by identifying key points of the care recipient or care staff from an image. [Figure 8] Figure 1 is a flowchart illustrating the processing procedure for the information processing method in the server (control unit). [Figure 9] Figure 6 is a schematic diagram illustrating the keypoint detection results of an image containing multiple frames. [Figure 10] This figure illustrates the estimated posture, number of people, and behavior based on the supplemented keypoint detection results. [Figure 11] This is a schematic diagram illustrating the complementary keypoint detection results. [Figure 12] Figure 1 is a functional block diagram illustrating the functions of the control unit when the server functions as a machine learning device. [Figure 13] Figure 12 is a flowchart illustrating the processing steps of the learning method in the machine learning device shown. [Modes for carrying out the invention]

[0029] Hereinafter, an information processing apparatus, an information processing program, a machine learning apparatus, and a machine learning program according to an embodiment of the present invention will be described with reference to the drawings. In the drawings, the same elements are denoted by the same reference numerals, and redundant explanations are omitted. Also, the dimensional ratios in the drawings are exaggerated for illustrative purposes and may differ from the actual ratios.

[0030] <Embodiment> [Overall configuration of information processing system 10] Figure 1 is a block diagram illustrating the schematic configuration of an information processing system 10 according to one embodiment. The information processing system 10 includes, for example, a camera 100, a server 200, a communication network 300, and a mobile terminal 400. The camera 100 is connected to the server 200 via the communication network 300 so that they can communicate with each other. The mobile terminal 400 may be connected to the communication network 300 via an access point 310. The server 200 corresponds to one specific example of the information processing device of this embodiment. Note that the camera 100 may perform some or all of the functions of the server 200, which will be described later. In this case, the camera 100 can constitute the information processing device either alone or together with the server 200.

[0031] (Photography device 100) Figure 2 is a block diagram illustrating the schematic configuration of the imaging device 100 shown in Figure 1. The imaging device 100 includes a control unit 110, a communication unit 120, and a camera 130, which are interconnected by a bus 101. At least one imaging device 100 is installed, for example, on the ceiling or wall of the room of the person receiving care 510. The following example illustrates the case where one imaging device 100 is installed on the ceiling, but the number is not limited to one.

[0032] The control unit 110 consists of a CPU (Central Processing Unit) and memory such as RAM (Random Access Memory) and ROM (Read Only Memory), and performs control and calculation processing of each part of the imaging device 100 according to the information processing program.

[0033] The control unit 110 transmits multiple frames of images (for example, image 500 in Figure 6, described later) obtained by the camera 130 capturing a predetermined area to the server 200 or the like via the communication unit 120. The predetermined area is, for example, a three-dimensional area including the entire floor surface of the room of the person receiving care 510 (Figure 1).

[0034] The communication unit 120 includes, for example, an interface circuit (such as a LAN card) for communicating with a mobile terminal 400 or the like via a communication network 300.

[0035] Camera 130 is, for example, a wide-angle camera. Camera 130 is installed in a position that provides an overhead view of a predetermined area, specifically, on the ceiling of the room of the person receiving care 510, and photographs the predetermined area. The person receiving care 510 is, for example, a person who requires care or nursing from care staff. Camera 130 may also be a standard camera with a narrower field of view than a wide-angle camera.

[0036] For the sake of simplicity, the following explanation assumes that camera 130 is a wide-angle camera. Images captured by camera 130 may include the person being cared for 510, care staff, and objects. Objects may include, for example, a bed 610, a wheelchair 620, etc. Images captured by camera 130 may include still images and videos.

[0037] Camera 130 is, for example, a near-infrared camera that illuminates the shooting area with near-infrared light using an LED (Light Emitting Device), and can capture a predetermined area by receiving the reflected near-infrared light reflected by objects within the shooting area with a CMOS (Complementary Metal Oxide Semiconductor) sensor.

[0038] The image captured by camera 130 may be a monochrome image where the reflectivity of each pixel is determined by the near-infrared light. The imaging device 100 can capture the area as a video consisting of multiple temporally continuous captured images (frames) at a frame rate of, for example, 15 fps to 30 fps. Camera 130 may also use a visible light camera instead of a near-infrared camera, or both may be used in combination.

[0039] (Server 200) Figure 3 is a block diagram illustrating the schematic configuration of the server 200 shown in Figure 1. The server 200 includes a control unit 210, a communication unit 220, and a storage unit 230. The components of the server 200 are interconnected by a bus 201.

[0040] The basic configuration of the control unit 210 and the communication unit 220 is the same as that of the control unit 110 and the communication unit 120 of the imaging device 100, so redundant explanations will be omitted. The specific functions of the control unit 210 will be described later. The storage unit 230 is composed of, for example, RAM, ROM, SSD (Solid State Drive), etc. The SSD stores, for example, programs such as information processing programs and trained models, which will be described later.

[0041] (Mobile phone service 400) Figure 4 is a block diagram illustrating the schematic configuration of the mobile terminal 400 shown in Figure 1. The mobile terminal 400 includes a control unit 410, a wireless communication unit 420, a display unit 430, an input unit 440, and an audio input / output unit 450. Each component is interconnected by a bus 401. The mobile terminal 400 can be composed of, for example, a tablet computer, a smartphone, or a mobile phone. The control unit 410 has a basic configuration including a CPU, RAM, ROM, etc., similar to the configuration of the control unit 110 of the imaging device 100.

[0042] The wireless communication unit 420 has the function of performing wireless communication according to standards such as Wi-Fi and Bluetooth (registered trademark), and communicates wirelessly with each device via the access point 310 or directly. The wireless communication unit 420 receives event notifications from the server 200.

[0043] The display unit 430 and the input unit 440 are touch panels, and the display surface of the display unit 430, which is made of liquid crystal or the like, is provided with a touch sensor that serves as the input unit 440. The display unit 430 displays the actions of the person being cared for 510 received from the server 200. The actions of the person being cared for 510 may also be displayed by displaying the event notification described above.

[0044] The audio input / output unit 450 includes, for example, a speaker and a microphone. This audio input / output unit 450 enables voice communication between care staff and other mobile terminals 400 via the wireless communication unit 420.

[0045] [Server 200 Features] Next, the functions of the server 200, specifically the control unit 210, will be described. Figure 5 is a functional block diagram illustrating the functions of the control unit 210 when the server 200 functions as an information processing device. The control unit 210 functions, for example, as an image acquisition unit 211, a keypoint detection unit 212, an interpolation unit 213, a posture estimation unit 214, a person estimation unit 215, an action estimation unit 216, and an output unit 217.

[0046] The image acquisition unit 211 acquires images for multiple frames in which a predetermined area has been captured. Figure 6 is an example of an image 500 that includes multiple frames (A) to (F) acquired by the image acquisition unit 211. The image 500 may be, for example, a video captured sequentially (at times t1, t2, t3, t4, t5, t6) in a living room by the camera 100.

[0047] The diagram illustrates, for example, a caregiver 520 transferring a care recipient 510 to a wheelchair 530 in a nursing care setting. More specifically, in frames (A) and (B), the caregiver 520 moves closer to the care recipient 510; in (C), the caregiver 520 explains to the care recipient 510 how to transfer to the wheelchair 530; in (D) and (E), the transfer takes place; and in (F), the transfer to the wheelchair 530 is completed.

[0048] The image acquisition unit 211 acquires the image 500 from the imaging device 100, for example, by receiving it via the communication unit 220. If the image 500 captured by the imaging device 100 is already stored in the storage unit 230 or the like, the image acquisition unit 211 may acquire the image 500 by reading it from the storage unit 230 or the like. The image 500 captured by the imaging device 100 may also be stored in an external storage device or the like. Furthermore, the image 500 acquired by the image acquisition unit 211 may be, for example, an image that has undergone batch processing, and the image acquisition unit 211 may acquire the image 500 offline.

[0049] The keypoint detection unit 212 detects keypoints for the care recipient 510 and care staff 520 from the image 500, which includes multiple frames of images acquired by the image acquisition unit 211, and outputs them as keypoint detection results. Alternatively, the keypoint detection unit 212 can also receive keypoint detection results for multiple frames of the target from outside the server 200. The keypoint detection unit 212 functions as a keypoint acquisition unit.

[0050] Furthermore, the number of frames in the keypoint detection results acquired by the keypoint detection unit 212 is not fixed, but can be changed by the interpolation method of the interpolation unit 213. For example, the keypoint detection unit 212 may be configured to change the number of frames according to the processing method of the interpolation unit 213, which will be described later. Alternatively, the user may be able to arbitrarily set the number of frames.

[0051] For example, as will be described later, if the interpolation unit 213 performs interpolation using a machine learning-trained model, the number of frames may be set according to the configuration of the trained model. For example, if the trained model is an auto-encoder (AE) or variational auto-encoder (VAE) that reconstructs keypoints, the number of frames required for feature extraction may be set. Also, if the trained model is a transformer model, the number of frames required for the input sequence may be set. Furthermore, if the interpolation unit 213 performs interpolation using a method other than machine learning, the number of frames appropriate for that interpolation process will be set.

[0052] As shown in Figure 7, keypoints 700 may include, for example, the two-dimensional or three-dimensional coordinates of feature points (joint points) 710 of the eyes, nose, neck, shoulders, elbows, wrists, hips, knees, ankles, etc. of the care recipient 510 and care staff 520. The keypoint detection result may be the detection result of multiple joint points 710, or the detection result of skeletal information including multiple joint points 710 and nodes 720 connecting the joint points 710.

[0053] The keypoint detection unit 212 can detect the keypoints 700 of each of the care recipient 510 and the care staff 520 using a known method such as OpenPose (https: / / arxiv.org / abs / 1812.08008). OpenPose is software that can simultaneously detect the keypoints of multiple people.

[0054] Furthermore, the keypoint detection unit 212 may be configured to detect keypoints 700 by performing object detection (person detection) on the image 500 and individually estimating the posture of each region of the detected care recipient 510 and care staff 520.

[0055] For example, the keypoint detection unit 212 estimates the person rectangles 730 of the care recipient 510 and the care staff 520 from the image 500, and estimates keypoints 700 for each of the estimated person rectangles 730, thereby obtaining keypoints 700 and person rectangles 730.

[0056] The person rectangle 730 is the region that contains the key point 700 of the care recipient 510 or care staff 520 in the image 500, and can reflect the position, size, and posture of the care recipient 510 and care staff 520, respectively. For example, if the key point 700 is 2D data and does not contain information about depth (height), displaying the person rectangle that contains the key point 700 will change how the person (care recipient 510 and care staff 520) appears depending on their position in the depth direction. That is, the person closer to the viewer will appear larger, while the person further away will appear smaller. In this way, by having a person rectangle 730 that represents the size of the person in addition to the key point 700, it is possible to respond to changes in apparent size depending on the position in the depth direction. On the other hand, if the key point 700 is 3D data and contains information about depth (height), the appearance is not affected by the position in the depth direction, so there is no need to display the person rectangle.

[0057] The human rectangle 730 can be estimated, for example, using a pre-trained model of a neural network that has been trained to estimate the human rectangle from an image. Examples of pre-trained models for estimating the human rectangle from an image include R-CNN, Fast R-CNN, Faster R-CNN (https: / / arxiv.org / abs / 1506.01497), YOLO (https: / / arxiv.org / abs / 1506.02640), and SSD (https: / / arxiv.org / abs / 1512.02325).

[0058] Keypoint 700 is estimated using a pre-trained neural network model that has been trained to estimate keypoints from a person rectangle. Examples of pre-trained models for estimating keypoints from a person rectangle include Deep Pose (https: / / arxiv.org / abs / 1312.4659) and ResNet (https: / / arxiv.org / abs / 1512.03385).

[0059] Alternatively, the keypoint detection unit 212 may be configured to capture a predetermined area without a person using the shooting device 100, store it as a background image, and calculate the person rectangle 730 based on the difference between the captured image of the predetermined area with a person and the background image (background difference method). Or, the keypoint detection unit 212 may be configured to calculate the person rectangle 730 based on the difference between the captured image and the average of past captured images (time difference method).

[0060] Thus, the keypoint acquisition unit 112 acquires keypoint detection results by detecting them from the image 500 or by receiving them from an external source. The keypoint acquisition unit 112 can also acquire a person rectangle 730 containing the keypoint 700 by estimating it or by receiving it from an external source.

[0061] The interpolation unit 213 uses the keypoint detection results for multiple frames of the target to interpolate any undetected (i.e., missing) keypoints in the keypoint detection results and transmits the interpolated keypoint detection results to the output unit 217 as the interpolation result. For example, the interpolation unit 213 uses the keypoint detection results for an image 500 containing images from multiple frames to interpolate any undetected keypoints in the said keypoint detection results.

[0062] More specifically, the interpolation unit 213 inputs the keypoint detection results for multiple frames to be detected into a trained model and uses the trained model to interpolate the undetected keypoints in the keypoint detection results. The trained model is generated by inputting the keypoint detection results for multiple frames, including missing frames in which at least some keypoints are missing, and the ground truth data corresponding to those keypoint detection results into the training model and performing machine learning with the ground truth data as the target.

[0063] The learning model can be a generative model that extracts features from keypoint detection results across multiple frames and reconstructs undetected keypoints based on the extracted features (i.e., using multidimensional data of keypoint detection results across multiple frames as input). In this embodiment, the features are multidimensional data (tensors) containing information about the pose and position of a person across multiple frames. For example, the generative model can be implemented using an AE or VAE that takes keypoints of a person across multiple frames as input and reconstructs keypoints in undetected frames between frames.

[0064] Furthermore, the interpolation unit 213 may be configured to interpolate undetected keypoints using a trained model learned by the transformer. In the transformer, the learning model may be a transformer model that takes keypoint detection results for multiple frames as an input sequence and infers the reconstructed keypoint detection results. For example, the transformer model is a trained model that has been pre-trained on the task of inferring undetected keypoints from keypoint detection results. In the transformer, the learning model (transformer model) is input with keypoint detection results for multiple frames, including missing frames in which at least some keypoints are missing, and the ground truth data corresponding to those keypoint detection results, and a trained model is generated by performing machine learning with the ground truth data as the target.

[0065] Furthermore, the completion method performed by the completion unit 213 is not limited to a method that uses machine learning to perform the completion process.

[0066] The posture estimation unit 214 uses the completion results from the completion unit 213 to estimate the posture of each person (care recipient 510 and / or care staff 520) in a specific image of image 500. The posture estimation results are transmitted to the output unit 217.

[0067] The person estimation unit 215 uses the interpolation results to estimate the number of people included in a specific image of image 500. The person estimation result is transmitted to the output unit 217.

[0068] The behavior estimation unit 216 uses the completion results to estimate the behavior of each person (care recipient 510 and / or care staff 520). The behavior estimation results are transmitted to the output unit 217.

[0069] The output unit 217 outputs the completion results. In addition to the completion results, the output unit 217 also outputs estimated results for posture, number of people, and behavior. Details of these estimated results will be described later.

[0070] [Processing on Server 200] Next, a specific example of the processing performed by the control unit 210, that is, the information processing method of the present invention, will be described using Figures 8 to 11. Figure 8 is a flowchart illustrating the processing procedure of the information processing method in the server shown in Figure 1. Note that if some or all of the functions shown in Figure 8 are performed by the imaging device 100, this flowchart may be executed by the control unit 110 of the imaging device 100 in accordance with the information processing program. Figure 9 is a schematic diagram illustrating the keypoint detection results of an image containing multiple frames shown in Figure 6. Figure 10 is a diagram illustrating the estimated results of posture, number of people, and actions estimated based on the interpolation results, and Figure 11 is a schematic diagram illustrating the interpolation results.

[0071] First, an image of the room is acquired (step S101). The image acquisition unit 211 acquires image 500 by receiving image data of the room from the camera 100.

[0072] Next, the keypoint detection results for multiple frames of the target are obtained (step S102). As shown in Figure 9, the keypoint detection unit 212 detects keypoints from the image 500 and obtains keypoint detection results for multiple frames. In frames (C) and (D) in Figure 9, some keypoints of the care recipient 510 are not detected because the feet or lower body of the care recipient 510 are hidden in the shadow of the care staff 520. Also, in the frame shown in Figure (E), almost the entire body of the care recipient 510 is hidden in the shadow of the care staff 520, so all keypoints of the care recipient 510 are not detected.

[0073] Next, the undetected keypoints in the keypoint detection results are supplemented (step S103). As shown in Figure 10, the supplementation unit 213 uses the keypoint detection results for multiple frames to supplement the undetected keypoints in the keypoint detection results using the trained model. For example, the supplementation unit 213 uses the keypoint detection results for the five frames (A) to (F) in Figure 6 to supplement the undetected keypoints of the care recipient 510 in frames (C) to (E) in the same figure.

[0074] Furthermore, the posture estimation unit 214 uses the completion results to estimate the posture of both the care recipient 510 and the care staff 520. For example, in frames (A) and (B), the posture of the care recipient 510 is estimated to be "sitting," and the posture of the care staff 520 is estimated to be "standing." On the other hand, in frames (C) and (D) where the key points are completed, the posture of the care recipient 510 is estimated to be "sitting," and the posture of the care staff 520 is estimated to be "standing (crouching)." Also, in frame (E) where the key points are completed, the posture of the care recipient 510 is estimated to be "sitting," and the posture of the care staff 520 is estimated to be "standing (bent forward)." Furthermore, in frame (F), the posture of the care recipient 510 is "sitting," and since only the care recipient 510 is in the room, the posture of the care staff 520 is not estimated.

[0075] Furthermore, the person estimation unit 215 uses the completion results to estimate the number of people included in the specific image 500. In frames (A) and (B), there are two people, the care recipient 510 and the care staff 520, so the estimated number of people is "2". Similarly, in frames (C) to (E) where the key points have been completed, the estimated number of people is "2". On the other hand, in frame (F), only the care recipient 510 is in the room, so the estimated number of people is "1".

[0076] Furthermore, the behavior estimation unit 216 uses the completion results to estimate the behavior of the care staff 520. For example, in frames (A) and (B), the behavior of the care staff 520 is estimated to be "nursing care". Similarly, in frames (C) to (E) where the key points have been completed, the behavior of the care staff 520 is estimated to be "nursing care". On the other hand, in frame (F), since only the person receiving care 510 is in the room, the behavior of the care staff 520 is estimated to be "non-nursing care".

[0077] Next, the completion results are output (step S104). As shown in Figure 11, the output unit 217 displays the completion results on a display, for example. In this figure, the outline of the care staff 520 is shown as a dashed line to make the completed key points of the care recipient 510 easier to see, and the key points are also shown. In addition to the completion results, the output unit 217 can also display the estimated posture results, the estimated number of people, and the estimated behavior results shown in Figure 10.

[0078] As described above, in the flowchart shown in Figure 8, the control unit 210 acquires keypoint detection results for multiple frames of the target to be detected, and uses the keypoint detection results for multiple frames to fill in the missing keypoints in the keypoint detection results. The output unit 217 outputs the completion results and estimation results based on the completion results.

[0079] (Machine learning device) Next, we will describe the machine learning device that generates the trained model shown in Figure 10. Figure 12 is a functional block diagram illustrating the functions of the control unit 210 when the server 200 shown in Figure 1 functions as a machine learning device, and Figure 13 is a flowchart illustrating the processing steps of the learning method in the machine learning device shown in Figure 12.

[0080] As shown in Figure 12, the control unit 210 functions as a receiving unit 218 and a learning unit 219. The outline of the processing procedure for the learning method by the machine learning device, and the outlines of the functions of the receiving unit 218 and the learning unit 219 are as follows.

[0081] As shown in Figure 13, first, the system receives keypoint detection results for multiple frames to be detected, along with the corresponding ground truth data (step S201). The multiple frames may include missing frames in which at least some of the keypoints from the keypoint detection results are missing. The receiving unit 218 receives training data consisting of keypoint detection results and ground truth data from outside the server 200 via the storage unit 230 or the communication unit 220. It is desirable that the training data consists of, for example, several thousand to several hundred thousand frames.

[0082] Next, a trained model is generated (step S202). The learning unit 219 inputs the keypoint detection results for multiple frames to be detected, as well as the ground truth data, into the learning model, and generates a trained model by repeatedly training the learning model with the ground truth data as the target. The generated trained model is stored in the storage unit 230. The learning model can be the generative model or the transformer model described above.

[0083] [Effects and Effects of Information Processing System 10] As described above, in facilities, the imaging device 100 is sometimes installed on the ceiling of the room of the person receiving care 510. That is, the imaging device 100 photographs a predetermined area from above the person receiving care 510. As a result, in some images captured by the imaging device 100, the person receiving care 510 may overlap with the care staff 520 or other objects, making it impossible to detect the person receiving care 510 (or the care staff 520 or other objects), i.e., they may not be detected.

[0084] For example, when the care staff member 520 is transferring the person being cared for 510 from the bed 610 to the wheelchair 620, the distance between the care staff member 520 and the person being cared for 510 is close, and the field of view of the camera 130, which has an overhead view of the room from the ceiling, may be obstructed by the care staff member 520, making it impossible to detect the person being cared for 510. As a result, the image 500 may not be able to determine that the care staff member 520 is providing care to the person being cared for 510, which could lead to an error in estimating the actions of the care staff member 520.

[0085] According to the information processing device and information processing program of this embodiment, keypoint detection results for multiple frames of the target to be detected are acquired, and the undetected keypoints in the keypoint detection results are used with the keypoint detection results for the multiple frames. In other words, while conventional joint point interpolation techniques interpolate keypoints for a single still image, the information processing device and information processing program of this embodiment are techniques that interpolate undetected keypoints for keypoint detection results detected from an image 500 consisting of multiple frames. As a result, keypoints can be interpolated even for frames where the person being photographed is in a position not visible from the camera 100, i.e., where all keypoints are undetected.

[0086] Therefore, for example, in a nursing care setting, even if the care recipient 510 is partially or entirely hidden by another person or object and cannot be detected while being photographed, the posture and behavior of the care recipient 510 can be estimated using the complemented keypoints. Although the above mainly describes the case of complementing undetected keypoints of a care recipient 510 hidden by a care staff member 520, the present invention is not limited to such cases. The present invention can also be applied to cases where undetected keypoints of a care staff member 520 hidden by a care recipient 510, or undetected keypoints of a care recipient 510 hidden by an object, are complemented. Furthermore, even outside the nursing care field, in fields such as surveillance cameras and sports, if a person is not detected when posture estimation is performed from an image, the undetected keypoints can be complemented by using this technology. This improves the accuracy of posture estimation, behavior estimation, and number estimation of the detected target.

[0087] Furthermore, in this embodiment, the keypoint detection unit 212 acquires estimated keypoint detection results for multiple frames of still images captured by at least one imaging device 100. For example, the keypoint detection unit 212 acquires estimated two-dimensional keypoint detection results for multiple frames of still images captured by one imaging device 100 over a predetermined area. The interpolation unit 213 then uses the two-dimensional keypoint detection results for multiple frames to interpolate any undetected keypoints in the keypoint detection results. Therefore, it is not necessary to measure the positions (three-dimensional) of the joint points of the care recipient 510 and care staff 520 using special devices such as motion capture. For this reason, this embodiment can be applied even in fields such as nursing care and surveillance cameras where it is difficult to attach sensors for measuring three-dimensional data to the detection target and where three-dimensional data of the detection target cannot be acquired. On the other hand, the keypoint detection unit 212 can also acquire estimated three-dimensional keypoint detection results for multiple frames of still images captured by two imaging devices 100 over a predetermined area.

[0088] Furthermore, the keypoint detection unit 212 can acquire a person rectangle 730 containing the keypoint 700, in addition to the keypoint detection results. In the case of 2D data, the apparent size of a person changes depending on their position in the image. By simultaneously acquiring the person rectangle 730 containing the keypoint 700, it is possible to reconstruct keypoint detection results for multiple frames using 2D data that takes the apparent size into account.

[0089] The configuration of the information processing system 10 described above is intended to illustrate the main features of the above-described embodiment, and is not limited to the above configuration; various modifications can be made within the scope of the claims. Furthermore, it does not preclude configurations that are generally found in information processing systems.

[0090] For example, in the above embodiment, an example was described in which the information processing system 10 includes a camera 100, a server 200, a communication network 300, and a mobile terminal 400. However, the information processing system 10 may further include a terminal for the facility's information administrator (administrator terminal). In this case, the administrator terminal may correspond to a specific example of some or all of the information processing apparatus of the present invention.

[0091] Furthermore, the means and methods for performing the various processing tasks in the information processing system 10 described above can be implemented using either dedicated hardware circuits or a programmed computer. The program may be provided, for example, on a computer-readable recording medium such as a USB memory stick or a DVD (Digital Versatile Disc)-ROM, or it may be provided online via a network such as the Internet. In this case, the program recorded on the computer-readable recording medium is usually transferred to and stored in a storage unit such as a hard disk. The program may also be provided as a standalone application software, or it may be incorporated as a function into the software of a server or other device.

[0092] This application is based on Japanese Patent Application No. 2022-064277, filed on April 8, 2022, the disclosures of which are incorporated in their entirety by reference. [Explanation of Symbols]

[0093] 10 Information processing systems, 100 imaging devices, 110 Control unit, 120 Communications Department, 130 cameras, 200 servers, 210 Control unit, 211 Image acquisition unit, 212 Keypoint detection unit, 213 Supplementary section, 214 Posture estimation unit, 215 Population Estimation Department; 216 Behavior Estimation Department, 217 Output section, 218 Reception Department, 219 Learning Department, 220 Communications Department, 230 storage section, 300 communication networks, 400 mobile devices, 410 Control unit, 420 Wireless Communication Section, 430 display section, 440 Input section, 450 Audio input / output section, 500 images, 510 people receiving care, 520 care staff, 610 beds, 620 wheelchairs, 700 key points, 710 joint points, 720 nodes, 730 people rectangle.

Claims

1. A keypoint acquisition unit that acquires keypoint detection results for multiple frames to be detected, The system includes a interpolation unit that uses the keypoint detection results for the aforementioned multiple frames to fill in any undetected keypoints in the keypoint detection results, The keypoint acquisition unit is an information processing device that changes the number of frames of the keypoint detection result to be acquired according to the processing method of the interpolation unit.

2. A keypoint acquisition unit that acquires keypoint detection results for multiple frames to be detected, The system includes a interpolation unit that uses the keypoint detection results for the aforementioned multiple frames to fill in any undetected keypoints in the keypoint detection results, The interpolation unit inputs keypoint detection results for multiple frames, including missing frames in which at least some keypoints are missing, and the corresponding ground truth data into a learning model, and uses a trained model that has been machine-trained with the ground truth data as its target to interpolate the undetected keypoints in the keypoint detection results.

3. The keypoint acquisition unit acquires keypoint detection results in a video containing multiple temporally consecutive frames of images. The information processing apparatus according to claim 1 or 2, wherein the interpolation unit interpolates undetected key points in the video using the key point detection results in the video.

4. The information processing apparatus according to claim 1 or 2, wherein the key point acquisition unit acquires key point detection results estimated for multiple frames of still images captured by at least one imaging device.

5. The information processing apparatus according to claim 4, wherein the key point detection result consists of two-dimensional key points.

6. The information processing apparatus according to claim 5, wherein the keypoint acquisition unit further acquires a rectangle containing the keypoint in addition to the keypoint detection result.

7. The information processing device according to claim 1 or 2, wherein the keypoint detection result is the detection result of a plurality of joint points, or the detection result of skeletal information including joint points and nodes connecting the joint points.

8. The information processing apparatus according to claim 2, wherein the learning model is a generative model that extracts features from the keypoint detection results and reconstructs undetected keypoints based on the extracted features.

9. The information processing apparatus according to claim 2, wherein the learning model is a transformer model that takes the keypoint detection results for multiple frames as an input sequence and the reconstructed keypoint detection results for multiple frames as an inference result.

10. The information processing apparatus according to claim 1 or 2, further comprising a behavior estimation unit that performs behavior estimation using key points complemented by the complementation unit.

11. The information processing apparatus according to claim 1 or 2, further comprising a person estimation unit that performs person estimation using key points complemented by the complementation unit.

12. The information processing apparatus according to claim 1 or 2, further comprising a posture estimation unit that performs posture estimation using key points complemented by the complementation unit.

13. The information processing apparatus according to claim 1 or 2, wherein the keypoint acquisition unit detects keypoints in images for multiple frames of the target to be detected.

14. Procedure (a) for obtaining keypoint detection results for multiple frames to be detected, A process including (b) a step of using the keypoint detection results for the multiple frames to fill in any undetected keypoints in the keypoint detection results, The information processing program instructs a computer to perform a process in which, according to the processing method of the procedure (a) described above, the number of frames of the keypoint detection result to be acquired is changed.

15. Procedure (a) for obtaining keypoint detection results for multiple frames to be detected, A process including (b) a step of using the keypoint detection results for the multiple frames to fill in any undetected keypoints in the keypoint detection results, An information processing program that causes a computer to perform the following steps in (b): input keypoint detection results for multiple frames, including missing frames in which at least some keypoints are missing, and the ground truth data corresponding to the keypoint detection results into a learning model, and then use the trained model, which has been machine-trained using the ground truth data as its target, to complete the process of filling in the undetected keypoints in the keypoint detection results.

16. A receiving unit that receives keypoint detection results for multiple frames, including missing frames in which at least some keypoints are missing, and the correct answer data corresponding to said keypoint detection results, A machine learning device comprising: a learning unit that inputs the keypoint detection results for multiple frames and the ground truth data into a learning model, and generates a trained model by having the learning model machine learn using the ground truth data as the target.

17. A procedure (a) for receiving keypoint detection results for multiple frames, including missing frames in which at least some keypoints are missing, and the ground truth data corresponding to said keypoint detection results, A machine learning program that causes a computer to perform the following steps: (b) A procedure in which the keypoint detection results for multiple frames and the ground truth data are input into a learning model, and the learning model is trained using the ground truth data as the target to generate a trained model.

Citation Information

Patent Citations

  • Multi-person behavior recognition system based on key point detection and working method

    CN110929687A

  • Behavior analysis method and device based on human body key point detection

    CN111027481A

  • Image processing program and image processing device

    JP2020086819A

  • Object recognition device, object recognition method and object recognition program

    JP2020135551A

  • Target searching device and method, and electronic apparatus

    JP2021034015A