Monitoring entities in healthcare facilities
Patent Information
- Application Number
- JP2024527293
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-11-16
- Filing Date
- 2022-11-09
- Publication Date
- 2025-11-18
AI Technical Summary
Existing video analysis techniques struggle to automate the semantic understanding of complex and cluttered clinical environments in healthcare facilities, such as hospitals, due to challenges in object detection and tracking using conventional image processing and deep learning methods like YOLO and OpenPose.
The use of articulated joint models with keypoints and association fields, combined with machine learning processes, to fit and determine the position or pose of entities in medical facilities, enabling more accurate semantic understanding and automation of workflows.
Enhances the ability to automate clinical workflows by providing detailed tracking and interaction analysis of patients and equipment, improving resource allocation and event detection in healthcare settings.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present disclosure relates to monitoring entities (eg, people, clinicians, equipment) in a healthcare facility. [Background technology]
[0002] Workflows (also known as clinical workflows) are used in healthcare facilities (hospitals, clinics, etc.) to ensure that the right treatment is performed for each patient in a standardized way. This helps to ensure compliance with best practices and clinical guidelines in the healthcare facility. Workflows often specify a specific set of tasks or checks (items in the workflow) to be performed on a patient. Workflows can be used at all stages of a patient's care, for example there may be a workflow associated with admitting a patient to a healthcare facility, another workflow associated with triaging the patient, and subsequent workflows used depending on the specific problem or care pathway identified for the patient.
[0003] Workflow management (e.g., recording when actions in a workflow have been performed) is a significant and even significant overhead in healthcare facilities. Thus, automated analysis, optimization, and control of clinical workflows is an ongoing area of active research. Summary of the Invention [Problem to be solved by the invention]
[0004] In addition to workflow management, there are other tasks in a healthcare facility where it is desirable to automate, for example, equipment and / or patient tracking.
[0005] The present disclosure aims to address these and other problems. [Means for solving the problem]
[0006] Various projects aim to automate different aspects of workflow management. Previous work in this field has used infrared sensor tags to track patients, medical staff and equipment in hospitals, aiming to improve resource allocation and avoid supply bottlenecks, for example in emergency departments. However, such data is often of relatively coarse resolution in time and space, making the subsequent semantic understanding of clinical processes far from straightforward.
[0007] Another project proposes the use of in-hospital video (infrared and / or depth) data, which provides richer information, i.e. allows capturing the presence, location and activity of multiple people, e.g., healthcare providers and patients, as well as the use of medical equipment in great spatial and temporal detail. The equipment and devices in the room can be combined with information about the people in the image to give a complete picture of the situation. Video monitors directly capture events, e.g., a nurse changing an infusion pump, a nurse operating a monitor, a patient sitting in a chair for a while, etc. However, a key challenge is to automate such video analysis with computer algorithms.
[0008] In particular, clinical environments are often cluttered and highly complex scenes in which conventional, well-known image processing techniques for object detection and tracking tend to struggle or fail altogether.
[0009] Artificial intelligence (AI) techniques, and especially deep learning (DL) methods of large-scale neural networks, offer opportunities for real-time video analysis. For example, the YORO (You Only Look Once) algorithm described in the paper by J. Redmon, S. Divvala, R. Girshick and A. Farhadi ("You Only Look Once: Unified, Real-Time Object Detection," 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 779-788, doi: 10.1109 / CVPR.2016.91) allows real-time identification and tracking of objects in video streams. YOLO generates a bounding box that identifies the location of the desired object, but has the drawback that it is difficult to infer further semantic meaning from video images using YOLO alone.
[0010] Another deep neural network solution, the “OpenPose” algorithm (Hidalgo et al., “Single-Network Whole-Body Pose Estimation”, 2019), is capable of detecting humans in image and video data. OpenPose provides more information than YOLO, as it includes the location of key points in the image, including the position of the person’s head, shoulders, hips, and elbows (depending on the precision model used).
[0011] For example, the inventors have realized that algorithms such as OpenPose can be advantageously applied in medical facilities to extract semantic information from video images, allowing for a deeper understanding of what is happening in the hospital. As explained in more detail below, such semantic information can be used to update clinical workflows in a reliable and automated manner.
[0012] Thus, according to a first aspect of the present specification there is provided a method for use in monitoring a first entity in a healthcare facility, the method comprising: i) acquiring an image of a medical facility; ii) fitting a first joint model to a first entity in the image using a machine learning process, the first joint model having keypoints corresponding to joints and affinity fields indicating links between the keypoints; and iii) determining a position or pose of a first entity in the medical facility from relative positions of fitted keypoints of a first joint model in the images. has.
[0013] In some embodiments, the keypoints correspond to location coordinates and the association field corresponds to vectors connecting the coordinates of related keypoints.
[0014] In some embodiments, the first joint model is represented as a tuple of coordinates, where each coordinate in the coordinate tuple corresponds to a keypoint, and a tuple of vectors between different pairs of coordinates in the coordinate tuple, where each vector corresponds to an association field.
[0015] In some embodiments, the machine learning process involves the use of a neural network (i.e., an artificial neural network).
[0016] According to a second aspect, there is a computer program product having a computer readable medium having computer readable code embodied therein, the computer readable code being configured, upon execution of instructions by a suitable computer or processor, to cause the computer or processor to perform the method of the first aspect.
[0017] According to a third aspect, there is an apparatus for use in monitoring a first entity in a healthcare facility, the apparatus having a memory having instruction data representing a set of instructions, and a processor in communication with the memory and configured to execute the set of instructions, the set of instructions, when executed by the processor, causing the processor to: i) acquiring images of a medical facility; ii) using a machine learning process to fit a first joint model to a first entity in the image, where the first joint model has keypoints corresponding to the joints and an association field indicative of links between the keypoints; and iii) determining a position or pose of a first entity in the medical facility from relative positions of fitted keypoints of a first joint model in the images; Make them do this.
[0018] Thus, in embodiments herein, entities in a hospital are modeled in an articulated manner using an articulated model having key points and an association field. It is recognized that the flexibility of the articulated model is well suited to the complex and often cluttered scenes in medical facilities. Furthermore, it is recognized that many medical devices (such as ventilators) are advantageously fitted using the articulated model. The relative positions between the fitted key points and the association field allow the position and / or pose of the entities to be better determined. This is used in various scenarios, for example to provide a semantic understanding of video images of the hospital that are linked to workflows for workflow automation in the hospital.
[0019] These and other aspects will be apparent from and elucidated with reference to the embodiments described hereinafter.
[0020] [Brief description of the drawings]
[0021] Exemplary embodiments will now be described, by way of example only, with reference to the following drawings, in which: [Figure 1] FIG. 1 is an exemplary apparatus for monitoring a healthcare facility according to some embodiments herein. [Diagram 2] FIG. 2 is an exemplary method for monitoring a healthcare facility according to some embodiments herein. [Figure 3a] FIG. 3a shows an exemplary joint model of a bed. [Figure 3b] FIG. 3b shows an exemplary articulated model of a ventilator. [Figure 3c] Figure 3c shows an example of seizure detection using a human joint model. [Figure 4a] FIG. 4a shows an exemplary image (photograph represented as a line drawing) with two joint models superimposed on the patient in the image and the monitor in the image. [Figure 4b] FIG. 4b shows the interaction between the fitted joint model of the patient shown in FIG. 4a and the fitted joint model of the monitor. [Diagram 5] FIG. 5 shows an example image (photograph represented as a line drawing) of a bed and a monitor, with a first joint model fitted to the bed and a second joint model fitted to the monitor. [Figure 6] FIG. 6 shows an exemplary image (photographs represented as line drawings) of two patient beds side-by-side, each with a patient interacting with a respective clinician. [Figure 7] FIG. 7 illustrates a method according to some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0022] As mentioned above, the objective of the embodiments described herein is to provide improved semantic understanding of video streams in medical environments (hospitals, clinics, doctor's offices, dental offices, etc.), but particularly, but not exclusively, for use in automated workflow analysis.
[0023] 1, in some embodiments there is an apparatus 100 for use in monitoring a first entity in a healthcare facility, according to some embodiments herein. Typically the apparatus forms part of a computing device or system, such as, for example, a laptop, desktop computer or other computing device. In some embodiments the apparatus 100 may form part of a distributed computing arrangement or cloud.
[0024] The apparatus includes a memory 104 having instruction data representing a set of instructions 106, and a processor 102 (e.g., processing circuitry or logic circuitry) in communication with the memory and configured to execute the set of instructions. In general, the set of instructions, when executed by a processor, causes the processor to perform any of the embodiments of a method 200 as described below.
[0025] An embodiment of the device 100 may be an apparatus for use in monitoring a first entity in a healthcare facility. More specifically, the set of instructions, when executed by a processor, causes the processor to: i) obtain an image of the healthcare facility; ii) use a machine learning process to fit a first joint model to the first entity in the image, where the first joint model has keypoints corresponding to joints and an association field indicating links between the keypoints; and iii) determine a position or pose of the first entity in the healthcare facility from relative positions of the fitted keypoints of the first joint model in the image.
[0026] The processor 102 may have one or more processors, processing units, multi-core processors, or modules configured or programmed to control the device 100 in the manner described herein. In certain implementations, the processor 102 may have multiple software and / or hardware modules, each configured or programmed to perform individual or multiple steps of the methods described herein. The processor 102 may have one or more processors, processing units, multi-core processors, and / or modules configured or programmed to control the device 100 in the manner described herein. In some embodiments, for example, the processor 102 may have multiple (e.g., used simultaneously) processors, processing units, multi-core processors, and / or modules configured for distributed processing. It will be understood by those skilled in the art that such processors, processing units, multi-core processors, and / or modules may be located in different locations and perform different steps of the methods described herein and / or different portions of a single step.
[0027] The memory 104 is configured to store program code executed by the processor 102 to perform the methods described herein. Alternatively or additionally, one or more memories 104 may be external to the device 100 (i.e., separate or remote from the device). For example, one or more memories 104 may be part of another device. The memory 104 is used to store an image, a first joint model, a determined position or pose of a first entity, and / or any other information received, calculated or determined by the processor 102 of the device 100 or from any interface, memory or device external to the device 100. The processor 102 is configured to control the memory 104 to store the image, the first joint model, the determined position or pose of the first entity.
[0028] In some embodiments, memory 104 may have multiple sub-memories, each capable of storing instruction data, such as at least one sub-memory capable of storing instruction data representing at least one instruction of the set of instructions, while at least one other sub-memory may store instruction data representing at least one other instruction of the set of instructions.
[0029] It will be understood that Fig. 1 shows only the components required to illustrate this aspect of the disclosure, and that in a practical embodiment, the device 100 may have additional components to those shown. For example, the device may have an image capture unit used to capture images in a medical facility. The image capture unit may include any audiovisual equipment capable of taking images or videos, such as a camera, an infrared camera, etc. This image capture unit may be connected to the device, for example, via a wired or wireless connection.
[0030] In another example, the device may be configured to receive images from an image acquisition unit that is separate from device 100 (eg, via a wired or wireless connection).
[0031] In some examples, as described in more detail below, the apparatus 100 may further include a time-of-flight (ToF) camera. A ToF camera generates an image matrix where the value of each pixel is the distance / depth of the object from the camera. ToF cameras generally use infrared image sensors. Such devices can generate both "depth images" and traditional infrared intensity images simultaneously.
[0032] In other examples, device 100 may be configured to receive images and / or image matrices from a ToF camera that is separate from device 100 (eg, via a wired or wireless connection).
[0033] More generally, the apparatus 100 may further comprise a display. The display may comprise, for example, a computer screen and / or a screen on a mobile phone or tablet, for example for displaying the image and / or the fitted model to a user. The apparatus may further comprise a user input device, for example a keyboard, mouse or other input device, allowing a user to interact with the apparatus, for example to provide initial input parameters (e.g. selection of a model) used in the methods described herein. The apparatus 100 may comprise a battery or other power source for powering the apparatus 100, or a means for connecting the apparatus 100 to a mains power source.
[0034] The device may be used in a medical facility. Examples of medical facilities include, but are not limited to, hospitals, clinics, doctor's offices, dental offices, and veterinary clinics. As mentioned above, a medical facility may use workflows (also known as clinical workflows) to monitor activities occurring in the facility.
[0035] The device is for monitoring a first entity in a medical facility. In this sense, the first entity can be any object, person or animal in the medical facility. For example, the first entity can be a person in the medical facility, such as a patient, a care provider, a doctor, a nurse, a surgeon or a cleaner. As another example, the first entity can be a piece of equipment in the medical facility, such as a ventilator, a monitor, a SpO2 device or a vital signs monitor. As yet another example, the first entity can be any other inanimate object in the medical facility, such as a hospital bed, a chair, a wheelchair or a walking device.
[0036] As mentioned above, the device 100 has an image acquisition unit or receives an image from an image acquisition unit. The image may be any image of (e.g., the interior of) a medical facility. The image may be a photographic image. The image may be in color (e.g., an RGB image) or in black and white. The image may be an infrared image or any other type of image modality.
[0037] As will be explained below, this image may be a single frame or may be included in a video, for example as part of a sequence of video frames.
[0038] The image acquisition unit may form part of an audiovisual equipment, for example a video camera. Two or more cameras / video cameras may be implemented in the medical facility. Such cameras / video cameras may be arranged to continuously cover a part of the interior of the medical facility. As an example, the images or video streams may be acquired from a system similar to a CCTV (integrated closed-circuit television) system.
[0039] Referring to FIG. 2, there is a computer-implemented method 200 for use in monitoring a first entity in a healthcare facility. An embodiment of the method 200 may be performed by an apparatus such as, for example, the apparatus 100 described above.
[0040] Briefly, in a first step 202, the method 200 comprises i) acquiring an image of a medical facility. In a second step 204, the method 200 comprises ii) fitting a first joint model to a first entity in the image using a machine learning process, the first joint model having key points corresponding to joints and an association field indicating links between these key points. In a third step 206, the method comprises iii) determining a position or pose of the first entity in the medical facility from relative positions of the fitted key points of the first joint model in the image. As mentioned above, the step 202 "i) comprising acquiring an image of a medical facility" may be performed in different ways. For example, the image may be received from an image acquisition unit (e.g. in real time or near real time). In other examples, the image may be retrieved from a server, a database of images and videos, or the like.
[0041] The image may show part of a healthcare facility, for example part of the interior or exterior of a healthcare facility. The image may for example be an image of a ward, a clinic or an examination room.
[0042] In step 204, the method 200 includes ii) using a machine learning process to fit a first joint model to the first entity in the image, where the first joint model has keypoints corresponding to the joints and an association field indicating links between the keypoints.
[0043] In other words, the joint model is fitted to the first entity in the image. The joint model, also known as a "skeletal model" or articulated skeleton, has keypoints that correspond to joints in the model. The keypoints correspond to landmarks on the first entity. In general, each keypoint corresponds to a point, i.e., a specific location, on the first entity.
[0044] These key points correspond to (external) joints, e.g., flexible joints that allow rotation or translation of the structure represented by the association fields on either side of the key point. The joints may be, for example, pivot points on the first entity. For example, if the first entity is a person (e.g., a patient, a clinician, etc.), one or more key points are defined in the first joint model that correspond to one or more anatomical joints of the person. For example, one or more key points are defined that correspond to a hip joint, a shoulder joint, a knee, an elbow, or any other joint on the person's body. If the first entity is a device, one or more key points are defined in the first joint model that correspond to a joint or joint in the device. For example, if the first entity is a ventilator, a key point is defined that corresponds to where the mask is attached to the hose of the ventilator.
[0045] The first joint model further comprises key points corresponding to landmarks on the first entity in addition to the joints. For example, if the first entity is a person, the first joint model may further comprise key points for other landmarks of the person, e.g. specific anatomical features such as the person's eyes, nose or shoulders. If the first entity is a device, the key points may for example correspond to specific landmarks of the device, i.e. the device's outer edges (e.g. edges or corners) for example of a mask.
[0046] The first joint model may have one or more key points corresponding to joints and one or more key points corresponding to landmarks, e.g., a mixture of key point types. In general, when designing a joint model, key points should be selected from prominent and distinctive image features of the object, e.g., a person's eyes or the wheels of a hospital bed. Furthermore, for a joint model, the key points should coincide with the joints or hinges of the object. It will be understood that these are merely examples, and that depending on the nature of the first entity being modeled, key points may be defined at many different locations on this first entity.
[0047] The first articulated model further comprises an association field, the association field indicating or corresponding to links between the key-points.
[0048] One or more association fields in the first joint model correspond to physical links. For example, if the first entity is a person, examples of physical links include, but are not limited to, the thigh (represented by an association field placed between a key point corresponding to the hip joint and a key point corresponding to the knee) and the forearm (represented by an association field placed between a key point corresponding to the hand and a key point corresponding to the elbow). If the first entity is a ventilator, the physical link corresponds to the hose (placed between two key points corresponding to the mask and the base unit).
[0049] One or more of the association fields may also have a logical link. A logical link may have or represent a positional relationship between two key points (e.g., between locations on the first entity to which the key points correspond) even if the two key points are not directly connected (e.g., by a single corresponding device or anatomical structure). For example, in an example where the first entity is a person, a logical link exists between the jaw and the clavicle, even if the jaw and the clavicle are not directly connected, because a positional relationship exists between the jaw and the clavicle. Thus, in a human joint model, an association field may be defined between two key points corresponding to the jaw and the clavicle.
[0050] The first joint model may, for example, have one or more relationship fields corresponding to physical links and one or more relationship fields corresponding to logical links, such that it is a mixture of relationship field types.
[0051] Step 204 includes obtaining a first joint model, which may be obtained from a database of joint models or may be defined by, for example, a human engineer.
[0052] The first joint model can be represented as a tuple of coordinates (e.g., in a normalized coordinate system) and as a tuple of vectors between different pairs of coordinates in the tuple of coordinates, where each coordinate in the tuple of coordinates corresponds to a keypoint as described above, and each vector corresponds to an association field as described above, however, this is merely an example and one skilled in the art will appreciate that the first model may be represented in ways different from that described above.
[0053] As an example, if the first entity is a person (e.g., a patient, a doctor, a nurse, etc.), an appropriate first joint model is specified in the OpenPose paper by Hidalgo (2019).
[0054] As another example, if the first entity is a bed 302, a first joint model is defined as shown in Fig. 3a. Fig. 3a shows the joint model of the bed in the form of a directed graph. Its vertices are the keypoints, i.e., ["1-front right wheel", "2-front left wheel", ..., "12-headboard top right"]. The edges of this graph represent the connectedness of the keypoints, i.e., [["35" → "31"], ["36" → "32"], ..., ["42" → "41"].
[0055] As another example, if the first entity is a ventilator 304, a first articulated model is defined as shown in Fig. 3b. In Fig. 3b, there is a ventilator (box with key points 51...58) on a roll stand (43...46) with a table (47...50). The articulated ventilator tubes are given at 52-59-60.
[0056] The advantage of modeling entities as articulated models (or skeletons) is that articulated models are inherently able to adapt to object-specific pose changes / deformations. For example, these models can define the skeleton of a human body equally accurately no matter what the particular pose is, e.g., whether the arms are raised or not. Similarly, articulated models can be used to detect hospital beds whether the backrest is raised or not, since the backrest is an articulated component of the bed articulated model.
[0057] The first joint model is fitted to the first entity in the image using a machine learning process. An example of a suitable machine learning process is the OpenPose method as described in Hidalgo (2019). In Hidalgo, a first deep neural network converts the image data into a heatmap of skeletal keypoints and a part association field (PAF). The heatmap is trained to transform the image into part-affinity fields (Heatmaps). The heatmaps indicate the likelihood that each pixel in the image is a keypoint. One heatmap is generated for each keypoint in the articulated model.
[0058] The PAF is an image (consisting of two planes, one for the x component and one for the y component) where each pixel corresponds to a 2D vector (x and y components). Thus, the PAF is a 2D vector field.
[0059] Furthermore, a PAF always belongs to a pair of keypoints, e.g., "right elbow" and "right shoulder".
[0060] The (x,y) vector at a particular pixel location in the PAF encodes two pieces of information: (1) The magnitude of the vector (x 2 +y 2 ) 1 / 2 indicates the likelihood that the pixel belongs to a connection ("limb") between two keypoint instances (i.e., the "right upper arm" in this embodiment). (2) The direction of the (x,y) vector at that pixel encodes the orientation of that limb, i.e., which is the elbow and which is the shoulder.
[0061] This vector field information is then used to determine which pairs of candidate keypoints (as detected in the heatmap) are the highest and form a limb. This is done by computing the path integral of the part-association vector field from the location of one candidate keypoint to the location of another candidate keypoint.
[0062] It should also be noted that this coding principle can be applied in more than two dimensions, e.g. three spatial dimensions in the case of volumetric data, and / or the temporal dimension in the case of e.g. video data.
[0063] The heatmap and PAF are then fed into a (two-part) graph matching stage (a.k.a. skeletal parser), which generates a skeletal model-based description of the object-content of the image, using the method described in Hidalgo et al. (2019).
[0064] Thus, in other words, the machine learning process is performed in two steps. First, a first deep neural network is used to determine a first set of locations in the image that correspond to keypoints in the first joint model. This is similar to using a neural network to identify landmarks in an image. For example, the first deep neural network is trained on a corpus of training images to identify keypoints in the images using a supervised learning process. The first neural network can output a heatmap and / or relevance field as described above.
[0065] Next, a first graph fitting process (or "skeleton parser") is used that is capable of taking as input locations in the image that correspond to the keypoints and relationship fields in the first joint model and fitting the first joint model to the first entity in the image.
[0066] The output of a machine learning process (e.g., the output of a graph fitting process) can be, for example, a list of coordinates in the image of keypoint locations for each detected object in the image. Thus, the output can be a list of lists of 2D-locations. For example, if we model a person as person=[head, left shoulder, right shoulder] and there are two people in the image, the output will be of the following form: [Person 1, Person 2] = [[[Head X, Head Y], [Left shoulder X, Left shoulder Y], [Right shoulder X, Right shoulder Y]], [[Head X, Head Y], [Left shoulder X, Left shoulder Y], [Right shoulder If no instances of the object are found, in this example, the output is an empty list.
[0067] It will be appreciated that this is just one example and that the inputs and outputs of the first joint model may differ from those described in the above example.
[0068] Returning to step 206, the method further comprises iii) determining a position or pose of a first entity in the medical facility from the relative positions of the fitted keypoints of the first joint model in the image. Once the positions of the different keypoints have been obtained from the fitted model in the image, the positions of the keypoints can be used to determine the position and / or pose of a first entity in the image.
[0069] A pose may be determined from the fitted keypoints in various ways, for example, a series of if-then statements may be used to determine whether the positions of the fitted keypoints are consistent with a particular pose. For example, if the first entity is a person, then if the position of the keypoint corresponding to the hand in the image is higher than the position of the keypoint corresponding to the shoulder to which the hand is attached, then the person is determined to have raised their arm (i.e., is in an arms up pose). As another example, if the first entity is a bed, then the pose of the bed may be labeled as "reclined" or "upright" depending on the angle made by the keypoints corresponding to the head of the bed and the body or main part of the bed.
[0070] In other embodiments, a different machine learning model is trained and used to provide pose labels based on the fitted keypoints in step 206. For example, a convolutional neural network is trained to take the fitted keypoints and / or images as input and output labels describing the poses. Training is done in a supervised manner using a training dataset that has examples of fitted keypoint locations and corresponding ground truth pose labels.
[0071] The determined position and / or pose of the first entity is used to perform various tasks. Some examples are:
[0072] Equipment Tracking Often, carts,(EMR) computers are scattered throughout a department and finding them can take a significant amount of time.,Using this approach, a specific cart can be identified and,located in an automated manner.
[0073] Checking object position Certain patients may require a particular position (e.g., leg elevation) to ensure a safe and speedy recovery. As an example, an articulated model may be fitted to a bed and the angle of the bed may be determined from the fitted key points to ensure the patient is in an upright position. If the angle is outside of a predefined range, a signal is provided that the angle is not within the specified range.
[0074] As another example, a joint model can be fitted to a patient and the fitted keypoints can be used to determine if the patient has lifted a leg, for example as implemented in the following pseudocode: For each image in the video stream: Detecting people (skeleton) in an image Calculate the position of the patient's legs (using the position coordinates and viewpoint of the installed camera) if the ankles / knees / hips are a visible part of the detected skeleton If the position is out of tolerance, alert the care provider: “Patient position is incorrect!” If not, it will alert the care provider that "patient position cannot be monitored." Do this.
[0075] If a particular pose or patient position is required as part of a medical workflow, this is used as a trigger to perform the method 200 for that particular patient and that particular position / pose. This is used to automate aspects of this workflow.
[0076] Turning now to other embodiments, as mentioned above, in some embodiments the images are frames of a video and the method may further comprise repeating steps i), ii) and iii) for a sequence of frames of the video and determining a change in pose or change in position of the first entity over the sequence of frames. Thus, method 200 may be repeated to process the video data and monitor entities in the medical facility over time.
[0077] Analysis of the video data is used to determine the change in position or pose of the first entity over time. For example, the change in position is determined by comparing the positions of fitted keypoints in the first joint model from one frame of the video to another. Similarly, the change in pose is determined by determining a pose of a first image of the video, determining a pose of a second image of the video, and determining the change in pose as the difference between the pose of the first image and the pose of the second image.
[0078] In some embodiments, the position or pose (or change in position or pose) is used to determine if an event has occurred for the first entity. For example, if the first entity is a medical device, the change in position or pose of the medical device is used to determine if the medical device has moved from a first position to a second position. As another example, if the first entity is a person, the event may include the person getting out of bed, the person having a seizure, or the person remaining in one position for longer than a predetermined time threshold.
[0079] For example, seizure detection may involve tracking the position of a body part over time and raising an alarm when, for example, large vibrations occur.
[0080] This is illustrated in Figure 3c, which shows a person 306 and a fitted joint model 308 fitted to the person 306. The positions of the key points can be plotted against time. When a seizure occurs, vibrations at the key point positions, indicated by circles 310, are detected as shown in the graph of the left shoulder position against time.
[0081] For example, detection of a medical condition such as a seizure may be used to trigger an automatic update of the patient's medical record (e.g., with details such as time, date, location and / or duration of the seizure). Detection of a medical condition (e.g., a seizure) may further trigger a new workflow to be implemented for the patient. Additionally, facial recognition is used to ensure the correct medical record is updated. In this manner, method 200 may be used for automated record management and workflow management.
[0082] Referring now to other embodiments, the method 200 may further comprise fitting different joint models to different entities in the images of the video and determining interactions between these different entities from relative positions of respective fitted key-points of these models. For example, the above steps i) to iii) may be repeated for each entity in the image or sequence of images.
[0083] In other words, the method 200 may further comprise fitting a second joint model to a second entity in the image using a machine learning process. The second joint model comprises a (second set) of key points and a (second set) of association fields indicating links between the (second set) of key points. The method then comprises determining interactions between the first and second entities in the image from the relative positions of the fitted key points of the first joint model and the fitted key points of the second joint model. It will be appreciated that the method may be further extended to a third entity and / or subsequent entities in the image.
[0084] Although the joint model has been described above with respect to the first joint model, it will be understood that the details apply equally to the second joint model, e.g., the second joint model also has keypoints corresponding to the joints and an association field indicating the links between these keypoints.
[0085] The first and second joint models can be the same type of model or different types of models depending on the type of interaction being monitored, for example, in an interaction between a patient and a doctor, the first and second joint models are both human joint models.
[0086] The extension of the method 200 to the first and second entities can be achieved in various ways. For example, if OpenPose (as described in Hidalgo (2019)) is used in step 204, what is proposed here requires an extension of the network stage to include the detection of (new) keypoint types specific to the entity(ies) that must be detected. Of course, along with extending the neural network architecture itself, the training dataset must also be extended so that it includes appropriate instantiations of the objects and a sufficient coverage of the different poses that occur naturally.
[0087] Furthermore, a skeletal parser must be written such that it parses the skeletal model of a specific (new) object from the heatmap and the part association field, i.e., it requires a tailored model description. Different options exist to achieve this. 1) Using multiple custom OpenPose systems used in parallel for different kinds of skeletal models, e.g., multiple different dedicated systems are fed with the same video data. In other words, a second deep neural network (different from the first deep neural network) is trained and used to determine a second set of positions in the image that correspond to keypoints in the second joint model, and a second graph fitting process is used to fit the second joint model to a second entity in the image, taking as input the positions in the image that correspond to the keypoints and the relevance field in the second model. 2) The OpenPose systems share the trunk of the neural network processing and only use different skeletal model parsers. In other words, the first deep neural network is further trained and used to determine a second set of positions in the image that correspond to keypoints in the second joint model (e.g., the same deep neural network is trained and used to determine the positions of the keypoints of the first joint model in the image and the keypoints of the second joint model in the image). Then, to fit the second joint model to a second entity in the image, a second (e.g., different) graph fitting process is used for the second entity, taking as input the positions in the image that correspond to the relevance field and keypoints in the second model.
[0088] Determining the interaction between the first entity and the second entity is performed, for example, by detecting overlap between fitted keypoints of the first joint model and fitted keypoints of the second joint model.
[0089] In some embodiments, depth information may be used to enhance understanding of an interaction between a first entity and a second entity in an image, and, for example, to ensure that the interaction is a real interaction and not just a coincidence of an overlap between two entities in an image (e.g., due to the camera angle). When depth images are used to localize the skeletons of objects / people, the 3D coordinates of these objects / people can be derived. Over time, the 3D coordinates form a trajectory flow of the object / person. This gives more insight into the semantics of the scene. For example, if the skeleton of a patient's hand / arm is raised, it may indicate that the patient is trying to ask for help or to attract attention. Depth information is also useful to resolve ambiguities and distinguish between occlusions and interactions.
[0090] Thus, the method 200 further comprises determining depth information associated with the fitted keypoints in the first joint model and the fitted keypoints in the second joint model. Determining an interaction between the first entity and the second entity in the image can then be further based on the depth information.
[0091] Depth information is determined, for example, using a ToF camera, which generates an image matrix where the value of each pixel is the distance / depth of the object from the camera. ToF cameras generally use a specific infrared imaging sensor, so that such devices can simultaneously generate both depth images and conventional infrared intensity images.
[0092] In general, the depth information is matched to the fitted keypoints according to the following pseudocode: For each image in the video stream, Detect all objects (skeletons) in a 2D image (e.g. infrared or RGB image) Search for skeleton keypoints detected in depth images Construct 3D position data of skeleton keypoints from 2D and depth information Further processing with 3D skeleton data (which is more robust than 2D) Do this.
[0093] As an example of the "processing further" step in the above example, interactions between a first entity (the bed) and a second entity (the patient) can be used to determine if the patient has left the bed. If the 2D or 3D distance between the keypoints of the first joint model (for the patient) and the fitted keypoints of the second joint model (for the bed) exceeds a predefined threshold distance, then it is determined that the patient has left the bed. This can be accomplished, for example, according to the following pseudocode: For each image in the video stream, Detecting human (skeleton) in an image Detecting beds (skeleton) in images If 3D distance (from skeleton to bed) > threshold, Alerts care provider that "Patient is out of bed!" Do this.
[0094] As another example, depth information can be used to determine, for example, that a ventilator is connected to a patient by a (articulated) tube, rather than simply that the ventilator is present in the room, as would result from traditional object detection in photos / videos.
[0095] This is illustrated in FIG. 4a, which shows a line drawing of an image of a medical facility. The line drawing represents a photograph of the medical facility. It will be understood that in reality the image is a photograph and may for example be in color. The image shows a patient 402 in a bed 404. The patient 402 is being ventilated by a ventilator 406. A first joint model 408 of a person is fitted to the patient 402. A second joint model 410 of the ventilator is fitted to the ventilator 406. FIG. 4b shows the same fitted joint models 408 and 410 shown in FIG. 4a. In FIG. 4b, the ventilation tube of the ventilator is represented by the relationship field 412 and the stand associated with the ventilator is represented by the relationship field 416. The junction between these two relationship fields is represented by the key point 414. Other key points on the ventilator include for example the point 418 which represents the point between the stand 416 and the screen (the square structure above the stand). The corners of the screen are represented by key points such as 420, and the edges are represented by association fields (e.g., 422). In this example, the overlap between skeletons 408 and 410 can be used to determine that a ventilator 406 is connected to a patient 402. If 3D coordinates are acquired (e.g., via a ToF camera), these can be used to verify that the patient is connected to the ventilator, for example, to exclude accidental overlaps. FIG. 5 shows another line drawing of an image of a medical facility. This line drawing represents a photograph of a medical facility. It will be understood that in reality, the image is a photograph, and may be, for example, in color. This example shows articulated models 504 and 508 fitting to a bed 502 and medical equipment 506, respectively.
[0096] In general, the interaction between the person's fitted joint model and the object's fitted joint model can provide insights that can be used to update the clinical workflow. For example, the intersection / overlap of two skeletons can be interpreted as an interaction occurring. As an example, the method 200 can be used to register events or data such as how often a nurse (first entity) operates an equipment (second entity) such as a mobile patient monitor, when the nurse (first entity) interacts with an infusion pump (second entity), or whether a patient's bed is adjusted to a sitting or sleeping position.
[0097] In an embodiment where the first entity is a medical device and the second entity is a patient, the method 200 is used to determine that an event has occurred between the device and the patient. Examples include, but are not limited to, detecting a device attached to a patient or detecting a device being used to perform a medical procedure on the patient. Again, this information can be used to update one or more workflows or medical records associated with the patient.
[0098] In an embodiment where the first entity is a clinician and the second entity is a patient, the method 200 is used to determine a first interaction between the clinician and the patient. Examples include, but are not limited to, a first interaction that is contact between the clinician and the patient or a medical procedure that the clinician performs on the patient. Facial recognition is further used to link an interaction between a patient and a medical professional to the correct patient and medical professional.
[0099] The following pseudocode shows an example in which interactions between each of the above mentioned patients (first entity), clinicians (second entity) and / or medical devices (third entity) are counted and used to provide statistical data for use in analyzing and optimizing workflow efficiency. For each image in the video stream, Detect all objects (skeleton) in the image (e.g. medical equipment, beds, etc.) Detect all human (skeleton) in an image option, (1) A marker attached to human clothing (e.g., a QR code) (2) Face recognition for each detected human region in the image to classify people into "nurses", "patients" and "visitors" Calculate the geometric distance from the person to the detected object (e.g. person in bed → patient, person near medical equipment → nurse) Counting interactions between people and objects (e.g., a nurse interacting with a patient, a bed, or a piece of medical equipment) Calculate interaction frequency statistics and use them to analyze and optimize workflow efficiency Do this.
[0100] The output of this process is shown in FIG. 6, which shows a line drawing of an image of a medical facility. The line drawing represents a photograph of the medical facility. It will be understood that in reality the image is a pixelated photograph and may be, for example, in color. The image shows a first bed 602 and a first patient 606 in the first bed interacting with a first person 610. The image further shows a second bed 614 with a second patient 618 in it. A second person 622 is interacting with the second patient 618. In this example, the result of the step "Detect all objects (skeleton) in the image (e.g., medical equipment, beds, etc.)" is fitting a joint model 604 to the bed 602 and a joint model 616 to the bed 614. The result of the step “Detect all humans (skeleton) in image” is to fit joint model 608 to first patient 606, fit joint model 612 to first person 610, fit joint model 620 to second patient 618, and fit joint model 624 to second person 622.
[0101] Reference is now made to FIG. 7, which shows a flow chart of a method for determining workflow metrics in a healthcare facility using method 200. The method is performed by an apparatus such as, for example, apparatus 100 described above. In this example, in step 702, an image is acquired (e.g., as illustrated in FIG. 6) according to step 202 of method 200 described above. The image 702 is then fed to a machine learning process 704 that is used to fit one or more joint models to one or more entities (e.g., patient, monitoring equipment) in the image. In this embodiment, the OpenPose machine learning process is used, as described above with respect to step 204 of method 200. In this embodiment, the OpenPose method is extended with new joint models corresponding to hospital equipment, beds and machines (as described above). In step 706, the positions and / or poses of one or more entities in the healthcare facility are determined from the relative positions of the fitted keypoints of the first joint model in the image (as described above with respect to step 206 above). Then, in step 708, the positions and poses are analyzed according to the methods described above to determine activities, events and scene situations. In step 710, metrics (eg, contact metrics between the patient and caregiver, or duration of sitting in position, etc.) are derived and used to update items in the patient's workflow.
[0102] Turning now to other embodiments, as described above, the method 200 is used to update and manage a workflow. More generally, the method 200 further comprises using a position or pose of a first entity to determine whether an item in a clinical workflow has been performed and updating the workflow using the result of this determination. For example, a position, pose or event related to the first entity (patient) or an interaction between the first entity (patient) and a second entity (medical professional, medical equipment or other object in a medical facility) is matched to an item in the workflow and / or triggers an update to said item.
[0103] As discussed above, in other embodiments, if a patient medical event or condition (e.g., seizure detection, fall detection, etc.) is detected using method 200, this can trigger, for example, updating the patient's medical record with data regarding the detected event or condition.
[0104] In other examples, an item in a clinical workflow may trigger the execution of method 200. For example, as described above, if an item in a clinical workflow indicates that the patient should be in a particular position (legs elevated, reclined, upright, etc.), this may trigger the execution of method 200 to determine whether the patient is in the particular position.
[0105] Similarly, if an item in the clinical workflow indicates that a medical procedure should be performed on a patient, the method 200 is used to identify whether an interaction between a clinician (a first entity) and a patient (a second entity) corresponds to a medical procedure being performed, and the clinical workflow is updated accordingly.
[0106] Referring now to other embodiments, a computer program product is provided having a computer readable medium having computer readable code embodied therein, which, when executed by a suitable computer or processor, is configured to cause the computer or processor to perform the methods described herein.
[0107] It will therefore be understood that the present disclosure also applies to computer programs adapted to carry out the embodiments, in particular computer programs on or in a carrier, which may be in the form of source code, object code, code intermediate source and object code, for example in a partially compiled form, or in any other form suitable for use in carrying out the methods according to the embodiments described herein.
[0108] It will also be appreciated that such programs can have many different architectural designs. For example, program code implementing the functionality of the method or system may be subdivided into one or more subroutines. Many different ways of distributing the functionality among these subroutines will be apparent to those skilled in the art. The subroutines are stored together in one executable file to form a self-contained program. Such an executable file comprises computer executable instructions, such as, for example, processor instructions and / or interpreter instructions (e.g., Java interpreter instructions). Alternatively, one or more or all of the subroutines are stored in at least one external library file and linked with the main program statically or dynamically, e.g., at run-time. The main program includes at least one call to at least one of the subroutines. The subroutines also have function calls to each other.
[0109] A carrier of a computer program is any entity or device capable of carrying the program. For example, the carrier includes a data storage device, for example a ROM, such as a CD-ROM or a semiconductor ROM, or a magnetic recording medium, for example a hard disk. Furthermore, the carrier may be a transmissible carrier, for example an electrical or optical signal, conveyed via an electrical or optical cable or by radio or other means. When the program is embodied in such a signal, the carrier may be constituted by said cable or other device or means. Alternatively, the carrier may be an integrated circuit in which the program is embedded, the integrated circuit being adapted to perform, or used for the performance of, the relevant method.
[0110] Variations to the disclosed embodiments can be understood and implemented by those skilled in the art in implementing the principles and techniques described herein, from a study of the drawings, the disclosure and the appended claims. In the claims, the word "comprising" does not exclude other elements or steps, nor does it exclude a plurality of them, even if a plurality is not stated. A single processor or other unit may fulfill the functions of several items recited in the claims. The mere fact that certain means are recited in mutually different dependent claims does not indicate that a combination of these means cannot be used to advantage. A computer program can be stored or distributed on a suitable medium, for example an optical storage medium or a solid-state medium, supplied together with or as part of other hardware, but can also be distributed in other forms, for example via the Internet or other wired or wireless telecommunication systems. Any reference signs in the claims should not be interpreted as limiting the scope.
Claims
1. 1. A computer-implemented method for use in monitoring a first entity in a healthcare facility, the method comprising: acquiring an image of the medical facility; using a machine learning process to fit a first joint model to the first entity in the image, the first joint model having keypoints corresponding to joints and an association field indicating links between the keypoints; and determining a position or pose of the first entity in the medical facility from the relative positions of the fitted keypoints of the first joint model in the image; and fitting a first joint model to the first entity in the image using the machine learning process includes: determining a first set of locations in the image that correspond to the keypoints of the first joint model using a first deep neural network; and using a first graph fitting process that takes as input the relevance field of the first joint model and the locations in the image that correspond to the keypoints to fit the first joint model to the first entity in the image. having Computer-implemented methods.
2. The method of claim 1 , wherein the keypoints correspond to location coordinates and the association field corresponds to vectors connecting the coordinates of related keypoints.
3. The first joint model is a tuple of coordinates, each coordinate in the tuple of coordinates corresponding to a keypoint; and a tuple of vectors between different pairs of coordinates in the tuple of coordinates, each vector corresponding to an association field; The method of claim 1 , wherein:
4. The method of claim 1 , wherein the machine learning process comprises the use of a neural network.
5. the image is a frame of a video; The method comprises: repeating steps i), ii) and iii) for a sequence of frames of said video; and determining a change in pose or a change in position of the first entity over the sequence of frames; The method of claim 1 further comprising:
6. the position or orientation is used to determine whether an event has occurred for the first entity; The first entity is a person and the event is The person is out of bed, the person is having a seizure, or the person remains in one position for longer than a predetermined time threshold; That is, or The first entity is a medical device and the event is: the medical device is being moved from a first location to a second location; the medical device is attached to a patient; or the medical device is being used to perform a medical procedure on a patient; That is, The method of claim 1.
7. using the machine learning process to fit a second joint model to a second entity in the image, the second joint model having keypoints corresponding to joints and an association field indicating links between the keypoints; determining an interaction between the first entity and the second entity in the image from the relative positions of fitted keypoints of the first joint model and fitted keypoints of the second joint model; and determining depth information associated with fitted keypoints of the first joint model and fitted keypoints of the second joint model; and determining an interaction between the first entity and the second entity in the image is further based on the depth information; The method of claim 1.
8. The first entity is a clinician and the second entity is a patient, and the first interaction comprises: Contact between the clinician and the patient; or The medical procedure the clinician is performing on the patient The method of claim 7, wherein
9. using the first deep neural network to determine a second set of locations in the image that correspond to the keypoints of the second joint model; and using a second graph fitting process that takes as input the relevance field of the second joint model and the locations in the image that correspond to the keypoints to fit the second joint model to the second entity in the image. The method of claim 7 further comprising:
10. determining a second set of locations in the images that correspond to the keypoints of the second joint model using a second deep neural network; and using a second graph fitting process that takes as input the relevance field of the second joint model and the locations in the image that correspond to the keypoints to fit the second joint model to the second entity in the image. The method of claim 7 further comprising:
11. The method of claim 1 , wherein the position or pose of the first entity is used to determine whether an item in a clinical workflow has been performed and to update the workflow with the results of the determination.
12. 10. The method of claim 1, wherein the method is triggered by an item in a clinical workflow, and the position or posture of the first entity is used to determine whether the item has been performed and to update the workflow with the results of the determination.
13. 13. A computer program having a computer readable medium having computer readable code embodied therein, the computer readable code being configured, when executed by a suitable computer or processor, to cause the computer or processor to perform a method according to any one of claims 1 to 12.
14. 1. An apparatus for use in monitoring a first entity in a healthcare facility, the apparatus comprising: a memory having instruction data representing a set of instructions; a processor in communication with the memory and configured to execute the set of instructions; the set of instructions, when executed by the processor, cause the processor to: acquiring an image of the medical facility; using a machine learning process to fit a first joint model to the first entity in the image, wherein the first joint model has keypoints corresponding to joints and an association field indicating links between the keypoints; and determining a position or pose of the first entity in the medical facility from the relative positions of the fitted keypoints of the first joint model in the images; Let them do this, Fitting a first joint model to the first entity in the image using the machine learning process includes: determining a first set of locations in the image that correspond to the keypoints of the first joint model using a first deep neural network; and using a first graph fitting process that takes as input the relevance field of the first joint model and the locations in the image corresponding to the keypoints to fit the first joint model to the first entity in the image; having Device.
15. an image acquisition unit for acquiring said image, and / or a Time of Flight camera for acquiring image depth information for the fitted keypoints of the first entity in the image; 15. The apparatus of claim 14, further comprising: