Clinical Activity Recognition Using Multiple Cameras
A multi-camera system automatically recognizes clinical activities using AI and computer vision to enhance patient care and operational efficiency in dynamic environments like operating rooms and ICUs.
Patent Information
- Application Number
- JP2023571952
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-06-10
- Filing Date
- 2022-05-27
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2042-05-27
AI Technical Summary
Existing camera systems for monitoring clinical activities are unreliable and time-consuming due to manual video review, limited subject and movement variations, and inadequate environmental coverage, especially in dynamic settings like operating rooms and ICUs.
A context-aware system using multiple cameras to automatically recognize clinical activities by determining key points, tracking objects, and calculating workflow information without wearable devices, employing AI and computer vision to monitor staff and patient activities in real-time.
Enhances patient care and hospital efficiency by providing automated, scalable, and accurate monitoring of clinical workflows, reducing staff costs and improving operational efficiency.
Smart Images

Figure 0007767464000001 
Figure 0007767464000002 
Figure 0007767464000003
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Patent Application No. 17 / 344,730, entitled "CLINICAL ACTIVITY RECOGNITION WITH MULTIPLE CAMERAS," filed June 10, 2021 (Client Reference Number: SYP339212US01), which is incorporated by reference herein for all purposes as if set forth in its entirety.
[0002] This application is related to U.S. Patent Application Serial No. 17 / 344,734 (SYP339216US01), filed June 10, 2021, entitled "POSE RECONSTRUCTION BY TRACKING FOR VIDEO ANALYSIS," which is incorporated herein by reference for all purposes as if set forth in its entirety. [Background technology]
[0003] Some camera systems can capture video of people, analyze their movements, and generate metadata image or video datasets. Identifying human actions captured by the system's camera video requires a person to manually review the video. Manual monitoring and event reporting can be unreliable and time-consuming, especially when video camera positions and angles change and do not provide sufficient coverage. Multiple cameras can also be used in a controlled environment. However, variations in subject matter, movement, and background can still be significantly limited. Summary of the Invention [Means for solving the problem]
[0004] Embodiments generally relate to recognizing clinical activities using multiple cameras. The embodiments described herein can be applied to recognizing human activities in clinical environments such as operating rooms, intensive care units (ICUs), hospital rooms, emergency rooms, etc. The embodiments provide a context-aware system to provide better patient care and greater hospital efficiency.
[0005] In some embodiments, a system includes one or more processors and logic encoded in one or more non-transitory computer-readable storage media for execution by the one or more processors, wherein the logic, when executed, is operable to cause the one or more processors to perform operations including acquiring a plurality of videos of a plurality of objects in an environment, determining one or more key points for each object of the plurality of objects, recognizing activity information based on the one or more key points, and calculating workflow information based on the activity information.
[0006] In some embodiments, the environment is a surgical room. In some embodiments, the plurality of videos are captured by at least two video cameras. In some embodiments, the activity information includes pose information. In some embodiments, the logic, when executed, is further operable to cause the one or more processors to perform operations including: recognizing one or more objects that are people in the environment, tracking each person's path in the environment, and identifying one or more activities for each person. In some embodiments, the logic, when executed, is further operable to cause the one or more processors to perform operations including: recognizing one or more inanimate objects in the environment, tracking the location of each recognized inanimate object in the environment, and associating the one or more inanimate objects with each person. In some embodiments, the workflow information includes surgical workflow information.
[0007] Some embodiments provide a non-transitory computer-readable storage medium having stored thereon program instructions that, when executed by one or more processors, are operable to cause the one or more processors to perform operations including acquiring a plurality of videos of a plurality of objects in an environment, determining one or more key points for each of the plurality of objects, recognizing activity information based on the one or more key points, and calculating workflow information based on the activity information.
[0008] Further with respect to the computer-readable storage medium, in some embodiments, the environment is a surgical room. In some embodiments, the plurality of videos are captured by at least two video cameras. In some embodiments, the activity information includes pose information. In some embodiments, the instructions, when executed, are further operable to cause the one or more processors to perform operations including: recognizing one or more objects that are people in the environment, tracking each person's path in the environment, and identifying one or more activities for each person. In some embodiments, the instructions, when executed, are further operable to cause the one or more processors to perform operations including: recognizing one or more inanimate objects in the environment, tracking the location of each recognized inanimate object in the environment, and associating the one or more inanimate objects with each person. In some embodiments, the workflow information includes surgical workflow information.
[0009] In some embodiments, a method includes obtaining a plurality of videos of a plurality of objects in an environment, determining one or more key points for each object of the plurality of objects, recognizing activity information based on the one or more key points, and calculating workflow information based on the activity information.
[0010] In some embodiments, the environment is a surgical room. In some embodiments, the plurality of videos are captured by at least two video cameras. In some embodiments, the activity information includes pose information. In some embodiments, the method further includes recognizing one or more objects that are people in the environment, tracking each person's path in the environment, and identifying one or more activities for each person. In some embodiments, the method further includes recognizing one or more inanimate objects in the environment, tracking the location of each recognized inanimate object in the environment, and associating the one or more inanimate objects with each person. In some embodiments, the workflow information includes surgical workflow information.
[0011] A further understanding of the nature and advantages of particular implementations disclosed herein may be realized by reference to the remaining portions of the specification and the accompanying drawings. [Brief explanation of the drawings]
[0012] [Figure 1] FIG. 1 is a block diagram of an example environment for recognizing clinical activities using multiple cameras that can be used in implementations described herein. [Figure 2] 1 is an example flow diagram for recognizing clinical activity using multiple cameras, according to some embodiments. [Figure 3] 1 is an example flow diagram for recognizing clinical activity using multiple cameras, according to some embodiments. [Figure 4] FIG. 1 is a block diagram of an example environment for recognizing clinical activities using multiple cameras and overlapping regions that can be used in implementations described herein. [Figure 5A] FIG. 1 is a flow diagram for recognizing clinical activities using a top-down approach that can be used in the implementations described herein. [Figure 5B] FIG. 1 is a flow diagram for recognizing clinical activities using a bottom-up approach that can be used in the implementations described herein. [Figure 6]FIG. 1 is a block diagram of an example environment for recognizing clinical activities that can be used in the implementations described herein. [Figure 7] FIG. 1 is a block diagram of an example user interface used in clinical activity recognition that can be used in the implementations described herein. [Figure 8] FIG. 1 is a block diagram of an example network environment that can be used in the implementations described herein. [Figure 9] FIG. 1 is a block diagram of an example computer system that can be used to implement the methods described herein. DETAILED DESCRIPTION OF THE INVENTION
[0013] Embodiments described herein enable, facilitate, and manage the recognition and monitoring of clinical activities using multiple cameras. In some embodiments, a system acquires multiple videos of multiple objects in an environment. The system determines one or more key points for each of the multiple objects. The system recognizes activity information based on the one or more key points. The system further calculates workflow information based on the activity information.
[0014] Although the embodiments disclosed herein are described in the context of the object or subject being a human, these embodiments may also be applied to other objects, such as animals, mechanical devices, and the like, that are capable of performing various actions in an environment, such as a clinical environment.
[0015] 1 is a block diagram of an example environment 100 for recognizing clinical activities using multiple cameras that can be used in the implementations described herein. As described in further detail herein, the system 102 is a context-aware system that provides better patient care and greater hospital efficiency. In some implementations, the environment 100 includes the system 102 in communication with a client 104 over a network 106. The network 106 can be any suitable communications network, such as a Wi-Fi network, a Bluetooth network, the Internet, or the like.
[0016] In various embodiments, environment 100 can be any environment in which activity involving one or more people and / or one or more objects is recognized, monitored, and tracked. In various embodiments, environment 100 can be any clinical environment. For example, in some embodiments, environment 100 can be an operating room. In other embodiments, environment 100 can be an intensive care unit (ICU), a hospital room, an emergency room, etc.
[0017] The activity area 110 can be the operating area of a surgical suite. In some embodiments, the activity area 110 can be the entire surgical suite. In various embodiments, the system 102, client 104, and network 106 can be local to the environment, remote (e.g., in the cloud), or a combination thereof.
[0018] In various embodiments, video is captured by at least two video cameras. For example, as shown, system 102 monitors activity of object 108 using physical video cameras 112, 114, 116, and 118 that capture video of object 108 within activity area 110 at different angles.
[0019] As described in further detail herein, in various embodiments, an object 108 may represent one or more people. For example, in various scenarios, an object 108 may represent one or more of clinicians, such as a doctor and a nurse, one or more assistants, a patient, etc. In various embodiments, an object 108 may also represent one or more inanimate objects. For example, in various scenarios, an object 108 may represent one or more hospital beds, surgical instruments, surgical tools, etc. An object 108 may also represent multiple people, multiple inanimate objects, or a combination thereof. The specific type of object may vary and depend on the particular implementation. In various embodiments, an object 108 may also be referred to as a subject 108, a person 108, a target user 108, or any inanimate object 108.
[0020] In various embodiments, the system utilizes a vision-based approach that is efficient in that it does not require the subject to have wearable devices. The vision-based approach is also highly scalable for different system configurations. In various embodiments, the system automatically and accurately recognizes activities in clinical environments (e.g., operating rooms, emergency rooms, etc.), enabling understanding of surgical or clinical workflows that are important for optimizing clinical activities. The system performs real-time monitoring of staff and patient activities to enhance patient outcomes and care and reduce staff costs.
[0021] In various embodiments, physical video cameras 112, 114, 116, and 118 are positioned at various locations to capture multiple video and / or still images from different perspectives of the same object, including from different angles and / or distances. The terms camera and video camera can be used interchangeably. These different perspectives make it easier to distinguish the appearance of different objects.
[0022] For ease of explanation, FIG. 1 shows one block each for system 102, client 104, network 106, and activity area 110. Blocks 102, 104, 106, and 110 may represent multiple systems, client devices, networks, and activity areas. Also, any number of people / subjects may be present in a given activity area. For example, in some embodiments, subject 108 may represent one or more different subjects. In other implementations, environment 100 may not have all of the components shown and / or may have other elements, including other types of elements instead of or in addition to the elements shown herein.
[0023] Although the embodiments described herein are performed by system 102, in other embodiments, the implementation of the embodiments described herein may be facilitated by any suitable component or combination of components associated with system 102, or any suitable processor or processors associated with system 102.
[0024] 2 is an example flow diagram for recognizing clinical activities using multiple cameras, according to some embodiments. Referring to both FIGS. 1 and 2, the method begins at block 202, where a system, such as system 102, acquires multiple videos of multiple objects in an environment. In various embodiments, the cameras record the videos, and the videos can be stored in any suitable storage location. In various embodiments, the video sequences are captured from multiple cameras, which can be configured with predetermined camera parameters (including pre-calibrated ones). Such camera parameters can include one or more intrinsic matrices, one or more extrinsic matrices, etc.
[0025] In block 204, the system determines one or more key points for each object in the environment. In various embodiments, the system utilizes a vision-based technique using multiple cameras, which is beneficial in that it does not require wearable devices. The system is also highly scalable for different system configurations.
[0026] In various embodiments, a system provides a skeletal-based activity recognition technique that helps staff better recognize various situations during surgery to improve the efficiency of clinical procedures. For example, in various embodiments, the system can use keypoints in performing pose estimation. For example, when a staff member, such as a doctor, nurse, or other clinician, escorts a patient into an operating room, the system identifies keypoints such as major body parts (e.g., head, torso, legs, arms, etc.), joints (neck, shoulders, elbows, wrists, knees, ankles, etc.), equipment, bed, etc.
[0027] In various embodiments, the system can use artificial intelligence (AI), deep machine learning, and computer vision techniques to detect, identify, and recognize key points from the video and associate each key point with an object (e.g., a staff member's head, a patient's torso, etc.). The system uses these techniques to identify, classify, measure, monitor, and track the movements and paths of the key points. As mentioned above, no handcrafted features or wearable devices are required. The use of multiple cameras makes the system robust to environmental changes. The use of multiple cameras also reduces object occlusions in complex and crowded environments.
[0028] In block 206, the system recognizes activity information based on one or more key points. In various embodiments, the activity information includes pause information. For example, the system can detect and recognize that a clinician is walking a patient to a bed. The system can detect and recognize that a patient is lying down. The system can then detect and recognize that a person, such as a staff member, is pushing the bed on which the patient is lying. The system can detect whether the person is moving the bed while the patient is present in the bed. As described in more detail herein, the system can also detect when one or more people are entering or leaving a room and / or when equipment and / or supplies are being brought into and moved around the room.
[0029] In various embodiments, the system may utilize AI, deep machine learning, and computer vision techniques to recognize specific activity information, such as movements associated with walking, movements associated with carrying a device, movements associated with operating a device, and movements associated with taking notes. In various embodiments, the system may also utilize AI, deep machine learning, and computer vision techniques to associate activity information, including a subject's body position and movements, with specific objects. The system may utilize these and other techniques to distinguish between different objects. As illustrated herein, the system utilizes multiple cameras to capture video of different objects in a given environment at different angles and distances relative to the objects.
[0030] Such activity awareness enables understanding of surgical and / or other clinical workflows, which is important for optimizing hospital utilization. Real-time monitoring of activity within the clinical environment enhances patient outcomes and care and reduces staff costs.
[0031] In block 208, the system calculates workflow information based on the activity information. In various embodiments, the workflow information includes activity information for one or more objects (e.g., people, equipment, etc.) in the environment. For example, the workflow information may represent a surgery from start to finish, which may include when each person (e.g., clinician, patient, etc.) enters the room, preparation activities, surgical activities, cleanup activities, etc. The workflow information may also include a timeline and specific activities that occur during the timeline. Further example embodiments related to workflow information are described in further detail herein, for example, in connection with FIG. 7.
[0032] As described herein, the system recognizes one or more objects that are people within an environment and also identifies one or more activities of each person. In various embodiments, the system also tracks the path of each person within the environment. For example, the system can detect specific movements, including the path of a person as they enter or exit a given room or space. For example, the system can detect specific movements, including the path of a person as they walk within a given environment (e.g., an operating room). For example, the system can track the path taken by a staff member as they move a patient to a specific location and / or orientation within a given environment.
[0033] In various embodiments, the system recognizes one or more inanimate objects in the environment. The system tracks the location of each inanimate object recognized in the environment. For example, the system may detect trays of surgical tools, beds on which patients reside, and various other equipment, along with their locations and orientations within the environment (e.g., an operating room). The system also associates one or more inanimate objects with each person. For example, if a given person (e.g., a clinician, assistant, or other staff member) handles a particular inanimate object (e.g., a tray of surgical tools), the system may associate the inanimate object with the particular person (e.g., the assistant).
[0034] In various embodiments, the workflow information includes surgical workflow information. For example, the system can generate a list of objects (e.g., one or more people present in an environment, one or more inanimate objects entering or exiting the environment, one or more inanimate objects, etc.). The system can then determine actions associated with each object as described herein. For example, the system can detect, recognize, and store information related to a nurse escorting a patient into an operating room, a nurse assisting a patient to lie down, a doctor entering a room, a team of personnel preparing the patient and equipment for surgery, a doctor performing surgery including various surgical procedures, post-surgery cleanup, etc. These are just a few examples, and the specific actions associated will vary depending on the particular implementation.
[0035] In various embodiments, the system also organizes the actions of the workflow chronologically and stores timing information (e.g., timestamps, etc.) associated with each action. The workflow information can include a list of detected objects, the relationships between various different objects, and a timeline of different actions. Thus, the system determines start and stop times for the overall procedure. The system also determines start and stop times for stages within the overall procedure. Such stages can include, for example, a setup stage, a surgical stage, a reporting stage, a cleanup stage, etc.
[0036] In various embodiments, such workflow information is useful for personnel (e.g., administrators, doctors, nurses, etc.) to analyze actions taken within a workflow. The system can determine whether each action is appropriate or inappropriate, normal or abnormal, quick or time-consuming, etc. The system can flag certain activities that appear to be inappropriate, abnormal, time-consuming, etc.
[0037] In various embodiments, the system can generate a report presenting the workflow information. The system can calculate one or more recommendations based on the workflow information. The recommendations can be based on flags associated with particular activities, as described herein. For example, the system can determine that a particular setup procedure takes an unusually long time compared to other similar setup procedures. The system can flag the action and / or the person associated with the action in the report. In various embodiments, a user or staff member can verify such determination and / or modify the workflow for greater efficiency and / or effectiveness. Accordingly, the embodiments described herein are advantageous in that the generated workflow information can be used to improve the timing of different actions, understand complex situations, and the like. Further example embodiments related to reports are described in further detail herein, for example, in connection with FIG. 7.
[0038] 3 is an example flow diagram for recognizing clinical activity using multiple cameras, according to some embodiments. Referring to both FIGS. 1 and 3, the method begins at block 302, where a system, such as system 102, acquires video from multiple video cameras. As described herein, the multiple cameras record video and the videos can be stored in any suitable storage location. In various embodiments, video sequences are captured from multiple cameras, which can be configured with predetermined camera parameters (including pre-calibrated ones). Such camera parameters can include one or more intrinsic matrices, one or more extrinsic matrices, etc.
[0039] In block 304, the system performs pose estimation. Such pose estimation may include pose information for one or more people, including staff and patients. Such pose estimation may be performed using any suitable multi-person pose estimator or keypoint detector (e.g., alpha pose estimator, high-resolution network, etc.).
[0040] In block 306, the system performs data fusion using multiple cameras. Robust and accurate data fusion from multiple cameras can handle complex and crowded environments. In various embodiments, data fusion is the process of relating or fusing a person's pose from one camera with the same person's pose from other cameras. After data fusion, the system reconstructs the 3D poses of all objects (e.g., staff, patients, etc.) in the virtual 3D space given the multiple 2D corresponding poses.
[0041] In various embodiments, multiple cameras allow the system to address objects with self-occlusion and inter-object occlusion. For example, significant self-occlusion and inter-object occlusion can result from other people or large clinical equipment partially or completely blocking a given object from a given camera.
[0042] Multiple cameras simplify the monitoring task by providing more views of the object being monitored. The use of multiple cameras provides distinguishable appearance information, allowing the system to recognize faces even when they are covered by masks and / or when staff and patients are wearing similar clothing.
[0043] The system recognizes clinical activities in block 308. In various embodiments, the system may utilize a general skeleton-based activity classifier, which may include graphics core next (GCN) methods, recurrent neural network (RNN) methods, etc.
[0044] At block 310, the system generates workflow information including clinical activities. In various embodiments, the workflow information can include the paths of objects (e.g., staff, patients, inanimate objects, etc.) and the activities of such objects (e.g., staff, patients, etc.). For example, in some embodiments, the system can identify and recognize the possibility that one object (e.g., staff, etc.) is guiding another object (e.g., patient, etc.) to an operating room. Such information can further be used for many applications in the medical field, such as medical monitoring, improving operating room efficiency, etc. Accordingly, the system automatically recognizes staff, patients, and various objects in the environment, identifies their activities and movements, and monitors and tracks their paths.
[0045] 4 is a block diagram of an example environment 400 for recognizing clinical activity using multiple cameras and overlapping regions that can be used in the implementations described herein. Environment 400 includes cameras 402, 404, and 406. In various embodiments, cameras 402-406 can be located in different locations.
[0046] In various embodiments, cameras 402-406 can be positioned at different locations such that their fields of view overlap. As shown, the fields of view of cameras 402, 404, and 406 overlap at overlap region 408. When a given object or objects (e.g., staff, patients, etc.) are positioned in overlap region 408, each of cameras 402, 404, and 406 can capture footage of the given object or objects.
[0047] In various embodiments, cameras 402-406 are configured and pre-calibrated to avoid occlusion and enable 3D reconstruction of objects in the environment. In various embodiments, objects used for calibration are simultaneously visible to all cameras. Although three cameras are shown, there can be any number of cameras in environment 400. The specific number of cameras can depend on the particular environment. In various embodiments, the system uses cameras 402-406 to monitor objects, such as floor tiles, to calibrate patterns in the environment. Alternative camera calibration methods can also be used, including the commonly used checkerboard pattern or the use of red-green-blue-depth (RGB-D) cameras.
[0048] 5A and 5B are flow diagrams for two-dimensional (2D) pose estimation of multiple people in a clinical environment. The embodiments described herein identify and locate the body joints of all people in a given image to estimate the pose of the multiple people. As described below in connection with FIGS. 5A and 5B, the embodiments can include top-down and bottom-up approaches.
[0049] 5A is a flow diagram for recognizing clinical activity using a top-down approach that can be used in the implementations described herein. Referring to both FIG. 1 and FIG. 5A, the method begins at block 502 where a system, such as system 102, samples an image.
[0050] The system detects people in block 504. The system can utilize a general object detector to detect staff (e.g., clinicians, assistants, etc.) and patients.
[0051] The system estimates keypoints in block 506. The system uses a keypoint detector to estimate keypoints such as the head, limbs, joints, etc. of each person.
[0052] 5B is a flow diagram for recognizing clinical activity using a bottom-up approach that can be used in the implementations described herein. Referring to both FIG. 1 and FIG. 5B, the method begins at block 512 where a system, such as system 102, samples an image.
[0053] The system estimates keypoints in block 514. As shown herein, the system uses a keypoint detector to estimate keypoints such as the head, limbs, joints, etc. of each person.
[0054] In block 516, the system associates the keypoints, for example, the system associates the keypoints with poses and estimates the 2D pose by connecting the associated keypoints.
[0055] In some embodiments, the system can achieve further gains by tracking people and keypoints in image space, refining regions of interest, removing redundant pose(s) with non-maximum suppression, and using enhanced heatmap decoding to enhance keypoint detection.
[0056] 6 is a block diagram of an example environment 600 for recognizing clinical activities that can be used in the implementations described herein. Shown are cameras 602 and 604 capturing video footage of objects or subjects 606 and 608. Objects 606 and 608 can be, for example, personnel in an operating room, or personnel and a patient in an operating room, etc.
[0057] In various embodiments, the system performs data fusion and clinical action recognition, including skeleton-based activity recognition. As described above, in various embodiments, data fusion is the process of relating or fusing a person's pose from one camera with the same person's pose from other cameras. After data fusion, the system reconstructs the 3D poses of all objects (e.g., staff, patients, etc.) in the virtual 3D space given the multiple 2D corresponding poses.
[0058] The system recognizes each staff and patient action based on their skeletal pose. Such actions can include standing, walking, crouching, sitting, etc. The system can utilize a behavior classifier to recognize such actions. Compared to RGB images or depth maps, the system's process is more robust to visual noise, such as background objects and irrelevant objects (e.g., clothing textures). Another approach is to recognize actions directly from images or depth maps. In some embodiments, the system can achieve further gains by tracking poses in the reconstructed 3D space and extracting skeletal features from both spatial and temporal space.
[0059] 7 is a block diagram of an example user interface 700 used in clinical activity recognition that can be used in the implementations described herein. The surgical workflow analysis shows workflow information related to three objects or subjects. In this particular example embodiment, the workflow information is related to two staff members (denoted as Nurse1 and Nurse2) and one assistant (denoted as Asst1). The number of objects or subjects can vary and depends on the particular implementation. For example, there can be workflow information related to patients, clinical and / or surgical equipment, tools, and / or supplies, etc.
[0060] In this example embodiment, the surgical workflow analysis relates to equipment delivery. As shown, equipment delivery takes 60 minutes. For example, one staff member, Nurse 1, takes 10 minutes to deliver an energy device and 50 minutes to deliver an endoscope. Another staff member, Nurse 2, takes 20 minutes to deliver certain tools and 40 minutes to deliver medical supplies. An assistant, Asst 1, takes 20 minutes to move equipment out of the operating room (OR), 20 minutes to deliver an ultrasound machine, and 20 minutes to set up the endoscope. Although three objects or subjects are shown, Nurse 1, Nurse 2, and Asst 1, any number of objects may be shown in user interface 700.
[0061] In various embodiments, as described herein, a system recognizes, monitors, and tracks various objects, including people and inanimate objects. The system identifies individual actions performed by each person. These actions can include movements such as those shown in Figure 6. The actions can also include actions performed by each person with respect to inanimate objects, such as clinical and / or surgical equipment, tools, and / or supplies.
[0062] The embodiments described herein have a variety of applications. Such applications may include, for example, analysis of clinical staff and patient journey information and activities (e.g., walking, standing, etc.). Other applications may include intelligent surgical workflow analysis, robotic-assisted surgery, improved efficiency and optimization of operating rooms, medical monitoring, and improved patient safety.
[0063] The embodiments described herein provide various advantages. For example, the system recognizes and analyzes human activities and behaviors in clinical environments (such as operating rooms, ICUs, patient rooms, and emergency rooms). This enables automated monitoring of hospital operations, including efficiency understanding, analysis, optimization, and abnormal behavior alerts. Furthermore, by leveraging people's pose skeletons, the embodiments utilize a deep learning-based framework for multi-camera, multi-person activity recognition without the need for wearable devices or specific poses required by many existing motion capture systems.
[0064] 8 is a block diagram of an example network environment 800 that can be used in some implementations described herein. In some implementations, the network environment 800 includes a system 802 that includes a server device 804 and a database 806. For example, the system 802 can be used to implement the system 102 of FIG. 1 and to perform the embodiments described herein. The network environment 800 also includes client devices 810, 820, 830, and 840 that can communicate with the system 802 and / or can communicate with each other directly or through the system 802. The network environment 800 also includes a network 850 that enables the system 802 and the client devices 810, 820, 830, and 840 to communicate. The network 850 can be any suitable communication network, such as a Wi-Fi network, a Bluetooth network, the Internet, etc.
[0065] For ease of explanation, Figure 8 shows one block each for system 802, server device 804, and network database 806, and four blocks for client devices 810, 820, 830, and 840. Blocks 802, 804, and 806 may represent multiple systems, server devices, and network databases. Also, there may be any number of client devices. In other implementations, environment 800 may not have all of the components shown and / or may have other elements, including other types of elements instead of or in addition to the elements shown herein.
[0066] Although the embodiments described herein are performed by server device 804 of system 802, in other embodiments, the implementation of the embodiments described herein may be facilitated by any suitable component or combination of components associated with system 802, or any suitable processor or processors associated with system 802.
[0067] In various embodiments described herein, the processor of system 802 and / or the processor of any of client devices 810, 820, 830, and 840 causes elements (e.g., information, etc.) described herein to be displayed within a user interface on one or more display screens.
[0068] FIG. 9 is a block diagram of an example computer system 900 that can be used in some implementations described herein. For example, computer system 900 can be used to implement server device 804 of FIG. 8 and / or system 102 of FIG. 1, as well as to perform the embodiments described herein. In some implementations, computer system 900 can include a processor 902, an operating system 904, memory 906, and an input / output (I / O) interface 908. In various implementations, processor 902 can be used to implement the various functions and features described herein and to perform implementations of the methods described herein. Although processor 902 is described as performing the implementations described herein, the described steps can also be performed by any suitable component or combination of components of computer system 900, or any suitable processor or processors associated with computer system 900 or any suitable system. The implementations described herein can be performed on a user device, a server, or a combination thereof.
[0069] Computer system 900 includes a software application 910, which may be stored on memory 906 or any other suitable storage location or computer-readable medium. The software application 910 provides instructions that enable processor 902 to perform the implementations and other functions described herein. The software application may also include engines, such as a network engine, that perform various functions related to one or more networks and network communications. The components of computer system 900 may be implemented by one or more processors, or any combination of hardware devices, as well as any combination of hardware, software, firmware, etc.
[0070] 9 illustrates one block for each of processor 902, operating system 904, memory 906, I / O interfaces 908, and software applications 910. These blocks 902, 904, 906, 908, and 910 may represent multiple processors, operating systems, memories, I / O interfaces, and software applications. In various implementations, computer system 900 may not have all of the components illustrated and / or may have other elements, including other types of elements instead of or in addition to the elements illustrated herein.
[0071] Although described with respect to specific embodiments, these specific embodiments are illustrative only and not limiting, and the concepts illustrated in these examples may be applied to other examples and implementations.
[0072] In various implementations, software for execution by one or more processors is encoded on one or more non-transitory computer-readable media, which, when executed by the one or more processors, performs the implementation and other functions described herein.
[0073] The routines of particular embodiments may be implemented using any suitable programming language, including C, C++, Java, assembly language, etc. Different programming techniques may be used, such as procedural or object-oriented. The routines may be executed on a single processing unit or on multiple processors. While steps, operations, or computations may be shown in a particular order, this order may be changed in different particular embodiments. In some particular embodiments, multiple steps shown herein as sequential may be executed simultaneously.
[0074] Certain embodiments may be implemented in a non-transitory computer-readable storage medium (also referred to as a machine-readable storage medium) used by or connected to an instruction execution system, apparatus, or device. Certain embodiments may also be implemented in the form of control logic in software or hardware, or a combination thereof. The control logic, when executed by one or more processors, may perform the implementations and other functions described herein. For example, tangible media, such as hardware storage devices, may be used to store the control logic, which may include executable instructions.
[0075] Certain embodiments may be implemented using programmable general-purpose digital computers and / or using application specific integrated circuits, programmable logic devices, field programmable gate arrays, optical, chemical, biological, quantum, or nanoengineered systems, components, and mechanisms. In general, the functionality of certain embodiments may be achieved by any means known in the art. Distributed networked systems, components, and / or circuits may also be used. Communication or transfer of data may be by wire, wireless, or any other means.
[0076] A "processor" may include any suitable hardware and / or software system, mechanism, or component that processes data, signals, or other information. A processor may include a general-purpose central processing unit, multiple processing units, a system having dedicated circuitry or other systems for implementing functions. Processing need not be limited to a geographic location or have time limitations. For example, a processor may perform its functions in "real time," "offline," "batch mode," etc. Portions of the processing may also be performed by different (or the same) processing systems at different times and in different locations. A computer may be any processor in communication with a memory. Memory may be any suitable data storage, memory, and / or non-transitory computer-readable storage medium, including electronic storage such as random access memory (RAM), read-only memory (ROM), magnetic storage (such as a hard disk drive), flash, optical storage (such as a CD or DVD), magnetic or optical disk, or other tangible medium suitable for storing instructions (e.g., program or software instructions) to be executed by a processor. For example, tangible media such as hardware storage may be used to store control logic, which may include executable instructions. The instructions may also be provided in or as electrical signals, such as in the form of software as a service (SaaS) delivered from a server (eg, a distributed system and / or a cloud computing system).
[0077] It will also be understood that one or more of the elements shown in the drawings / figures may be implemented in a more separate or integrated manner, or may be removed or inoperative in some cases, when useful according to a particular application. It is also within the spirit and scope of the present invention to implement a program or code storable on a machine-readable medium that enables a computer to perform any of the methods described above.
[0078] As used throughout this specification and the claims that follow, the terms "a" and "the" include their plural references unless the context clearly indicates otherwise. Also, as used throughout this specification and the claims that follow, the meaning of "in" includes the meanings "in" and "on," unless the context clearly indicates otherwise.
[0079] While specific embodiments have been described herein, it will be understood that the above disclosure is susceptible to modification, various changes, and substitutions, and that in some instances, some features of a specific embodiment may be used without the corresponding use of other features without departing from the scope and spirit of the described embodiment. Accordingly, many modifications may be made to adapt a particular situation or material to the basic scope and spirit. [Explanation of symbols]
[0080] 100 Environment 102 System 104 Client 106 Network 108 objects 110 Activity Area 112~118 Video camera
Claims
1. 1. A system comprising: one or more processors; one or more non-transitory computer-readable storage media having stored thereon program instructions for execution by the one or more processors; wherein the program instructions, when executed, acquiring a plurality of videos of a plurality of objects in an operating room; determining one or more key points for each object of the plurality of objects; Recognizing one or more objects that are people in the operating room; tracking the path of the one or more key points of each person in the operating room based on skeleton-based activity recognition; identifying one or more activities for each person based on a skeletal posture of each person to determine pose information; Recognizing one or more inanimate objects within the operating room; tracking the location of each inanimate object recognized within the operating room; associating the one or more inanimate objects with each person in the operating room; Recognizing activity information of one or more objects that are people in the operating room and one or more objects that are inanimate objects in the operating room based on the tracking of the positions of the one or more inanimate objects and the tracking of the positions of the one or more key points of each person in the operating room and the pose information; identifying one or more activities of each person in the operating room based on the movement of each person in the operating room and each inanimate object associated with each person in the operating room; identifying each person in the operating room as a patient or staff member; determining whether a particular staff member moved the patient from the first location to the second location based on the particular staff member moving an inanimate object associated with the patient from the first location to the second location; and operable to cause the one or more processors to perform operations including A system characterized by:
2. the plurality of videos are captured by at least two video cameras; The system of claim 1 .
3. A non-transitory computer-readable storage medium having stored thereon program instructions that, when executed by one or more processors, acquiring a plurality of videos of a plurality of objects in an operating room; determining one or more key points for each object of the plurality of objects; Recognizing one or more objects that are people in the operating room; tracking the path of the one or more key points of each person in the operating room based on skeleton-based activity recognition; identifying one or more activities for each person based on a skeletal posture of each person to determine pose information; Recognizing one or more inanimate objects within the operating room; tracking the location of each inanimate object recognized within the operating room; associating the one or more inanimate objects with each person in the operating room; Recognizing activity information of one or more objects that are people in the operating room and one or more objects that are inanimate objects in the operating room based on the tracking of the positions of the one or more inanimate objects and the tracking of the positions of the one or more key points of each person in the operating room and the pose information; identifying one or more activities of each person in the operating room based on the movement of each person in the operating room and each inanimate object associated with each person in the operating room; identifying each person in the operating room as a patient or staff member; determining whether a particular staff member moved the patient from the first location to the second location based on the particular staff member moving an inanimate object associated with the patient from the first location to the second location; 10. A computer-readable storage medium operable to cause the one or more processors to perform operations including:
4. the plurality of videos are captured by at least two video cameras; The computer-readable storage medium of claim 3 .
5. 1. A computer-implemented method comprising: acquiring a plurality of videos of a plurality of objects in an operating room; determining one or more key points for each object of the plurality of objects; Recognizing one or more objects that are people in the operating room; tracking the path of the one or more key points of each person in the operating room based on skeleton-based activity recognition; identifying one or more activities for each person based on a skeletal posture of each person to determine pose information; Recognizing one or more inanimate objects within the operating room; tracking the location of each inanimate object recognized within the operating room; associating the one or more inanimate objects with each person in the operating room; Recognizing activity information based on tracking the positions of the one or more inanimate objects and tracking the positions of the one or more key points of each person in the operating room and the pose information; identifying one or more activities of each person in the operating room based on the movement of each person in the operating room and each inanimate object associated with each person in the operating room; identifying each person in the operating room as a patient or staff member; determining whether a particular staff member moved the patient from the first location to the second location based on the particular staff member moving an inanimate object associated with the patient from the first location to the second location; A method comprising:
6. the plurality of videos are captured by at least two video cameras; The method of claim 5.
Citation Information
Patent Citations
System and method for protocol adherence
US20120154582A1
Workflow assistant for image guided procedures
US20190090954A1