Pose Reconstruction by Tracking for Video Analysis
The described system addresses the inefficiencies in analyzing human movements by tracking subjects across multiple cameras and reconstructing 3D models, enabling accurate and automated recognition of clinical activities in dynamic environments.
Patent Information
- Application Number
- JP2023571174
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-06-10
- Filing Date
- 2022-05-27
- Publication Date
- 2025-06-30
- Estimated Expiration
- 2042-05-27
AI Technical Summary
Existing camera systems struggle with efficiently analyzing human movements in videos, particularly in environments with changing camera positions and angles, leading to unreliable and time-consuming manual monitoring.
A system that uses one or more processors to obtain videos of a subject performing actions, track the subject across multiple cameras, and reconstruct a three-dimensional (3D) model based on the video analysis, enabling efficient pose reconstruction and action recognition.
The system enables automatic and accurate recognition of clinical activities in environments like operating rooms, reducing manual effort and improving efficiency by providing real-time monitoring and 3D model reconstruction.
Smart Images

Figure 0007700272000001 
Figure 0007700272000002 
Figure 0007700272000003
Abstract
Description
Technical Field
[0001] [Cross - Reference to Related Applications] This application claims priority based on U.S. Patent Application No. 17 / 344,734, entitled "POSE RECONSTRUCTION BY TRACKING FOR VIDEO ANALYSIS" (Client Reference No.: SYP339216US01), filed on June 10, 2021, and this document is incorporated herein by reference as if the entire text thereof were set forth herein for all purposes.
[0002] This application is related to U.S. Patent Application Serial No. 17 / 344,730, entitled "CLINICAL ACTIVITY RECOGNITION WITH MULTIPLE CAMERAS" (SYP339214US01), filed on June 10, 2021, and this document is incorporated herein by reference as if the entire text thereof were set forth herein for all purposes.
Background Art
[0003] Some camera systems can capture videos of people, analyze their movements, and generate metadata images or video datasets. To identify human actions captured by the system's camera videos, a person needs to manually review the videos. Manual monitoring and event reporting are unreliable and may take a significant amount of time, especially when the position and angle of the video cameras change and cannot provide sufficient coverage. Using multiple cameras in a controlled environment is also possible. However, the variations in subjects, movements, and backgrounds may still be significantly limited.
Summary of the Invention
Means for Solving the Problems
[0004] Embodiments generally relate to pose reconstruction by tracking for video analysis. In some embodiments, a system includes one or more processors and logic encoded on one or more non-transitory computer-readable storage media for execution by the one or more processors. The logic, when executed, causes the one or more processors to perform operations including: obtaining a plurality of videos of at least one subject performing at least one action in an environment; tracking the at least one subject across at least two cameras; and reconstructing a three-dimensional (3D) model of the at least one subject based on the plurality of videos and the tracking of the at least one subject.
[0005] Regarding the system further, in some embodiments, the plurality of videos obtained are two-dimensional (2D) videos. In some embodiments, the environment is an operating room. In some embodiments, the logic is further operable, when executed, to cause the one or more processors to perform operations including determining one or more keypoints of the at least one subject. In some embodiments, the logic is further operable, when executed, to cause the one or more processors to perform operations including determining pose information associated with the at least one subject. In some embodiments, the logic is operable, when executed, to cause the one or more processors to perform operations including determining pose information associated with the at least one subject based on triangulation. In some embodiments, the logic is further operable, when executed, to cause the one or more processors to perform operations including reconstructing a 3D model of the at least one subject based on the plurality of videos, where the plurality of videos are two-dimensional (2D) videos.
[0006] In some embodiments, a non-transitory computer-readable storage medium having program instructions is provided. The instructions, when executed by one or more processors, cause the one or more processors to perform operations including: obtaining a plurality of videos of at least one subject performing at least one action in an environment; tracking at least one subject across at least two cameras; and reconstructing a three-dimensional (3D) model of at least one subject based on the plurality of videos and the tracking of the at least one subject.
[0007] Regarding the computer-readable storage medium further, in some embodiments, the plurality of videos obtained are two-dimensional (2D) videos. In some embodiments, the environment is an operating room. In some embodiments, the logic is further operable to cause the one or more processors to perform operations including determining one or more key points of at least one subject at runtime. In some embodiments, the logic is further operable to cause the one or more processors to perform operations including determining pose information related to at least one subject at runtime. In some embodiments, the logic is operable to cause the one or more processors to perform operations including determining pose information related to at least one subject based on triangulation at runtime. In some embodiments, the logic is further operable to cause the one or more processors to perform operations including reconstructing a 3D model of at least one subject based on the plurality of videos, where the plurality of videos are two-dimensional (2D) videos.
[0008] In some embodiments, a method includes: obtaining a plurality of videos of at least one subject performing at least one action in an environment; tracking at least one subject across at least two cameras; and reconstructing a three-dimensional (3D) model of at least one subject based on the plurality of videos and the tracking of the at least one subject.
[0009] Regarding the method, in some embodiments, the plurality of videos obtained are two-dimensional (2D) videos. In some embodiments, the environment is an operating room. In some embodiments, the logic is further operable to cause one or more processors to perform operations including determining one or more key points of at least one subject at runtime. In some embodiments, the logic is further operable to cause one or more processors to perform operations including determining pose information related to at least one subject at runtime. In some embodiments, the logic is operable to cause one or more processors to perform operations including determining pose information related to at least one subject based on triangulation at runtime. In some embodiments, the logic is further operable to cause one or more processors to perform operations including reconstructing a 3D model of at least one subject based on a plurality of videos at runtime, and the plurality of videos are two-dimensional (2D) videos.
[0010] By referring to the remainder of the specification and the attached drawings, the characteristics and advantages of the specific implementations disclosed herein can be further understood.
Brief Description of the Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
[0012] The embodiments described herein enable, facilitate, and manage pose reconstruction by tracking for video analysis. In various embodiments, the system obtains video of at least one subject performing at least one action within an environment. The system tracks at least one subject across at least two cameras. The system further reconstructs a three-dimensional (3D) model of at least one subject based on the video and the tracking of the at least one subject.
[0013] FIG. 1 is a block diagram of an example environment 100 for recognizing clinical activities using multiple cameras that can be used in the implementations described herein. As will be described in more detail herein, system 102 is a context-aware system that provides better patient treatment and higher hospital efficiency. In some implementations, environment 100 includes a system 102 that communicates with a client 104 via a network 106. Network 106 can be any suitable communication network such as a Wi-Fi network, a Bluetooth network, the Internet, or the like.
[0014] In various embodiments, environment 100 can be any environment in which activities involving one or more people and / or one or more objects are recognized, monitored, and tracked. In various embodiments, environment 100 can be any clinical environment. For example, in some embodiments, environment 100 can be an operating room. In other embodiments, environment 100 can be an intensive care unit (ICU), a hospital room, an emergency room, etc.
[0015] Activity area 110 can be the surgical area of an operating room. In some embodiments, activity area 110 can be the entire operating room. In various embodiments, system 102, client 104, and network 106 can be local to the environment, remote (e.g., in the cloud), or a combination thereof.
[0016] In various embodiments, video is captured by at least two video cameras. For example, as illustrated, system 102 monitors the activities of object 108 using physical video cameras 112, 114, 116, and 118 that capture video of object 108 within activity area 110 from different angles.
[0017] As will be described in more detail herein, in various embodiments, object 108 can represent one or more people. For example, in various scenarios, object 108 can represent one or more of clinicians such as doctors and nurses, one or more assistants, one or more of patients, etc. In various embodiments, object 108 can also represent one or more inanimate objects. For example, in various scenarios, object 108 can represent one or more hospital beds, surgical instruments, surgical tools, etc. Also, object 108 can represent a plurality of people or a plurality of inanimate objects, or a combination thereof. The specific type of object can vary and depends on the specific implementation. In various embodiments, object 108 can also be referred to as subject 108, person 108, target user 108, or any inanimate object 108.
[0018] In various embodiments, the system utilizes an efficient vision-based approach in that the subject does not need to have a wearable device. Also, the vision-based approach is highly scalable for different settings of the system. In various embodiments, the system enables an understanding of surgical or clinical workflows important for the optimization of clinical activities by automatically and accurately recognizing activities in a clinical environment (e.g., operating room, emergency room, etc.). The system performs real-time monitoring of staff and patient activities to enhance patient outcomes and care and reduce staff costs.
[0019] In various embodiments, physical video cameras 112, 114, 116, and 118 are placed at various locations to capture multiple video images and / or still images from different viewpoints of the same object, including different angles and / or different distances. The terms camera and video camera can be used synonymously. These different viewpoints make it easier to distinguish the appearances of different objects.
[0020] To facilitate the description, FIG. 1 shows one block for each of system 102, client 104, network 106, and activity area 110. Blocks 102, 104, 106, and 110 can also represent multiple systems, client devices, networks, and activity areas. Also, any number of people / subjects can be present in a given activity area. For example, in some embodiments, subject 108 can represent one or two or more different subjects. In other implementations, environment 100 can also not have all of the components shown in the figure and / or can have other elements including other types of elements instead of or in addition to the elements shown herein.
[0021] Although the embodiments described herein are executed by system 102, in other embodiments, the execution of the embodiments described herein can be facilitated by any suitable component or combination of components associated with system 102, or any suitable one or more processors associated with system 102.
[0022] FIG. 2 is an example flowchart for pose reconstruction by tracking for video analysis, according to some embodiments. Referring to both FIGS. 1 and 2, the method begins at block 202 where a system, such as system 102, acquires a plurality of videos of at least one subject performing at least one action within the environment. In various embodiments, a camera can record the video and store the video in any suitable storage location. In various embodiments, a video sequence is captured from a plurality of cameras that can be configured with predetermined camera parameters (including pre-calibrated ones). Such camera parameters can include one or two or more intrinsic matrices, one or two or more extrinsic matrices, and the like.
[0023] In block 204, system 102 tracks at least one subject across at least two cameras. In various embodiments, the video obtained is a two-dimensional (2D) video. In various embodiments, the system avoids cross-view association ambiguity by processing 2D video information from multiple cameras. Noise and incomplete 2D poses due to occlusion complicate the association of a given pose from different cameras, which may further affect the reconstruction of poses in 3D space. The system can track from camera to camera without losing sight of each individual object by utilizing multiple cameras.
[0024] In various embodiments, the system determines one or more key points of each object or subject tracked via a video camera. The system also determines pose information associated with each object. The system determines pose information based also on each respective key point associated with each object. In various embodiments, the system determines pose information associated with at least one subject based on triangulation. Further embodiments related to key points, pose information, and triangulation are described in more detail herein.
[0025] In block 206, system 102 reconstructs a three-dimensional (3D) model of at least one subject based on the video and the tracking of at least one subject. In various embodiments, the system reconstructs a 3D model of an object or subject based on a video that is a 2D video. The reconstruction of the 3D model can be applied to various areas. For example, such areas can be applied to action understanding in the medical or sports fields, surveillance and security, retail or manufacturing, etc. Specific applications can be diverse and depend on the particular implementation.
[0026] Although steps, operations, or calculations may be shown in a particular order, the order may be changed in a particular implementation. Depending on the particular implementation, other step orders are possible. In some particular implementations, multiple steps shown herein as sequential may be executed simultaneously. Also, some implementations may not have all of the steps and / or may have other steps instead of or in addition to the steps shown herein.
[0027] Figure 3 is an example of a flowchart for reconstructing multi-view poses according to some embodiments. The following details describe a pose reconstruction and tracking framework according to some embodiments. Referring to both FIGS. 1 and 3, the method begins at block 302 where a system such as system 102 obtains camera parameters. In various embodiments, the camera parameters can include the intrinsic and extrinsic matrices of each camera within the system, depending on the configuration of the environment.
[0028] At block 304, system 102 calculates two-dimensional (2D) pose information. In various embodiments, the system can utilize a general keypoint estimator to calculate the 2D pose information and can use either a top-down approach or a bottom-up approach.
[0029] At block 306, system 102 performs matching of the 2D poses. In various embodiments, pose matching maintains and tracks the identity of each target object captured in a consistent video across multiple cameras. In various embodiments, the system can apply one or more metrics to the matching. Examples of metrics can include epipolar constraints, Euclidean distance and algorithm for data association, Hungarian algorithm, and the like.
[0030] In an exemplary scenario, the system can associate 2D poses of the same person across different camera views by using geometric constraints, cycle-consistent constraints, etc. Thus, if a person moves out of the field of view of one camera, the same person is captured in the field of view of another camera within the same environment. In various embodiments, the system can track the movement and pose of a person based on the detection and knowledge of parts of the person, such as limb joints, height, positions of joints and limbs, and the trajectory of the person.
[0031] The embodiments described herein reduce the computation by using pose tracking information in 3D space, as opposed to conventional methods of associating poses frame-by-frame across cameras.
[0032] In block 308, the system 102 obtains back-projected 2D pose information. In various embodiments, the system can obtain the back-projected 2D pose information by projecting 3D pose information from block 310 onto the image plane. In various embodiments, the tracking information from 3D space gives pointers for pose matching in block 306 for the current frame.
[0033] In block 310, the system 102 reconstructs 3D poses. In various embodiments, the system determines the 3D position of the pose based on multiple 2D corresponding poses and triangulation. Embodiments regarding triangulation are described in more detail herein, for example, in relation to FIG. 7.
[0034] Figure 4 is a block diagram of an environment example 400 for recognizing clinical activities using a plurality of cameras and overlapping regions, which can be used in the implementations described in this specification. Environment 400 includes cameras 402, 404, and 406. In various embodiments, cameras 402-406 can be arranged at different positions.
[0035] In various embodiments, cameras 402-406 can be arranged at different positions such that their fields of view overlap. As shown, the fields of view of cameras 402, 404, and 406 overlap in overlapping region 408. When a given one or more objects (e.g., staff, patients, etc.) are placed in overlapping region 408, each of cameras 402, 404, and 406 can capture footage of the given one or more objects.
[0036] In various embodiments, cameras 402-406 are set and pre-calibrated to avoid occlusion and enable 3D reconstruction of objects in the environment. In various embodiments, the objects used for calibration are visible to all cameras simultaneously. Although three cameras are shown, any number of cameras can be present within environment 400. The specific number of cameras can depend on the specific environment. In various embodiments, the system monitors objects such as floor tiles using cameras 402-406 to calibrate patterns within the environment. Another camera calibration method can also be used, including the use of a commonly used checkerboard pattern or a red-green-blue-depth (RGB-D) camera.
[0037] Figure 5 is a block diagram of an environment example 500 for recognizing clinical activities, which can be used in the implementations described in this specification. Cameras 502 and 504 are shown that capture video footage of objects or subjects 506 and 508. Objects 506 and 508 can be, for example, staff in an operating room, or staff and patients in an operating room.
[0038] In various embodiments, the system performs data fusion and clinical action recognition that includes skeleton-based activity recognition. As described above, in various embodiments, data fusion is the process of associating or fusing the pose of a person from one camera with the pose of the same person from another camera. After data fusion, the system reconstructs the 3D poses of all objects (e.g., staff, patients, etc.) in the virtual 3D space given a plurality of 2D corresponding poses.
[0039] The system recognizes the actions of each staff member and patient based on the skeletal poses. Such actions can include standing, walking, crouching, sitting, etc. The system can use an action classifier to recognize such actions. The system's process is robust to visual noise such as background objects and irrelevant objects (e.g., texture of clothing, etc.) compared to RGB images or depth maps. Another way can be to directly recognize actions from the image or depth map. In some embodiments, the system can achieve further gains by tracking the poses in the reconstructed 3D space and extracting skeletal features from both the spatial and temporal spaces.
[0040] FIG. 6 is an example flowchart for determining multi-view poses according to some embodiments. Referring to both FIG. 1 and FIG. 6, the method starts at block 602 where a system such as system 102 obtains the back-projected 2D pose information.
[0041] At block 604, system 102 obtains the estimated poses. The system collects the estimated poses of each object detected within the camera.
[0042] At block 606, system 102 discovers the corresponding poses. Such corresponding poses can include different poses of the same object (e.g., a person) captured by different cameras.
[0043] In block 608, system 102 performs pose matching. For example, the system matches the poses of the same object (e.g., a person) from different cameras. In some embodiments, the system executes a pose matching step when the pose does not match any of the existing tracklets. A tracklet can be defined as a fragment of the trajectory followed by a moving object, constructed by an image recognition system.
[0044] In various embodiments, the system can apply one or more metrics to the matching. Examples of metrics can include epipolar constraints, Euclidean distances and algorithms for data association, the Hungarian algorithm, etc.
[0045] In block 610, system 102 provides a match result. The match result indicates all the poses of each specific object (e.g., a person).
[0046] FIG. 7 is an example flowchart for providing a reconstructed pose according to some embodiments. Referring to both FIG. 1 and FIG. 7, the method starts at block 702 where a system such as system 102 performs 2D pose matching.
[0047] In block 704, system 102 selects a plurality of pairs of views from the 2D poses. In various embodiments, the system obtains each pair from different cameras. In various embodiments, the selection of the plurality of pairs of views can be based on two conditions. In some embodiments, the first condition can be to select a pair of views based on the reprojection error being less than a predetermined threshold. In some embodiments, the second condition can be to select a pair of views based on the confidence score being higher than a predetermined threshold. For example, a high confidence score can be related to less occlusion, and a low confidence score can be related to more occlusion. This selection can be achieved by minimizing the reprojection error and maximizing the confidence score for accurate 3D reconstruction.
[0048] As described below, the method provides a reconstructed pose according to two series of steps. The first series of steps is related to blocks 706, 708, and 710. The system executes these steps if the set of pairs of views is not empty. The second series of steps is related to blocks 712, 714, and 716. The system executes these steps if no pair of views is selected.
[0049] In block 706, system 102 selects two views. In various embodiments, the system selects two views having the highest-ranked confidence score and the lowest-ranked reprojection error. As described below in connection with block 708, the system can perform triangulation using the two views for 3D pose reconstruction.
[0050] In block 708, system 102 performs triangulation. In various embodiments, the system can utilize adaptive triangulation. Triangulation can be used to obtain 3D pose information based on a given 2D correspondence pose in a multi-view framework. In some embodiments, instead of performing reconstruction across all cameras, the system can adaptively select a subset of camera views for 3D pose reconstruction. For example, the system can determine the cameras that capture a given target object in order to minimize calculations. Other cameras that do not capture the given object are unnecessary and thus not used for collecting information about this particular object. By using only the cameras that capture the object, it is ensured that the calculations performed by the system are sufficient but not excessive.
[0051] In block 710, system 102 provides the reconstructed poses. In various embodiments, the system determines the 3D positions of the poses of the same object (e.g., clinician, patient, etc.) based on multiple 2D corresponding poses and triangulation. The system determines the poses from the video feeds of multiple cameras to reconstruct the 3D poses of each object.
[0052] As described above, the second series of steps is associated with blocks 712, 714, and 716. The system performs these steps when a pair of views is not selected.
[0053] In block 712, system 102 performs triangulation. In various embodiments, system 102 performs triangulation in the same manner as step 708 described above.
[0054] In block 714, system 102 integrates the poses. For example, in various embodiments, the system aggregates the poses of each object from different viewpoints of different cameras that capture each object (e.g., clinician, patient, etc.).
[0055] In block 716, system 102 provides a reconstructed pose. In various embodiments, system 102 performs triangulation as in step 710 described above.
[0056] The embodiments described herein provide various advantages. For example, embodiments efficiently estimate the 3D poses of all persons in an environment using a calibrated set of cameras. Embodiments can be built on top of any real-time multi-person 2D pose estimation system, and such embodiments are robust to occlusions that may frequently occur in practical applications.
[0057] The embodiments described herein are simple yet effective in 3D multi-camera multi-target pose reconstruction. Also, the embodiments described herein provide a cost-effective solution for pose matching that functions as an important step for further 3D pose reconstruction.
[0058] FIG. 8 is a block diagram of a network environment example 800 that can be used in some implementations described herein. In some implementations, network environment 800 includes a system 802 that includes a server device 804 and a database 806. For example, system 802 can be used to implement system 102 of FIG. 1 and to execute the embodiments described herein. Network environment 800 also includes client devices 810, 820, 830, and 840 that can communicate with system 802 and / or with each other directly or via system 802. Network environment 800 also includes a network 850 that enables system 802 and client devices 810, 820, 830, and 840 to communicate. Network 850 can be any suitable communication network such as a Wi-Fi network, a Bluetooth network, the Internet, etc.
[0059] To facilitate the description, FIG. 8 shows one block for each of the system 802, the server device 804, and the network database 806, and four blocks for the client devices 810, 820, 830, and 840. The blocks 802, 804, and 806 can also represent a plurality of systems, server devices, and network databases. Also, any number of client devices can exist. In other implementations, the environment 800 may not have all of the components shown in the figure and / or may have other elements including other types of elements instead of or in addition to the elements shown herein.
[0060] Although the embodiments described herein are executed by the server device 804 of the system 802, in other embodiments, the execution of the embodiments described herein can be facilitated by any suitable component or combination of components associated with the system 802, or any suitable one or more processors associated with the system 802.
[0061] In various embodiments described herein, the processor of the system 802 and / or the processors of any of the client devices 810, 820, 830, and 840 cause the elements (e.g., information, etc.) described herein to be displayed within a user interface on one or more display screens.
[0062] FIG. 9 is a block diagram of an example computer system 900 that can be used in some implementations described herein. For example, computer system 900 can be used to implement the server device 804 of FIG. 8 and / or the system 102 of FIG. 1, as well as to execute the embodiments described herein. In some implementations, computer system 900 can include a processor 902, an operating system 904, a memory 906, and an input / output (I / O) interface 908. In various implementations, the processor 902 can be used to implement the various functions and features described herein, as well as to execute the implementation of the methods described herein. Although the processor 902 is described as executing the implementations described herein, the steps described can also be executed by any suitable component or combination of components of computer system 900, or by any suitable one or more processors associated with computer system 900 or any suitable system. The implementations described herein can be executed on a user device, on a server, or in a combination thereof.
[0063] Computer system 900 can include a software application 910 that can be stored on memory 906, or on any other suitable storage location, or on a computer-readable medium. Software application 910 provides instructions that enable processor 902 to execute the implementations described herein and other functions. The software application can also include an engine, such as a network engine, that executes various functions related to one or more networks and network communications. The components of computer system 900 can be implemented by any combination of one or more processors, or any combination of hardware devices, as well as any combination of hardware, software, firmware, and the like.
[0064] To facilitate the description, FIG. 9 shows one block for each of processor 902, operating system 904, memory 906, I / O interface 908, and software application 910. These blocks 902, 904, 906, 908, and 910 can also represent multiple processors, operating systems, memories, I / O interfaces, and software applications. In various implementations, computer system 900 may not have all of the components shown in the figure and / or may have other elements including other types of elements instead of or in addition to the elements shown herein.
[0065] Although the description has been made with respect to specific embodiments, these specific embodiments are merely illustrative and not limiting. The concepts illustrated in these examples can also be applied to other examples and implementations.
[0066] In various implementations, software for execution by one or more processors is encoded on one or more non-transitory computer-readable media. This software, when executed by one or more processors, performs the implementations and other functions described herein.
[0067] For the implementation of the routines of specific embodiments, any suitable programming language can be used, including C, C++, Java, assembly language, etc. Different programming techniques such as procedural or object-oriented can be used. These routines can be executed on a single processing device or multiple processors. Although steps, operations, or calculations may be shown in a specific order, this order can be changed in different specific embodiments. In some specific embodiments, multiple steps shown herein as sequential can also be executed simultaneously.
[0068] Certain embodiments can be implemented on a non-transitory computer-readable storage medium (also referred to as a machine-readable storage medium) used by or connected to an instruction execution system, apparatus, or device. Certain embodiments can also be implemented in the form of control logic in software or hardware or a combination thereof. The control logic, when executed by one or more processors, can perform the implementations and other functions described herein. For example, a tangible medium such as a hardware storage device can be used for storing control logic that can include executable instructions.
[0069] Certain embodiments can be implemented by using a programmable general-purpose digital computer and / or by using application-specific integrated circuits, programmable logic devices, field programmable gate arrays, optical, chemical, biological, quantum, or nanoengineering systems, components, and mechanisms. In general, the functions of certain embodiments can be realized by any means well known in the art. Distributed, networked systems, components, and / or circuits can also be used. The communication or transfer of data can be by wire, wireless, or any other means.
[0070] A "processor" can include any suitable hardware and / or software system, mechanism, or component that processes data, signals, or other information. The processor can include a general-purpose central processing unit, multiple processing units, a dedicated circuit for realizing functions, or a system having other systems. The processing need not be restricted by geographical location or have time limitations. For example, the processor can execute its functions in "real time", "offline", "batch mode", etc. Some of the processing can also be executed by different (or the same) processing systems at different times and in different locations. A computer can be any processor that communicates with a memory. The memory can be any suitable data storage, memory, and / or non-transitory computer-readable storage medium, including an electronic storage device such as random access memory (RAM), read-only memory (ROM), magnetic storage devices (such as hard disk drives), flash memory, optical storage devices (such as CDs or DVDs), magnetic or optical disks, or other tangible media suitable for storing instructions (such as program or software instructions) executed by the processor. For example, a tangible medium such as a hardware storage device can be used to store control logic that can include executable instructions. The instructions can also be provided as electrical signals, for example, included in electrical signals in the form of service-type software (SaaS) distributed from a server (such as a distributed system and / or a cloud computing system).
[0071] Also, when useful according to a specific application, it will be understood that one or more of the elements shown in the drawings / figures can be implemented in a more separated or integrated form, or in some cases removed or made inoperable. Implementing a program or code that can be stored on a machine-readable medium and enables a computer to execute any of the above-described methods is also included in the spirit and scope of the present invention.
[0072] As used throughout this specification and the following claims, the articles "a" and "the" include plural referents unless the context clearly dictates otherwise. Also, as used throughout this specification and the following claims, the meaning of "in" includes the meaning of "in" and "on" unless the context clearly dictates otherwise.
[0073] Having described specific embodiments hereinabove, it is intended that the above disclosure be considered as an example only, and that numerous modifications, various changes and substitutions be possible, and that in some instances, some features of a particular embodiment may be used without the use of corresponding other features, without departing from the scope and spirit described. Accordingly, many modifications may be made to adapt a particular situation or material to the basic scope and spirit.
Description of Reference Numerals
[0074] 100 Environment 102 System 104 Client 106 Network 108 Subject 110 Activity Area 112 - 118 Video Camera
Claims
Claim 1. One or more processors, Program instructions stored in one or more computer-readable storage media for execution by the one or more processors, Comprising, the program instructions, when executed, Obtaining, from at least two cameras each having a perspective selected based on different conditions, a plurality of videos of at least one subject performing at least one action in an environment; Tracking the at least one subject based on the plurality of videos obtained by the at least two cameras; Determining information regarding the posture of the at least one subject detected by the at least two cameras; Determining information regarding corresponding postures from information regarding different postures of the same at least one subject detected by the at least two cameras; Matching information regarding corresponding postures of the same at least one subject; Integrating information regarding the postures of the at least one subject that have been matched from different perspectives of adaptively selected different combinations of the at least two cameras; Generating, from a two-dimensional (2D) image, a three-dimensional (3D) model to be processed by the one or more processors when recognizing the action of the at least one subject based on the integrated information regarding the corresponding postures of the at least one subject; Being operable to cause the one or more processors to perform operations including the above; A system characterized by the above. Claim 2. The plurality of videos to be obtained are two-dimensional (2D) videos, The system according to claim 1. Claim 3. The environment is an operating room, The system according to claim 1. Claim 4. The program instructions are further operable to cause the one or more processors to perform operations including determining one or more keypoints, which are feature points to be tracked when determining the posture of the at least one subject, using a keypoint estimator when executed; The system according to claim 1. Claim 5. The program instructions are operable to cause the one or more processors to perform operations including, at runtime, estimating information regarding a two-dimensional pose of the at least one subject based on images acquired from a plurality of viewpoints by the at least two cameras, and determining information regarding a three-dimensional pose of the at least one subject by performing triangulation on the estimated information regarding the two-dimensional pose. The system according to claim 1.
6. The program instructions are further operable to cause the one or more processors to perform operations including, at runtime, generating a 3D model processed by the one or more processors when recognizing an action of the at least one subject, based on information integrating information regarding a pose of the at least one subject included in the plurality of videos, from a two-dimensional (2D) image, and the plurality of videos are two-dimensional (2D) videos. The system according to claim 1.
7. A computer-readable storage medium storing program instructions, the program instructions, when executed by one or more processors, acquiring a plurality of videos of at least one subject performing at least one action in an environment from at least two cameras each having a viewpoint selected based on different conditions; tracking the at least one subject based on the plurality of videos acquired by the at least two cameras; determining information regarding a pose of the at least one subject detected by the at least two cameras; determining information regarding corresponding poses from information regarding different poses of the at least one subject that are the same object detected by the at least two cameras; matching information regarding corresponding poses of the at least one subject that are the same object; integrating information regarding the poses of the at least one subject that are matched from different viewpoints of the at least two cameras in a plurality of adaptively selected different combinations; Generating a three-dimensional (3D) model processed by the one or more processors from a two-dimensional (2D) image when recognizing an action of the at least one subject based on information regarding corresponding postures of the at least one integrated subject; A computer-readable storage medium operable to cause the one or more processors to execute operations including the above. **Claim 8** The plurality of videos obtained are two-dimensional (2D) videos. The computer-readable storage medium according to claim 7. **Claim 9** The environment is an operating room. The computer-readable storage medium according to claim 7. **Claim 10** The program instructions are further operable to cause the one or more processors to execute operations including determining, when executed, one or more keypoints that are feature points to be tracked when determining the posture of the at least one subject using a keypoint estimator. The computer-readable storage medium according to claim 7. **Claim 11** The program instructions are operable to cause the one or more processors to execute operations including estimating information regarding a two-dimensional posture of the at least one subject based on images obtained from a plurality of viewpoints by the at least two cameras when executed, and performing triangulation on the estimated information regarding the two-dimensional posture to determine information regarding a three-dimensional posture of the at least one subject. The computer-readable storage medium according to claim 7. **Claim 12** The program instructions are further operable to cause the one or more processors to execute operations including generating a 3D model processed by the one or more processors when recognizing an action of the at least one subject from a two-dimensional (2D) image based on information integrating information regarding the posture of the at least one subject included in the plurality of videos, and the plurality of videos are two-dimensional (2D) videos. The computer-readable storage medium according to claim 7. **Claim 13** A method of implementing program instructions on a computer, wherein the program instructions, when executed by the computer, Obtaining a plurality of videos of at least one subject performing at least one action in an environment from at least two cameras, each having a perspective selected based on different conditions; Tracking the at least one subject based on the plurality of videos acquired by the at least two cameras; Determining information regarding the posture of the at least one subject detected by the at least two cameras; Determining information regarding corresponding postures from information regarding different postures of the at least one subject who is the same target detected by the at least two cameras; Matching the information regarding the corresponding postures of the at least one subject who is the same target; Integrating the information regarding the postures of the at least one subject that has been matched from different perspectives of the at least two cameras in a plurality of adaptively selected different combinations; Generating a three-dimensional (3D) model processed by the one or more processors when recognizing the action of the at least one subject based on the integrated information regarding the corresponding postures of the at least one subject from a two-dimensional (2D) image; A method characterized by including the above.
14. The plurality of videos to be acquired are two-dimensional (2D) videos. The method according to claim 13.
15. The environment is an operating room. The method according to claim 13.
16. Further including determining one or more keypoints, which are feature points to be tracked when determining the posture of the at least one subject using a keypoint estimator. The method according to claim 13.
17. Further including estimating information regarding the two-dimensional posture of the at least one subject based on images acquired from a plurality of perspectives by the at least two cameras, and performing triangulation on the estimated information regarding the two-dimensional posture to determine information regarding the three-dimensional posture of the at least one subject. The method according to claim 13.
Citation Information
Patent Citations
Motion-controlled body capture and reconstruction
US20150294492A1
Workflow assistant for image guided procedures
US20190090954A1