Vision-based monitoring system for validating activity sequences
The vision monitoring system with multiple cameras and a visual processor addresses human-robot interaction inefficiencies by accurately tracking human actions and sequences, enhancing safety and efficiency in manufacturing processes.
Patent Information
- Application Number
- DE102015104954
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2014-04-10
- Filing Date
- 2015-03-31
- Publication Date
- 2025-08-14
- Estimated Expiration
- 2035-03-31
AI Technical Summary
Existing human-robot interaction systems lack the ability to efficiently track human actions and behaviors in dynamic workspaces, leading to potential collisions and inefficiencies in manufacturing processes.
A vision monitoring system using multiple cameras and a visual processor to capture and analyze human actions, integrating 3D tracking and activity logging, with pattern matching and actuation signals to validate action sequences and provide alarms for deviations.
Enhances safety and efficiency by accurately tracking human movements and actions, preventing collisions, and ensuring adherence to predefined task sequences in manufacturing environments.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present invention relates generally to vision monitoring systems for tracking people and validating a sequence of actions. BACKGROUND
[0002] Factory automation is used in many assembly contexts. To enable more flexible manufacturing processes, systems are needed that allow robots and humans to interact naturally and efficiently to perform tasks that are not necessarily repetitive. Human-robot interaction requires a new level of machine awareness that goes beyond the typical record / playback control style, where all parts start at a known location. Thus, the robot control system must understand the human's position and behavior and subsequently adapt the robot's behavior based on the human's actions.
[0003] The document US 2011 / 0 050 878 A1 discloses a human monitoring system according to the preamble of claim 1.
[0004] The document US 2012 / 0 214 594 A1 discloses methods and systems for recognizing human movements using a skeleton model, in which a database of movement prototype features is correlated with received skeleton movement data in order to determine match probabilities which are used to determine the most likely movement prototype.
[0005] The publication Weilun, Lao ea; "Automatic surveillance analyzer using trajectory and body-based modeling", In: International Conference on Consumer Electronics, 10.01.2009, IEEE Conference, ICCE '09, ISBN 978-1-4244-4701-5, pp. 1-2 discloses a device for analyzing video images that can detect multiple people in the video images, estimate their trajectories, and classify their poses.
[0006] In the paper "Anomaly Detection Using Motion Patterns Computed from Optical Flow" by Parvathy, R. ea., in: 2013 Third International Conference on Advances in Computing and Communications, IEEE Conference, August 29, 2013, pp. 58-61, a method for monitoring crowds of people to detect anomalies is disclosed. For this purpose, an optical flow is calculated as a vector field, where each vector represents a direction and amount of movement. This defines trajectories that are summarized and evaluated using statistical methods to detect anomalies.
[0007] The publication Krüger, Jörg ea; “Image-Based 3D Surveillance in Man-Robot Cooperation”, In: 2nd IEEE International Conference on Industrial Informatics, 2004. INDIN '04.2004, pp. 411-420 discloses a multi-camera image processing system that monitors a collaborative workspace in which robots and humans work together and avoids potential collisions.
[0008] In the paper Ding, H. ea; “Structured Collaborative Behaviour of Industrial Robots in Mixed Human-Robot Environments”, 2013 IEEE International Conference on Automation Science and Engineering (CASE), pp. 1101-1106, a finite state machine for automatic response to exceptional conditions and for resuming operation for industrial robots that collaborate with humans is disclosed. SUMMARY
[0009] A human monitoring system includes multiple cameras and a visual processor. The multiple cameras are arranged around a workspace, with each camera configured to capture a video stream containing multiple individual images, with the multiple individual images being time-synchronized between the respective cameras.
[0010] The visual processor is configured to identify the presence of a human in the workspace from the plurality of individual images, generate a motion track of the human in the workspace, generate an activity log of one or more activities performed by the human across the motion track, and compare the motion track and the activity log to an activity template defining a plurality of required actions. If one or more actions in the activity template are not performed in the workspace, the processor provides an alarm. In one embodiment, the alarm may include a video synopsis of the generated motion track.
[0011] The visual processor may generate the activity log of the one or more activities performed by the human throughout the motion tracking by pattern matching a detected representation of the human against a trained pose database. This pattern matching may include the use of a support vector machine and / or a neural network. The pattern matching may further detect the pose of the human, where the pose includes standing, walking, reaching, and / or crouching.
[0012] The visual processor is configured to identify the presence of a human in the workspace by detecting the human in each of the multiple individual images, mapping a representation of the detected human from the multiple views into a common coordinate system, and determining an intersection point of the mapped representations. Each of the multiple individual images is obtained from a different camera selected from the multiple cameras and represents a different view of the workspace.
[0013] The visual processor may further receive an actuation signal from a tool, the actuation signal indicating the tool currently being used. The visual processor is configured to use the actuation signal to confirm the performance of one of the one or more activities performed by the human.
[0014] According to the invention, each of the multiple required actions is specified at a respective location in the workspace. The visual processor can provide an alarm if the one or more actions in the activity template are not performed in the workspace at their respective specified location. The system is adapted to several different ways of executing a sequence of tasks, and it validates the motion tracking and activity log as long as a final human path and the performed activities conform to pre-specified targets at pre-specified locations in the workspace.
[0015] The above features and advantages and other features and advantages of the present invention will be readily apparent from the following detailed description of the best modes for carrying out the invention when taken in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS Fig. Figure 1 is a schematic block diagram of a human monitoring system. Fig. Figure 2 is a schematic representation of multiple imaging devices positioned around a work area. Fig. Figure 3 is a schematic block diagram of an activity monitoring process. Fig. Figure 4 is a schematic process flow diagram for detecting human movement using multiple imaging devices positioned around a work area. Fig. Figure 5A is a schematic representation of an image frame containing a sliding window input to a pattern matching algorithm that traverses the image frame in an image coordinate space. Fig. Figure 5B is a schematic representation of an image frame containing a sliding window input to a pattern matching algorithm that traverses the image frame in a corrected coordinate space. Fig. Figure 5C is a schematic representation of the single image from Fig. 5B, where the sliding window input is selected from a specific region of interest. Fig. Figure 6 is a schematic diagram illustrating one way of combining multiple representations of a detected human, each from a different camera, into a common coordinate system. Fig. 7 is a high-level schematic flow diagram of a method for performing activity sequence monitoring using the human monitoring system. Fig. 8 is a schematic detailed flowchart of a method for performing activity sequence monitoring using the human monitoring system. Fig. Figure 9 is a schematic diagram of the human monitoring system used across multiple work areas. Fig. Figure 10 is a schematic representation of three-dimensional localization using multiple sensor views. DETAILED DESCRIPTION
[0016] In the drawings, in which like reference numerals are used to identify similar or identical components throughout the different views, Fig. 1 schematically shows a block diagram of a human monitoring system 10 for monitoring a work area of an assembly process, a manufacturing process, or a similar process. The human monitoring system 10 includes a plurality of vision-based imaging devices 12 for capturing visual images of a particular work area. The plurality of vision-based imaging devices 12, as shown in Fig. 2 are positioned at different locations and at different heights so as to surround the automated moving equipment. Preferably, wide-angle lenses or similar devices with a wide field of view are used to visually cover more of the workspace area. All vision-based imaging devices are substantially offset from one another to capture an image of the workspace from a respective viewpoint that is substantially different from that of the other respective imaging devices. This allows different streaming video images to be captured from different viewpoints throughout the workspace to distinguish a person from the surrounding equipment. Due to visual obstructions (ieOcclusions) with objects and equipment in the workspace, the multiple viewpoints increase the probability of detecting the person in one or more images when occlusions are present within the workspace.
[0017] As in Fig. 2, a first vision-based imaging device 14 and a second vision-based imaging device 16 are positioned at overhead positions substantially spaced apart from each other so that each captures a high-angle view. The imaging devices 14 and 16 provide canonical high-angle or reference views. Preferably, the imaging devices 14 and 16 provide stereo-based three-dimensional scene analysis and scene tracking. The imaging devices 14 and 16 may include visual imaging, LIDAR detection, infrared detection, and / or any other type of imaging that can be used to detect physical objects within a region. Additional imaging devices may be positioned overhead and spaced apart from the first and second vision-based imaging devices 14 and 16 to obtain additional overhead views.For ease of description, the imaging devices 14 and 16 may be referred to generically as "cameras," although it should be recognized that such cameras need not be visible spectrum cameras unless otherwise indicated.
[0018] Various other vision-based imaging devices 17 ("cameras") are positioned on the sides or in the virtual corners of the monitored workspaces to capture medium-angle and / or low-angle views. Since the number of vision-based imaging devices is reconfigurable, as the system can operate with any number of imaging devices, it should be noted that more or fewer imaging devices than in Fig. 2; however, it is emphasized that the degree of integrity and redundant reliability increases as the number of redundant imaging devices increases. Each of the vision-based imaging devices 12 is spaced from another to capture an image from a viewpoint substantially different from another to produce a three-dimensional tracking of one or more people in the work area. The different views captured by the multiple vision-based imaging devices 12 together provide alternative views of the work area that enable the human monitoring system 10 to identify each person in the work area.These different viewpoints provide the opportunity to track each person across the entire workspace in three-dimensional space and improve the localization and tracking of each person as they move through the workspace to detect unwanted interactions between each person and the moving automated equipment in the workspace.
[0019] Again based on Fig. 1, the images captured by the plurality of vision-based imaging devices 12 are transmitted to a processing unit 18 via a communication medium 20. The communication medium 20 may be a communication bus, an Ethernet, or another communication connection (including a wireless one).
[0020] The processing unit 18 is preferably a host computer implemented with off-the-shelf components (not unlike a personal computer) or a similar device mounted in a housing suitable for its operating environment. The processing unit 18 may further include an image acquisition system (possibly including a frame grabber and / or network image acquisition software) used to capture image streams, process, and record the image streams as time-synchronized data. Multiple processing units may be interconnected in a data network using a protocol that ensures message integrity, such as Ethernet-Safe.Data indicating the status of an adjacent space monitored by other processing units, including alerts, signals, and tracking status data transmissions for people, objects moving from area to area, or zones spanning multiple systems, can be exchanged in a reliable manner. The processing unit 18 utilizes a primary processing routine and multiple sub-processing routines (i.e., one sub-processing routine for each vision-based imaging device). Each respective sub-processing routine is dedicated to a respective imaging device to process the images acquired by the respective imaging device. The primary processing routine performs multi-view integration to perform real-time monitoring of the workspace based on the cumulative acquired images as processed by each sub-processing routine.
[0021] In Fig. 1, detection of a worker in the work area is enabled by the subprocessing routines utilizing multiple databases 22 that collectively detect and identify people in the presence of other mobile equipment in the work area. The multiple databases store data used to detect objects, identify a person from the detected objects, and track an identified person in the work area. The various databases include, but are not limited to, a calibration database 24, a background database 25, a classification database 26, a vanishing point database 27, a tracking database 28, and a homography database 30. Data present in the databases is used by the subprocessing routines to detect, identify, and track people in the work area.
[0022] The calibration database 24 provides pattern-based camera calibration parameters (intrinsic and extrinsic) to correct distorted objects. In one configuration, the calibration parameters can be determined using a regular pattern, such as a checkerboard, displayed orthogonally to the camera's field of view. A calibration routine then uses the checkerboard to estimate the intrinsic and correction parameters that can be used to correct barrel distortion caused by the wide-angle lenses.
[0023] The background database 25 stores the background models for various views, which are used to decompose an image into its background and foreground component regions. The background models can be obtained by capturing images / video prior to installing any automated machinery or prior to arranging any dynamic objects in the workspace.
[0024] The classification database 26 contains a cascade of classifiers and related parameters to automatically classify humans and non-humans.
[0025] The vanishing point database 27 contains the vanishing point information for each of the camera views and is used to perform the vanishing point correction so that people appear upright in the corrected images.
[0026] The tracking database 28 maintains trajectories for each of the monitored humans, with new trajectories being added to the database as new humans enter the scene and deleted as they leave. Furthermore, the tracking database contains information about the appearance model for each human, so that existing trajectories can be easily mapped to trajectories in a different time step.
[0027] The homography database 30 contains the homography transformation parameters across the various views and across the canonical view. Appropriate data from the database(s) can be transferred to a system monitoring a neighboring area while a person walks into the area, enabling seamless transition of the person's tracking from area to area across multiple systems.
[0028] Each of the databases described above may contain parameters that are the result of various initialization routines executed during installation and / or maintenance of the system. The parameters may be stored, for example, in a format that is easily accessible to the processor during operation, such as an XML file format. In one configuration, the system may execute a lens calibration routine during the initial setup / initialization routine, such as by placing a checkerboard image within the field of view of each camera. Using the checkerboard image, the lens calibration routine may determine the required amount of correction necessary to remove any fisheye distortion. These correction parameters may be stored in the calibration database 24.
[0029] Following the lens calibration routine, the system can then determine the homography transformation parameters, which can be recorded in the homography database 30. This routine can include arranging reference objects within the workspace so that they can be viewed by multiple cameras. By correlating the location of the objects between the different views (and while knowing the fixed position of either the cameras or the objects), the various two-dimensional images can be mapped to 3D space.
[0030] Furthermore, by arranging multiple vertical reference markers at different locations within the workspace and analyzing how these markers are represented within each camera view, the vanishing point of each camera can be determined. The perspective nature of the camera can cause the representations of the respective vertical markers to converge on a common vanishing point, which can be recorded in the vanishing point database 27.
[0031] Fig. Figure 3 illustrates a block diagram of a high-level overview of the factory monitoring process flow that includes dynamic system integrity monitoring.
[0032] In block 32, data streams from the vision-based imaging devices 12, which capture the time-synchronized image data, are collected. In block 33, system integrity monitoring is performed. The vision processing unit checks the system's integrity for component failures and conditions that would prevent the monitoring system from operating properly and fulfilling its intended purpose. This "dynamic integrity monitoring" would detect these degraded or faulty conditions and trigger a mode of operation in which the system fails into a self-protective mode of operation where system integrity can be restored and process interaction can return to normal without any unintended consequences, except for the downtime necessary to perform repairs.
[0033] In one configuration, reference targets can be used for geometric calibration and integrity. Some of these reference targets, such as a blinking IR beacon in the field of view of one or more sensors, could be active. In one configuration, for example, the IR beacon can be made to blink at a particular rate. The monitoring system can then determine whether the beacon detection in the image actually coincides with the expected rate at which the IR beacon is actually blinking. If this is not the case, the automated equipment can fail into a safe mode, a faulty view can be discarded or disabled, or the equipment can be modified to operate in a safe mode.
[0034] Unexpected changes in the behavior of a reference target may also require the equipment to be modified to operate in the safe mode. For example, if a reference target is a moving object being tracked and disappears before the system detects it leaving the work area from an expected exit location, similar precautions may be taken. Another example of unexpected changes to a moving reference target is if the reference target appears at a first location and subsequently reappears at a second location at an inexplicably rapid rate (i.e., with a distance-to-time ratio that exceeds a predetermined threshold). If the visual processing unit in block 34 Fig. If block 3 determines that there are integrity problems, the system then enters the self-protective mode, in which warnings are activated and the system is shut down. If the visual processing unit determines that no integrity problems exist, blocks 35-39 are initiated sequentially.
[0035] In one configuration, the system integrity monitor 33 may include quantitatively assessing the integrity of each vision-based imaging device in a dynamic manner. For example, the integrity monitor may continuously analyze each video stream to measure the amount of noise within a stream or to identify discontinuities in the image over time. In one configuration, the system may use an absolute pixel difference and / or a global and / or local histogram difference and / or absolute edge differences to quantify the integrity of the image (i.e., to determine a relative "integrity score" ranging from 0.0 (no reliability) to 1.0 (ideally reliable)). The aforementioned differences may then be calculated either with respect to a predetermined reference frame / image (e.g.,one acquired during initialization routing) or relative to a frame acquired immediately before the frame being measured. When comparing to a pre-established reference frame / image, the algorithm can focus specifically on one or more background portions of the image (rather than dynamically changing foreground portions).
[0036] In block 35, background subtraction is performed, and the resulting images become the foreground regions. Background subtraction allows the system to identify those aspects of the image that are capable of motion. These portions of the individual images are then passed to subsequent modules for further analysis.
[0037] In block 36, the human check is performed to detect humans from the captured images. In this step, the identified foreground images are processed to detect / identify portions of the foreground that are most likely to be human.
[0038] In block 37, as previously described, an appearance comparison and tracking are performed, which identifies a person from the detected objects using various databases and tracks an identified person in the work area.
[0039] In block 38, three-dimensional processing is applied to the acquired data to obtain 3D range information for the objects in the workspace. The 3D range information enables the generation of 3D occupancy grids and voxels, which reduce false alarms and enable 3D object tracking. The 3D metrology processing can be performed, for example, using the overhead stereoscopic cameras (e.g., cameras 14, 16) or can be performed using voxel construction techniques from the projection of each angled camera 17.
[0040] In block 39, the compared trajectories are provided to a multi-view merging and object localization module. The multi-view merging module 39 can merge the different views to generate a probabilistic map of the location of each human within the workspace. Furthermore, the multi-view merging and object localization module receives information from the vision-based imaging devices as in Fig. 10 provides three-dimensional processing to determine the location, direction, speed, occupancy, and density of each human within the workspace. The identified humans are tracked for potential interaction with moving equipment within the workspace.
[0041] Fig. Figure 4 illustrates a process flowchart for detecting, identifying, and tracking people using the human monitoring system. In block 40, the system is initialized by the primary processing routine to perform multi-view integration in the monitored workspace. The primary processing routine initializes and starts the subprocessing routines. A respective subprocessing routine is provided for processing the data acquired by a respective imaging device. All subprocessing routines operate in parallel. To ensure that the acquired images are time-synchronized with each other, the following processing blocks, as described herein, are synchronized by the primary processing routine.The primary processing routine waits for each of the subprocessors to complete processing its respective acquired data before executing a multi-view integration. Preferably, the processing time for each subprocessor is no more than 100-200 ms. A system integrity check is also performed during system initialization (see also ). Fig. 3, Block 33). If it is determined that the system integration check has failed, the system immediately enables a warning and enters a self-protective mode in which the system is shut down until corrective action is taken.
[0042] Back in Fig. 4, in block 41, streaming image data is acquired by each vision-based imaging device. The data acquired by each imaging device is in pixel form (or is converted to pixel form). In block 42, the acquired image data is provided to an image buffer, where the images await processing to detect objects, and in particular, people, in the workspace beneath the moving automated equipment. Each acquired image is time-stamped so that each acquired image is synchronized for concurrent processing.
[0043] In block 43, autocalibration is applied to the acquired images to dewarp objects within the acquired image. The calibration database provides calibration parameters based on patterns for dewarping distorted objects. The image distortion caused by wide-angle lenses requires that the image be dewarped by applying camera calibration. This is necessary because any severe distortion of the image will render the homography mapping function between the imaging device views and the appearance models inaccurate. Imaging calibration is a one-time process; however, recalibration is required if the imaging device setting is changed. In addition, the image calibration is periodically checked by the dynamic integrity monitoring subsystem to detect conditions where the imaging device has been moved slightly outside its calibrated field of view.
[0044] In blocks 44 and 45, background modeling and foreground modeling, respectively, are initiated. Background training is used to distinguish background images from foreground images. The results are stored in a background database for use by each of the sub-processing routines to distinguish the background and foreground. All rectified images are background filtered to preserve foreground pixels within a digitized image. To distinguish the background in an acquired image, background parameters should be trained using images of an empty workspace viewing area so that the background pixels can be easily distinguished when moving objects are present. The background data should be updated over time.When a person is detected and tracked in the acquired image, the background pixels are filtered from the image data to detect foreground pixels. The detected foreground pixels are converted into blobs using connected component analysis with noise filtering and blob size filtering.
[0045] In block 46, a blob analysis is initiated. In a given workspace, not only a moving person can be detected, but also other moving objects such as robot arms, carts, or boxes. Thus, the blob analysis involves detecting all foreground pixels and determining which foreground images (e.g., blobs) are people and which are moving non-human objects.
[0046] A blob can be defined as an area of connected pixels (e.g., touching pixels). Blob analysis involves identifying and analyzing the specific area of pixels within the acquired image. The image distinguishes pixels based on a value. The pixels are then identified as either a foreground or a background. Pixels with a non-zero value are considered foreground, and pixels with a value of zero are considered background. Typically, blob analysis considers several factors, which may include, but are not limited to, the location of the blob, the area of the blob, the perimeter (e.g., edges) of the blob, the shape of the blob, the diameter, length, or width of the blob, and the orientation.Image or data segmentation techniques are not limited to 2D images, but can also leverage the output data from other sensor types that provide IR images and / or volumetric 3D data.
[0047] In block 47, as part of the blob analysis, human detection / human verification is performed to filter out non-human blobs from the human blobs. In one configuration, this verification can be performed using a swarming domain classifier technique.
[0048] In another configuration, the system may use pattern matching algorithms such as support vector machines (SVMs) or neural networks to perform pattern matching of foreground blobs with trained human pose models. Instead of attempting to process the entire image as a single entity, the system may instead process the individual image 60 using a localized sliding window 62 as generally described in Fig. 5A. This can reduce processing complexity and improve the robustness and specificity of detection. The sliding window 62 can then serve as the input to the SVM for identification.
[0049] The models that perform human detection can be trained using images of different people positioned in different postures (i.e., standing, crouching, kneeling, etc.) and facing in different directions. When training the model, the representative images can be provided in such a way that the person is generally aligned with the vertical axis of the image. As in Fig. However, as shown in Figure 5A, the body axis of an imaged person 64 may be angled in accordance with the perspective and the vanishing point of the point, which is not necessarily vertical. If the input to the detection model was a window aligned with the image coordinate system, the person's angled representation may negatively impact detection accuracy.
[0050] To account for the slanted nature of people in the image, the sliding window 62 can be used from a corrected space instead of the image coordinate space. The corrected space can map the perspective view to a rectangular view aligned with the ground plane. In other words, the corrected space can map a vertical in the workspace such that it is vertically aligned within an adjusted image. This is schematically shown in Fig. 5B, where a corrected sliding window 66 scans the individual image 60 and can map an angled person 64 onto a vertically oriented representation 68 provided in a rectangular space 70. This vertically oriented representation 68 can then ensure higher confidence detection when analyzed using SVM. In one configuration, the corrected sliding window 66 can be enabled by a correlation matrix that can map, for example, between a polar coordinate system and a rectangular coordinate system.
[0051] Although in one configuration the system can perform an exhaustive search across the entire image using the sliding window search strategy described above, the strategy can include searching for areas of the image where people are not physically located. Thus, in another configuration the system can restrict the search space to only a specific region of interest 72 (ROI), such as in Fig. 5C. In one configuration, the ROI 72 may represent the visible floor area within the frame 60 plus an edge tolerance to account for a person standing at the outermost edge of the floor area.
[0052] In yet another configuration, the computational requirements can be further reduced by prioritizing the search around portions of the ROI 72 where human blobs are expected to be found. In this configuration, the system can use engagement points to limit or prioritize the search based on supplementary information available to the image processor. This supplementary information can include motion detection within the frame, trajectory information from a previously identified human blob, and data fusion from other cameras in the multi-camera array. For example, after verifying a human location on the fused ground frame, the tracking algorithm generates a human trajectory and maintains the trajectory history across subsequent frames.If an environmental obstacle causes human localization to fail in one case, the system can quickly recover the human location by extrapolating the trajectory of the previously tracked human location to focus the corrected search within ROI 72. If the blob is not re-identified in multiple frames, the system can report that the target human has disappeared.
[0053] When the human blobs are again identified by Fig. 4 have been detected in the various views, a body axis estimation is performed in block 48 for each detected human blob. Using vanishing points (obtained from the vanishing point database) in the image, a main body axis line is determined for each human blob. In one configuration, the body axis line can be defined by two points of interest. The first point is a centroid of the identified human blob, and the second point (i.e., the vanishing point) is a respective point near a body bottom (i.e., not necessarily the blob bottom and possibly outside the blob). More precisely, the body axis line is a virtual line connecting the centroid to the vanishing point. As generally seen at 80, 82, and 84 from Fig. As shown in Figure 6, a respective vertical body axis line is determined for each human blob in each respective camera view. Generally, this line samples the human image on a line from the head to the shoe. A human detection score can be used to assist in determining a corresponding body axis. The score provides a level of confidence that a fit to the human has been made and that the corresponding body axis should be used. Each vertical body axis line is used via homography mapping to determine the human's location and is discussed in detail later.
[0054] Again based on Fig. 4, color profiling is performed in block 49. A driver appearance model is provided to compare the same person in each view. A color profile provides fingerprints and maintains the identity of the respective person across each captured image. In one configuration, the color profile is a vector of averaged color values of the body axis line with the blob's bounding box.
[0055] In blocks 50 and 51, a homography mapping and a multi-view integration routine are executed to coordinate the respective views and map the human location to a common plane. Homography (as used here) is a mathematical concept in which an invertible transformation maps objects from a coordinate system to a line or plane.
[0056] The homography mapping module 50 may include a body axis submodule and / or a synergy submodule. In general, the body axis submodule may use homography to map the detected / calculated body axis lines into a common plane viewed from an overhead perspective. In one configuration, this plane is a ground plane that coincides with the floor of the workspace. This illustration is schematically shown above the ground plane map at 86 in Fig. 6. After mapping into the common ground plane, the various body axis lines may intersect at or near a single location point 87 in the ground plane. In a case where the body axis lines do not intersect ideally, the system may use a least squares or least median squares approach to identify a best fit approximation of location point 87. This location point may represent an estimate of the location of the human's ground plane within the workspace. In another embodiment, location point 87 may be determined using a weighted least squares approach, where each line may be individually weighted using the integrity score determined for the frame / view from which the line was determined.
[0057] The synergy submodule may operate similarly to the body axis submodule in that it uses homography to map content from different image views into planes, each perceived from an overhead perspective. However, instead of mapping a single line (i.e., the body axis line), the synergy submodule maps the entire detected foreground blob onto the plane. More specifically, the synergy submodule uses homography to map the foreground blob onto a synergy map 88. This synergy map 88 is multiple planes, all parallel and each at a different elevation relative to the floor of the workspace. The detected blobs from each view may be mapped into each respective plane using homography. For example, in one configuration, the synergy map 88 may include a ground plane, a midplane, and a head plane.In other configurations, more or fewer levels can be used.
[0058] During the mapping of a foreground blob from each respective view to a common plane, there may be an area where multiple blob mappings overlap. In other words, when the pixels of a perceived blob in a view are mapped to a plane, each pixel of the original view has a corresponding pixel in the plane. When multiple views are all projected onto the plane, they are likely to intersect in an area, so that a pixel in the plane can be mapped to multiple original views from the intersection area. This area of within-plane coincidence reflects a high probability of the presence of a human at that location and height. In a similar manner to the body axis submodule, the integrity score can be used to weight the projections of the blobs from each view onto the synergy map 88.Thus, the clarity of the original image can influence the specific boundaries of the high probability region.
[0059] Once the blobs from each view have been mapped to the respective planes, the high probability regions can be isolated and regions can be grouped together along a common vertical axis. By isolating these high probability regions at different heights, the system can construct a bounding envelope that encapsulates the detected human shape. The position, velocity, and / or acceleration of this bounding envelope can then be used to modify the behavior of adjacent automated equipment, such as an assembly robot, or to provide a warning, for example, if a person were to enter or reach a defined protection zone. If a bounding envelope, for example,overlaps or impacts a specific restricted volume, the system can modify the performance of automated devices within the restricted volume (e.g., it can slow down or stop a robot). Furthermore, the system can anticipate the object's movement by monitoring its speed and / or acceleration, and it can modify the behavior of the automated device if a collision or interaction is anticipated.
[0060] In addition to simply identifying the bounding envelope, the entirety of the envelope (and / or the entirety of each plane) can be mapped down to the ground plane to determine a likely occupied ground area. In one configuration, this occupied ground area can be used to validate the location point 87 determined by the body axis sub-module. For example, location point 87 can be validated if it lies within a highly likely occupied ground area as determined by the synergy sub-module. Conversely, the system can identify an error or reject location point 87 if point 87 lies outside the area.
[0061] In another configuration, a principal axis may be drawn through the bounding envelope such that the axis is substantially vertical within the workspace (i.e., substantially perpendicular to the ground plane). The principal axis may be drawn at a central location within the bounding envelope and may intersect the ground plane at a second location. This second location may be merged with the location 87 determined via the body axis submodule.
[0062] In one configuration, the multi-view integration 51 may combine several different types of information to increase the probability of accurate detection. As shown in Fig. 6, for example, the information within the ground plane map 86 and the information within the synergy map 88 can be merged to form a consolidated probability map 92. To further refine the probability map 92, the system 10 can additionally merge 3D stereo or constructed voxel representations 94 of the workspace with the probability estimates. In this configuration, the 3D stereo can use scale-invariant feature transforms (SIFTs) to first obtain features and their counterparts. The system can then perform epipolar correction on both stereo pairs based on the known intrinsic camera parameters and the feature counterparts. A disparity map (depth map) can then be obtained in real time using a block matching technique, provided, for example, in OpenCV.
[0063] Similarly, voxel representation uses the image outlines obtained from background subtraction to generate a depth representation. The system projects 3D voxels onto all image planes (of the multiple cameras used) and determines whether the projection overlaps with the outlines (foreground pixels) in most images. Since certain images may be obscured due to robots or factory equipment, the system can use a matching scheme that does not directly require overlap of all images. The 3D stereo and 3D voxel results provide information about how the objects occupy 3D space, which can be used to improve the probability map 92.
[0064] The development of the probability map 92 by merging different data types can be achieved in several different ways. The simplest is a 'simple weighted mean integration' approach, which applies a weighting coefficient to each data type (i.e., to the body axis projection, to the synergy map 88, to the 3D stereo depth projection, and / or to the voxel representation). Furthermore, the body axis projection may further include Gaussian distributions across each body axis line, with each Gaussian distribution representing the distribution of blob pixels around the respective body axis line. When projected onto the ground plane, these distributions may overlap, in which case they may assist in determining the location point 87 or may be merged with the synergy map.
[0065] A second approach to fusion can use a 3D stereo and / or voxel depth map along with a foreground blob projection to pre-filter the image. Once pre-filtered, the system can perform multi-level body axis analysis within these filtered regions to provide higher-confidence body axis extraction in each view.
[0066] Again based on Fig. 4, in block 52, one or more trajectories can be assembled based on the determined multi-view homography information and the determined color profile. These trajectories can represent the orderly movement of a detected human through the entire workspace. In one configuration, the trajectories are filtered using Kalman filtering. In Kalman filtering, the state variables are the person's ground location and ground velocity.
[0067] In block 53, the system can determine whether a user's trajectory matches an expected or acceptable trajectory for a particular procedure. In addition, the system can also attempt to "predict" a person's intention to continue walking in a particular direction. This intent information can be used in other modules to calculate the final time and distance rate between the person and the detection zone (this is particularly important in improving zone detection latency with dynamic detection zones that follow the movement of equipment such as robots, conveyors, forklifts, and other mobile equipment).In addition, this is important information that can predict the person's movement into a neighboring monitored area, to which the person's data can be transmitted and in which the receiving system can prepare attention mechanisms to quickly detect the person's tracking in the entered monitored area.
[0068] If a person's particular activity is not validated or falls outside acceptable procedures, or if a person is anticipated to leave a predefined "safe zone," the system may issue an alarm in block 54, conveying the warning to the user. The alarm may, for example, be displayed on a display device while people are moving through the predefined safe zones, warning zones, and critical zones of the work area. The warning zone and critical zones (as well as any other zones desired to be configured in the system, including dynamic zones) are operating areas where alarms are provided, as initiated in block 54, when the person has entered the respective zone, and cause the equipment to slow down, stop, or otherwise avoid the person.The warning zone is an area where the person is initially warned that a person has entered an area and is sufficiently close to the moving equipment to cause the equipment to stop. The critical zone is a location (e.g., an envelope) determined within the warning zone. When the person is within the critical zone, a more critical warning can be issued so that the person is aware of their location in the critical zone or is requested to leave the critical zone. These warnings are provided to improve the productivity of the process system by preventing inconvenient equipment shutdowns caused by the unexpected entry of persons unaware of their proximity into the warning zones.These warnings are also silenced by the system during intervals of expected interaction, such as routine loading or unloading of the process. Furthermore, it is possible that a currently stationary person would be detected in the path of a dynamic zone moving in their direction.
[0069] In addition to warnings given to the person when they are in the respective zones, the warning can change or modify the movement of nearby automated equipment (e.g., stopping, accelerating, or decelerating the equipment) depending on the predicted path of the person (or possibly the dynamic zone) within the workspace. That is, the movement of the automated equipment operates according to a target routine that has predefined movements at a predetermined speed. By tracking and predicting the person's movements within the workspace, the movement of the automated equipment can be modified (i.e., slowing or accelerating) to avoid any potential contact with the person within the workspace zone.This allows equipment operation to be maintained without shutting down the assembly / manufacturing process. Current self-protecting operations are determined by the results of a task based on a risk assessment and typically require automated factory equipment to be completely stopped if a person is detected in a critical area. Startup procedures require an equipment operator to reset the controls to restart the assembly / manufacturing process. Such an unexpected halt in the process typically results in downtime and lost productivity. Activity sequence monitoring
[0070] In one configuration, the system described above can be used to monitor a series of operations performed by a user and verify whether the monitored process is being executed correctly. In addition to simply analyzing video streams, the system can also monitor the timing and use of auxiliary equipment such as torque guns, nut-tightening machines, or screwdrivers.
[0071] Fig. Figure 7 generally illustrates a method 100 for performing activity sequence monitoring using the above system. As shown, the input video is processed at 102 to generate an internal representation 104 that captures various types of information such as scene motion, activities, etc. The representations are used at 106 to learn classifiers, generating action labels and action similarity scores. This information is compiled at 108 and converted into a semantic description, which is then compared at 110 to a known activity template to generate an error-proofing score. A semantic synopsis and video synopsis are archived for future reference.If the comparison with the template produces a low score, indicating that the executed sequence is not similar to the expected work item progress, a warning is issued at 112.
[0072] This process can be used to validate an operator's activity by determining when and where certain actions are performed, along with their sequence. For example, if the system identifies that the operator reaches into a specific arranged box, walks toward a corner of a vehicle on the assembly line, kneels, and operates a nut-tightening machine, the system can determine that the operator has most likely attached a wheel to the vehicle. Conversely, if the sequence ends with only three wheels being attached, it can indicate / warn that the process was not completed because a fourth wheel is required. Similarly, the system can compare actions to a vehicle manifest to ensure that the required hardware options are installed for a specific vehicle.For example, if the system detects that the operator is picking up a bezel of the wrong color, the system can warn the user to verify the part before proceeding. In this way, the human monitoring system can be used as an error-prevention tool to ensure that required actions are performed during the assembly process.
[0073] The system can have sufficient flexibility to adapt to several different ways of performing a sequence of tasks and can validate the process as long as the final human trajectory and final activity logs achieve the pre-specified objectives at the pre-specified vehicle locations. Although the efficiency of whether a sequence of actions correctly met the objectives for an assembly station cannot be considered, it can be recorded separately. In this way, the actual trajectory and activity log can be compared with an optimized trajectory to quantify an overall deviation, which can be used to suggest process efficiency improvements (e.g., via a display or a printed activity report).
[0074] Fig. Figure 8 provides a more detailed block diagram 120 of the activity monitoring scheme. As shown, video data streams from the cameras are collected in block 32. These data streams are passed through a system integrity monitoring module at 33, which verifies that the footage is a normal operating regime. If the video streams deviate from the normal operating regime, an error is issued, and the system fails into a safe operating mode. The next step after the system integrity monitoring is a human detector tracker module 122, generally described above in Fig. 4. This module 122 takes each of the video streams and detects the moving people in the scene. If moving candidate blobs are available, the system can use classifiers to process and filter out the non-moving instances. The resulting output of this module is 3D people trajectories. The next step involves extracting appropriate representations from the 3D people trajectories at 124. The representation schemes are complementary and include image pixels 126 for appearance modeling of activities, spatiotemporal interest points (STIPs) 128 for representing scene motion, trajectories 130 for isolating actors from the background, and voxels 132 that integrate information across multiple views. Each of these representation schemes is described in more detail below.
[0075] Once the information has been extracted and represented in the above complementary forms at 104, the system extracts specific features and passes them through a corresponding set of pre-trained classifiers. A temporal SVM classifier 134 processes the STIP features 128 and generates action labels 136 such as standing, crouching, running, bending, etc. A spatial SVM classifier 138 processes the raw image pixels 126 and generates action labels 140. The extracted trajectory information 130 is used together with dynamic time warping action labels 142 to compare trajectories to typical expected trajectories and generate an action similarity score 144. A human pose estimation classifier 146 is trained to use a voxel representation 132 as input and generate a pose estimate 148 as output.The resulting combination of temporal pose, spatial pose, trajectory comparison pose, and voxel-based pose is input into a spatiotemporal signature 150, which becomes the building block for the semantic description module 152. This information is then used to decompose any activity sequence into atomic component actions and generate an AND-OR graph 154. The extracted AND-OR graph 154 is then compared to a prescribed activity list at 156, generating a fit score. A low fit score is used to issue a warning, indicating that the observed action is not typical and is instead anomalous. A semantic and visual synopsis is generated and archived at 158. Spatiotemporal interest points (STIPs) for representing actions
[0076] STIPs 128 are detected features that exhibit a significant local change in image properties over space and / or time. Many of these interest points are generated during the execution of an action by a human. Using STIPs 128, the system can attempt to determine which action is occurring within the observed video sequence. At 134, each extracted STIP feature 128 is passed through the set of SVM classifiers, and a voting mechanism determines which action the feature is most likely associated with. A sliding window then determines the detected action based on the classification of the detected STIPs within the time window in each frame. Since there are multiple views, the window considers all detected features from all views.The resulting information, in the form of one action per frame, can be compiled into a graph displaying the sequence of detected actions. Finally, the graph can be compared with the graph generated during the SVM training phase to verify the correctness of the detected action sequence.
[0077] In one example, STIPs 128 may be generated while observing a person moving across a platform to use a torque wrench in specific areas of the vehicle. This action may involve the person transitioning from a walking pose into one of many drilling poses, maintaining that pose for a short time, and transitioning back into a walking pose. Because STIPs are motion-based points of interest, those generated while entering and exiting each pose are what distinguishes one action from another. Dynamic Time Warping
[0078] Dynamic Time Warping (DTW) (performed at 142) is an algorithm for measuring the similarity between two sequences whose time or speed may vary. For example, similarities in walking patterns between two trajectories would be detected via DTW even if the subject walked slowly in one sequence and faster in another, or even if there were accelerations, decelerations, or multiple short stops, or even if two sequences shifted along the timeline during an observation. DTW can reliably determine an optimal fit between two given sequences (e.g., time series). The sequences are nonlinearly "warped" in the time dimension to determine a measure of their similarity independent of any specific nonlinear changes in the time dimension. The DTW algorithm uses a dynamic programming technique to solve this problem.The first step is to compare each point in one signal with each point in the second signal, creating a matrix. The second step is to work through this matrix, starting in the bottom left corner (corresponding to the beginning of both sequences) and ending in the top right corner (at the end of both sequences). For each cell, the cumulative distance is calculated by selecting the neighboring cell in the matrix to the left or below with the lowest cumulative distance, and this value is added to the distance of the focal cell. When this process is complete, the value in the top right cell represents the distance between the two sequence signals according to the most efficient path through the matrix.
[0079] The DTW can measure similarity using only lane labels or lane plus location labels. In a vehicle assembly context, six location labels can be used: FD, MD, RD, RP, FP, and Running, where F, R, and M represent the front, middle, and rear of the vehicle, and D and P represent the driver and passenger sides, respectively. The distance cost of the DTW is calculated as: Cost=αE+(1−α)L, 0≤α≤1, where E is the Euclidean distance between two points on the two trajectories and L is the histogram difference of locations within a given time window; α is a weight and is set to 0.8 if both trajectory and location labels are used for the DTW measurement. Otherwise, α is 1 for the trajectory-only measurement. Action labels using spatial classifiers
[0080] A single-image recognition system can be used to distinguish between a number of possible aggregate actions visible in the data, such as walking, bending, crouching, and reaching. These action labels can be determined using scale-invariant feature transforms (SIFT) and SVM classifiers. At the lowest level of most categorization techniques is a method for encoding an image in a way that is insensitive to the various perturbations that can occur in the image generation process (illumination, pose, viewpoint, and occlusions). It is well known in the field that SIFT descriptors are insensitive to illumination, robust to small changes in pose and viewpoint, and can be invariant to changes in scale and orientation.The SIFT descriptor is calculated within a circular image domain around a point at a specific scale, which determines the domain radius and the required image blur. After image blurring, the gradient orientation and gradient magnitude are determined, and a grid of spatial bins divides the circular image domain. The final descriptor is a normalized histogram (with Gaussian weighting decreasing from the center) of weighted gradient orientations separated by spatial bin. Thus, the descriptor has a size of 4 × 4 × 8 = 128 bins if the spatial bin grid is 4 × 4 and there are 8 orientation bins.While the locations, scales, and orientations of SIFT descriptors can be chosen in ways that are pose- and viewpoint-independent, most state-of-the-art categorization techniques use fixed scales and orientations and arrange the descriptors in a grid of overlapping domains. This not only accelerates performance, it also allows for very fast computation of all descriptors in an image.
[0081] For a visual category to be generalizable, there must be some visual similarity between members of the class and some dissimilarity compared to non-members. Furthermore, any large set of images will contain a wide variety of redundant data (walls, floor, etc.). This gives rise to the notion of "visual words"—a small set of prototype descriptors derived from the entire collection of training descriptors using a vector quantization technique such as k-means clustering. When the set of visual words—known as the codebook—is computed, images can be described solely in terms of which words appear where and with what frequencies. Here, k-means clustering is used to generate the codebook.This algorithm searches for k centers within the data space, each representing a collection of data points closest to it in that space. After the k cluster centers (the codebook) have been learned from training SIFT descriptors, the visual word of any new SIFT descriptor is simply the cluster center closest to it.
[0082] After an image has been decomposed into SIFT descriptors and visual words, these visual words can be used to generate a descriptor for the entire image, which is simply a histogram of all the visual words in the image. Optionally, images can be decomposed into spatial bins, and these image histograms can be spatially separated in the same way that SIFT descriptors are computed. This adds loose geometry to the process of learning actions from raw pixel information.
[0083] The final step of the visual category learning process is training a support vector machine (SVM) to distinguish among the classes given examples of their image histograms.
[0084] In the present context, the image-based technique can be used to recognize specific human actions such as bending, crouching, and reaching. Each "action" can comprise a collection of consecutive frames grouped together, and the system can use only the portion of an image in which the human of interest is present. Since there are multiple simultaneous views, the system can train one SVM per view, with each view's SVM scoring (or being trained on) each frame of an action. A vote count can then be calculated across all SVM frames across all views for a given action. The action is classified as the class with the highest total votes.
[0085] The system can then use the human tracker module to both determine where the person is at any time in any view and decide which frames are relevant for the classification process. First, the ground trajectories can be used to determine when the person performs an action of interest in the frame. Since the only way the person can move noticeably is by walking, any frames corresponding to large movements on the ground are assumed to contain images of the person walking. These images therefore do not need to be classified using the image-based categorization facility.
[0086] When analyzing a trajectory, long periods of little movement between periods of movement indicate frames in which the person performs an action other than walking. Frames corresponding to long periods of little movement are decomposed into groups, each of which constitutes an unknown action (or a labeled action if used for training). Within these frames, the human tracker provides a bounding box that specifies what portion of the image the person receives. As noted above, the bounding box can be specified in a corrected image space to enable more accurate training and recognition.
[0087] Once the frames and bounding boxes of interest have been identified by the human tracker, the procedure for training the SVMs is very similar to the conventional case. Within each action image bounding box, SIFT descriptors are computed across all frames and all views. Those images belonging to an action (i.e., that are temporally grouped together) are manually labeled within each view for SVM training. K-means clustering builds a codebook, which is then used to generate image histograms for each bounding box. The image histograms derived from a view are used to train their SVM. For example, in a system with six cameras, there are six SVMs, each classifying the three possible actions.
[0088] Given a new sequence, a number of unlabeled actions are extracted in the manner described above. These frames and bounding boxes are each classified using the appropriate view-based SVM. Each SVM generates scores for each frame of the action sequence. These scores are summed to calculate a cumulative score for the action across all frames and all views. The action (category) with the highest score is selected as the label for the action sequence.
[0089] At different times, the person may be occluded in a particular view while visible in others. Occluded views result in zero votes for all categories. Using one labeled sequence for training and four different sequences for testing results in increased accuracy. It is important to note that the same codebook developed during training is used at testing time, as otherwise the SVMs would not be able to classify the resulting image histograms.
[0090] The system can utilize a voxel-based reconstruction method that uses the moving foreground objects from the multiple views to reconstruct a 3D volume by projecting 3D voxels onto each of the image planes and determining whether the projection overlaps with the respective outlines of foreground objects. Once the 3D reconstruction is complete, the system can, for example, fit cylinder models to the different parts and use the parameters to train a classifier that estimates the human's pose.
[0091] The display and learning step in the block diagram from Fig. 6 are then combined with any external signals, such as those emitted by one or more additional tools (e.g., torque wrenches, nutrunners, screwdrivers, etc.), to form a spatiotemporal signature. This combined information is then used at 154 to construct AND-OR graphs. In general, AND-OR graphs can describe more complicated scenarios than a simple tree graph. The graph consists of two types of nodes: "Or" nodes, which are the same nodes in a typical tree graph, and "And" nodes, which allow a path along the tree to split into multiple simultaneous paths. This structure is used here to describe the acceptable sequences of actions that occur in a scene.In this context, the “And” nodes make it possible to describe events such as action A occurring, actions B and C occurring together, or D occurring, something that a standard tree graph cannot describe.
[0092] In another configuration, the system can use finite-state machines to describe user activity instead of the AND-OR graphs at 154. Finite-state machines are often used to describe systems with multiple states, along with the conditions for transitions between the states. After an activity recognition system temporally segments a sequence into elementary actions, the system can evaluate the sequence to determine whether it conforms to a set of confirmed action sequences. The set of confirmed sequences can also be learned from data, such as by constructing a finite-state machine (FSM) from training data and testing any sequence by passing it through the FSM.
[0093] Generating an FSM that represents the entire set of valid action sequences is straightforward. Given a set of training sequences (which have already been classified using the action recognition system), the nodes of the FSM are first generated by determining the union of all unique action labels across all training sequences. Once the nodes have been generated, the system can assign a direct edge from node A to node B if node B immediately follows node A in any training sequence.
[0094] Testing a given sequence is equally straightforward: The sequence is passed through the machine to determine whether it reaches the exit state. If it does, the sequence is valid; otherwise, it is not.
[0095] Furthermore, because the system knows the person's position when each activity is performed, it can incorporate spatial information into the structure of the FSM. This adds additional detail and the ability to evaluate an activity in terms of position, not just the sequence of events. Video synopsis
[0096] This video synopsis module 158 from Fig. 8 takes the input video sequences and represents dynamic activities in a very efficient and compact form for interpretation and archiving. The resulting synopsis maximizes information by showing multiple activities simultaneously. In one approach, a background view is selected, and foreground objects are extracted from selected frames and blended into the base view. Frame selection is based on the action labels obtained by the system and allows for the selection of those subsequences where a particular action of interest occurs. Multiple work spaces
[0097] The human monitoring system described herein thoroughly detects and monitors a person within the work zone from multiple different viewpoints, such that occlusion of a person in one or more of the views does not affect the tracking of the person. Furthermore, the human monitoring system can adjust and dynamically reconfigure the automated moving factory equipment to avoid potential interactions with the person within the work zone without requiring the automated equipment to stop. This may include determining and traversing a new walking path for the automated moving equipment.The human monitoring system can track multiple people within a work area and delegate tracking to other systems responsible for monitoring neighboring areas, with different zones being able to be defined for multiple locations within the work area.
[0098] Fig.Figure 9 shows a graphical representation of multiple workspaces. The sensing devices 12 for a respective workspace are coupled to a respective processing unit 18 dedicated to the respective workspace. Each respective processing unit identifies the proximity of and tracks people crossing into its respective workspace and communicates with another via a network connection 170 so that people can be tracked as they transition from one workspace to another. As a result, multiple visual monitoring systems can be linked to track people as they interact between the different workspaces.
[0099] It should be noted that the use of the vision monitoring system in a factory environment as described herein is only one example of where the vision monitoring system may be used, with this vision monitoring system having the capability of being applied in any application outside of a factory environment where the activities of people in an area are tracked and the movement and activity are logged.
[0100] The vision monitoring system is useful in the automated time and motion investigation of activities, which can be used to monitor function and provide data for use in improving the efficiency and productivity of work cell activity. This capability can also enable activity monitoring within a prescribed sequence, where deviations in the sequence can be identified and logged, and alerts can be generated to detect errors in tasks for humans. This "error protection" capability can be used to prevent task errors from propagating to downstream operations and causing quality and productivity problems due to errors in the sequence or proper material selection for the prescribed task.
[0101] It should be noted that a variation of the human monitoring capability of this system as described here is the monitoring of restricted-access areas that may contain essential activity involving automated or other equipment requiring only periodic service or access. This system would monitor the integrity of access controls to such areas and trigger alerts for unauthorized access. Since service or routine maintenance in this area may be necessary during off-shifts or other downtime, the system would monitor the authorized access and operations of a person (or persons) and trigger alerts locally and with a remote monitoring station if activity unexpectedly stops due to an accident or medical emergency. This capability could improve productivity for these types of tasks, where the system could be considered part of a "buddy system."
[0102] While the best modes for carrying out the invention have been described in detail, those skilled in the art to which this invention pertains will recognize various alternative designs and embodiments for practicing the invention within the scope of the appended claims. All matter contained in the above description or shown in the accompanying drawings is to be interpreted as illustrative and not restrictive.
Claims
[1] Human monitoring system (10) for monitoring a work area, the system (10) comprising: a plurality of cameras (12, 14, 16, 17) arranged somewhere in the work area, each camera (12, 14, 16, 17) being configured to record a video stream containing a plurality of individual images (60); a visual processor (18) configured to: to identify a presence of a person (64) in the work area from the plurality of individual images (60); to generate a motion tracking of the human (64) in the work area, wherein the motion tracking represents a position of the human (64) over a period of time; generate an activity log of one or more activities performed by the human (64) throughout the motion tracking; wherein the visual processor (18) is configured to: compare the motion tracking and activity log with an activity template that defines a plurality of required actions; each of the plurality of required actions is specified at a location in the workspace; and provide an alarm when one or more actions in the activity template are not executed in the workspace at their specified location; characterized by , that the system (10) is adapted to several different ways of executing a sequence of tasks and validates the motion tracking and the activity log as long as a final human walking path and the activities performed correspond to pre-specified targets at pre-specified locations in the work area. [2] The system (10) of claim 1, wherein the visual processor (18) is configured to generate the activity log of the one or more activities performed by the human (64) throughout the motion tracking by performing a pattern comparison of a pose of the human (64) against a trained pose database. [3] The system (10) of claim 2, wherein the pattern matching comprises using a support vector machine and / or a neural network. [4] The system (10) of claim 2, wherein the pattern comparison can further detect the pose of the human (64), and wherein the pose comprises standing, walking, stretching and / or crouching. [5] The system (10) of claim 1, wherein the alarm comprises a video synopsis of the generated motion tracking. [6] The system (10) of claim 1, wherein the visual processor (18) is configured to identify the presence of a human (64) in the work area by: detecting the human (64) in each of the plurality of individual images (60), wherein each of the plurality of individual images (60) is obtained from a different camera (12, 14, 16, 17) selected from the plurality of cameras (12, 14, 16, 17); maps a representation of the detected person (64) from several views into a common coordinate system; and an intersection point (87) of the depicted representations is determined. [7] System (10) according to claim 1, wherein the visual processor (18) is further configured to receive an actuation signal from a tool indicating the tool currently being used; and wherein the visual processor (18) is configured to use the actuation signal to confirm the performance of an activity of the one or more activities performed by the human (64). [8] Human monitoring system (10) for monitoring a work area, the system (10) comprising: a plurality of cameras (12, 14, 16, 17) arranged somewhere in the work area, each camera (12, 14, 16, 17) being configured to capture a video stream containing a plurality of individual images (60); a visual processor (18) configured to: to identify a presence of a person (64) in the work area from the plurality of individual images (60); to generate a motion tracking of the human (64) in the work area, wherein the motion tracking represents a position of the human (64) over a period of time; generate an activity log of one or more activities performed by the human (64) throughout the motion tracking by pattern matching a pose of the human (64) with a pose database; compare the motion tracking and activity log to an activity template that defines a plurality of required actions, each of the plurality of required actions being specified at a location in the workspace; and provide an alarm when one or more actions in the activity template are not executed in the workspace at their specified location; wherein the system (10) is adapted to several different ways of performing a sequence of tasks and validates the motion tracking and the activity log as long as a final human walking path and the activities performed conform to pre-specified targets at pre-specified locations in the work area. [9] System (10) according to claim 8, wherein the pattern matching comprises using a support vector machine and / or a neural network.
Citation Information
Patent Citations
Vision System for Monitoring Humans in Dynamic Environments
US20110050878A1
Motion recognition
US20120214594A1