Video surveillance system with cluster size estimation
Patent Information
- Application Number
- CN202210456420.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-05-14
- Filing Date
- 2022-04-27
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2042-04-27
Smart Images

Figure CN115346234B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates generally to video surveillance systems. More specifically, this disclosure relates to video surveillance systems capable of estimating crowd size using personnel detection algorithms. Background Technology
[0002] Multiple video surveillance systems employ cameras installed or otherwise deployed around monitored areas such as cities, parts of cities, facilities, or buildings. Video surveillance systems may also include mobile cameras, such as drones carrying cameras. Video surveillance systems may employ personnel detection algorithms to detect the presence of people within the video streams provided by the cameras. An improved approach is desired, using video streams from one or more cameras to estimate cluster sizes within a monitored area. Summary of the Invention
[0003] This disclosure relates to video surveillance systems, and more specifically to methods and systems for estimating the size of a crowd in a monitored area. In an example, a method is provided for estimating the number of people currently in a region of interest using a computing device having one or more processors. The region of interest is covered by a camera having a field of view (FOV) that captures the region of interest. An exemplary method includes receiving a video stream from the camera and analyzing the video stream to detect each of the multiple people in the region of interest and a measure of the size consumed by each of the multiple people detected in the FOV of the video stream. One or more adjacent pairs are identified from multiple possible pairs of people in the region of interest. In some cases, identifying one or more adjacent pairs includes, for each of the multiple possible pairs of people in the region of interest, comparing the size measure consumed by each of the pairs of people in the FOV of the video stream, and identifying the pairs of people as possible adjacent pairs when the size measures consumed by each of the pairs of people in the FOV of the video stream do not differ by more than a certain amount.
[0004] In some cases, identifying one or more neighbor pairs of a person further includes, for each of the possible neighbor pairs of a person, determining a measure of the distance between the person's counterparts in the field of view (FOV) of the video stream, and identifying the person's counterpart as one of the one or more neighbor pairs when the measure of the distance between the person's counterparts in the FOV of the video stream is less than a threshold distance, wherein the threshold distance depends on a measure of the size consumed by at least one of the person's counterparts in the FOV of the video stream. The identified one or more neighbor pairs are then clustered into one or more clusters. An estimated number of people in each of the one or more clusters is determined, wherein in some cases, the estimated number of people in each of the one or more clusters is based at least in part on the quantity of the identified one or more neighbor pairs associated with the corresponding cluster. A representation of the estimated number of people in each of the one or more clusters is then displayed on a monitor.
[0005] In another example, a method is provided for estimating the number of people currently in a region of interest using a computing device with one or more processors. The region of interest is covered by a camera having a field of view (FOV) that captures the region of interest. An exemplary method includes receiving a video stream from the camera and analyzing the video stream to detect each of a plurality of objects within the region of interest and a measure of the size consumed by each of the plurality of objects detected in the FOV of the video stream. One or more neighboring pairs of objects are identified from a plurality of possible pairs of objects within the region of interest. The detected plurality of objects are clustered into one or more clusters based at least in part on the identified one or more neighboring pairs. An estimated number of objects in each of the one or more clusters is estimated, wherein in some cases, the estimated number of objects in each of the one or more clusters is based at least in part on the quantity of the identified one or more neighboring pairs associated with the corresponding cluster. A representation of the estimated number of objects in each of the one or more clusters is displayed on a monitor.
[0006] In another example, a method is provided for estimating the number of people currently in a region of interest using a computing device with one or more processors. An exemplary method includes receiving a video stream from a camera and analyzing the video stream to find two or more individuals within the region of interest. Whether two of the two or more individuals are potentially eligible to be neighbors is determined at least in part based on their height in the field of view (FOV) between them. The exemplary method also includes determining, for individuals who are potentially eligible to be neighbors at least in part based on their height in the FOV between them, a distance in the FOV between the two individuals and an adaptive distance threshold relative to the two individuals, wherein in some cases, the adaptive distance threshold is at least in part based on the height in the FOV of at least one of the two individuals. When the distance is less than the adaptive distance threshold relative to the two individuals, a neighbor pair state is assigned to each of the two individuals. When the distance is greater than the adaptive distance threshold relative to the two individuals, a neighbor pair state is not assigned to each of the two individuals. The two or more individuals found within the region of interest are clustered into one or more clusters based at least in part on the assigned neighbor pair states.
[0007] The foregoing summary is provided to facilitate understanding of the innovative features unique to this disclosure and is not intended as a complete description. A full understanding of this disclosure can be obtained by considering the entire specification, claims, drawings, and abstract as a whole. Attached Figure Description
[0008] This disclosure can be more fully understood by considering the following description of various examples in conjunction with the accompanying drawings, in which:
[0009] Figure 1 This is a schematic block diagram of an exemplary video surveillance system;
[0010] Figure 2 It shows that it can be accessed via Figure 1 A flowchart illustrating the exemplary methods executed by an exemplary video surveillance system;
[0011] Figure 3 It shows that it can be accessed via Figure 1 A flowchart illustrating the exemplary methods executed by an exemplary video surveillance system;
[0012] Figure 4 It shows that it can be accessed via Figure 1 A flowchart illustrating the exemplary methods executed by an exemplary video surveillance system;
[0013] Figure 5 It shows that it can be accessed via Figure 1 A flowchart illustrating the exemplary methods executed by an exemplary video surveillance system;
[0014] Figure 6 This is a flowchart illustrating an exemplary method for determining adjacent pairs;
[0015] Figure 7 This is a schematic block diagram illustrating an example of actual neighbor pairs relative to the theoretical maximum number of neighbor pairs; and
[0016] Figure 8 This is a flowchart illustrating an example of applying a correction factor to the estimated number of people in a cluster.
[0017] While this disclosure is subject to various modifications and alternatives, its details have been shown by way of example in the accompanying drawings and will be described in detail. However, it should be understood that this disclosure is not intended to limit it to the specific examples described. Rather, it is intended to cover all modifications, equivalents, and alternatives that fall within the substance and scope of this disclosure. Detailed Implementation
[0018] The following description should be read with reference to the accompanying drawings, in which similar elements in different drawings are numbered in the same manner. The drawings are not necessarily drawn to scale and depict examples that are not intended to limit the scope of this disclosure. While examples of various elements are shown, those skilled in the art will recognize that many of the examples provided have suitable alternatives that can be utilized.
[0019] This document assumes that all numbers are modified by the term “about” unless otherwise explicitly stated. Expressions of numerical ranges using endpoints include all numbers contained within that range (e.g., 1 to 5 includes 1, 1.5, 2, 2.75, 3, 3.80, 4, and 5).
[0020] As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” include plural references, unless otherwise expressly stated. As used in this specification and the appended claims, the term “or” is generally used in its meaning to include “and / or,” unless otherwise expressly stated.
[0021] It should be noted that references to "one embodiment," "some embodiments," or "other embodiments" in the specification indicate that the described embodiments may include specific features, structures, or characteristics, but each embodiment may not necessarily include that specific feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Additionally, when a specific feature, structure, or characteristic is described in conjunction with an embodiment, it is conceivable that, whether explicitly described or not, that feature, structure, or characteristic may be applied to other embodiments, unless otherwise expressly stated otherwise.
[0022] Figure 1This is a schematic block diagram of an exemplary surveillance system 10 configured to provide monitoring of a monitored area (sometimes referring to a region of interest). The exemplary surveillance system 10 includes a plurality of cameras 12 (e.g., one or more) positioned within or otherwise covering at least a portion of the region of interest. In the example shown, the cameras are labeled 12a, 12b, 12c, and 12d, respectively. While a total of four cameras 12 are illustrated, it should be understood that the surveillance system 10 may include one camera, or hundreds or even thousands of cameras 12 in a setting such as a smart city. At least some of the cameras 12 may be fixed cameras (sometimes with pan, tilt, and / or zoom capabilities), meaning that these fixed cameras are each mounted in a fixed position. At least some of the cameras 12 may be mobile cameras configured to move back and forth within the monitored area. For example, at least some of the cameras 12 may be mounted within a drone configured to fly back and forth within the monitored area, thereby providing camera coverage at various locations and / or vertical positions within the monitored area. In some cases, mobile cameras can be driving cameras in emergency vehicles, body cameras of emergency personnel (such as police officers), and / or portable or wearable devices carried by citizens. These are just examples.
[0023] Each camera in camera 12 includes a field of view (FOV) 14. FOVs 14 are labeled 14a, 14b, 14c, and 14d, respectively. Some cameras in camera 12 may have a fixed FOV 14, which is affected by the camera's mounting position and method, the lens mounted on the camera, etc. Some cameras may have, for example, a 120-degree FOV or a 360-degree FOV. Some cameras in camera 12 may have an adjustable FOV 14. For example, some cameras in camera 12 may be pan-tilt-zoom (PTZ) cameras, which can adjust their FOV by adjusting one or more of the pan, tilt, and / or zoom of a particular camera 12.
[0024] The exemplary monitoring system 10 includes a computing device 16 configured to control at least some aspects of the operation of the monitoring system 10. For example, the computing device 16 may be configured to provide instructions to at least some of the cameras 12 to transmit video, such as, or to change one or more of the pan, tilt, and zoom functions of the camera 12, which is a PTZ camera. The computing device 16 may be configured to control the operation of one or more moving cameras that are part of the monitoring system 10. The computing device 16 includes a display 18, a memory 20, and a processor 22 operatively coupled to the display 18 and the memory 20. Although a single processor 22 is shown, it should be understood that the computing device 16 may include, for example, two or more different processors. In some cases, the computing device 16 may be an operator console, etc.
[0025] In some cases, camera 12 may communicate directly with computing device 16. In other cases, as shown, camera 12 may instead communicate directly with one or more edge devices 24. Two edge devices 24 are shown, labeled 24a and 24b, respectively. It should be understood that, for example, a substantially larger number of edge devices 24 may exist. As shown, cameras 12a, 12b, and 12c are operatively coupled to edge device 24a, and camera 12d (and possibly other cameras) is operatively coupled to edge device 24b. In some cases, edge devices 24 may provide some functionality that may otherwise be provided by computing device 16 and / or cloud-based server 26. When provided, cloud-based server 26 may be configured to send and receive information between edge devices 24 and computing device 16, and in some cases, provide processing capabilities to support the methods described herein.
[0026] In some cases, each edge device 24 may be an edge controller. In some cases, each edge device 24 may be configured to control the operation of each camera in the camera 12 operatively coupled to the specific edge device 24. The specific edge device 24 may be programmed or otherwise learn details relating to a specific camera 12 operatively coupled to the specific edge device 24. One or more of the cloud-based server 26, computing device 16, and / or edge device 24 may be configured to control at least some aspects of the operation of the monitoring system 11, such as providing instructions to at least some of the cameras 12 to transmit video, for example, or changing one or more of the pan, tilt, and zoom of the camera 12, which is a PTZ camera. One or more of the cloud-based server 26, computing device 16, and / or edge device 24 may be configured to control the operation of one or more mobile cameras that are part of the monitoring system 10. One or more of the cloud-based server 26, computing device 16, and / or edge device 24 may be configured to perform a number of different methods. Figures 2 to 5 This is a flowchart illustrating an exemplary method that can be coordinated by one or more of cloud-based servers 26, computing devices 16, and / or edge devices 24 and thus executed by monitoring system 10.
[0027] Figure 2 This is a flowchart illustrating an exemplary method 28 using a computing device (such as computing device 16, cloud-based server 26, and / or one or more edge devices from edge device 24) having one or more processors (such as processor 22) to estimate the number of people currently present in a region of interest. A camera (such as camera 12) covers the region of interest and has a field of view (FOV) that captures the region of interest. Exemplary method 28 includes receiving a video stream from the camera, as shown in box 30. The video stream is analyzed to detect each of the multiple people in the region of interest and a measure of the size consumed by each of the multiple people detected in the FOV of the video stream, as shown in box 32. Size can be any representation of the size of a person in the FOV, such as the height of a person, the width of a person, the height of a person's head, the width of a person's head, or any other suitable size measure in the FOV. Size measures can be expressed in any suitable unit, such as the number of pixels in the FOV. People closer to the camera are larger in the FOV of the video stream compared to people of similar size farther away from the camera. Therefore, even in the absence of any depth of field in the video stream, it can be determined with relatively good accuracy that two people of very different sizes in the FOV may not be close to each other and therefore should not be marked as adjacent pairs.
[0028] Identify one or more adjacent pairs from multiple possible pairs of multiple people within a region of interest, as shown in box 34. (Reference) Figure 3 An exemplary method for identifying one or more neighboring pairs is shown and described. The identified one or more neighboring pairs are then clustered into one or more clusters, as shown in box 36. Any of a variety of different clustering methods or algorithms can be used. As an example, density-based spatial clustering (DBSCAN) with noise application can be used to perform clustering. An estimated number of people in each of the one or more clusters is determined, wherein in some cases, the estimated number of people in each of the one or more clusters is based at least in part on the quantity of the one or more identified neighboring pairs associated with the corresponding cluster, as shown in box 38. A representation of the estimated number of people in each of the one or more clusters is then displayed on a monitor, as shown in box 40.
[0029] In some cases, the estimated number of people in each of one or more clusters can be determined, at least in part, based on the number of possible pairs of people associated with the identified adjacent pairs associated with the corresponding cluster.
[0030] In some cases, method 28 may further include overlaying a representation of each of one or more clusters of the video stream onto the display. Method 28 may further include overlaying a representation of the estimated number of people in each of the one or more clusters onto the display.
[0031] In some cases, such as indicated in box 42, method 28 may further include monitoring the estimated number of people in each of one or more clusters. When the estimated number of people in one or more clusters exceeds a cluster size threshold, an alert may sometimes be issued to the operator after the exceedance has persisted for at least a time interval exceeding a hold time threshold, as shown in box 44. For example, the hold time threshold may be a user-definable parameter. For instance, the hold time threshold may be set to one minute or five minutes.
[0032] Figure 3 This is a flowchart illustrating an example of identifying adjacent pairs, such as... Figure 2As mentioned in box 34. In this example, for each of the multiple possible pairs of people within a region of interest, identifying adjacent pairs involves several steps, as shown in box 35. Specifically, the size measure consumed by each person in the FOV is compared, as shown in box 37. When the size measure consumed by each person in the FOV of the video stream does not differ by a predetermined amount, the person pair is identified as a possible adjacent pair, as shown in box 39. In some cases, the size measure consumed by each person in the FOV of the video stream will not exceed a predetermined amount when the ratio (R) of the size measure consumed by the smaller person in the person pair to the size measure consumed by the larger person in the person pair is greater than a threshold ratio. As mentioned above, a person closer to the camera is larger in the FOV of the video stream than a similarly sized person farther away from the camera. Therefore, even without any depth of field in the video stream, it can be determined with relatively good accuracy that two people of very different sizes in the FOV may not be close to each other and therefore should not be marked as adjacent pairs. By determining a simple ratio (R) between the magnitude measure consumed by the smaller of the human pair and the magnitude measure consumed by the larger of the human pair, and determining whether this ratio (R) is greater than a threshold ratio, many possible human pairs can be easily eliminated from the considered possible neighbor pairs using a simple algorithm that does not consume significant processing resources.
[0033] For each of the possible adjacent pairs of people, as shown in box 41, several steps can be performed. In the example shown, a measure of the distance in the FOV of the video stream between the pairs of people is determined, as shown in box 43. When the measure of the distance between the pairs of people in the FOV of the video stream is less than a threshold distance, the pairs of people are identified as one or more adjacent pairs of people, where in some cases the threshold distance depends on a measure of the size consumed by at least one of the pairs of people in the FOV of the video stream, as shown in box 45. In some cases, the threshold distance depends on a measure of the size consumed by the smallest of the pairs of people in the FOV of the video stream, sometimes taken into account by programmable parameters.
[0034] A pair of people closer to the camera will appear larger in the field of view (FOV) of the video stream compared to a similarly sized pair further away from the camera. Therefore, a distance metric in the FOV (e.g., in pixels) between a pair of people further away from the camera will represent a greater physical separation between them compared to a similar distance metric in the FOV (e.g., in pixels) between a pair of people closer to the camera. Thus, a threshold distance is provided, which depends on a measure of the size consumed by at least one of the pairs of people (e.g., the smallest person) in the FOV of the video stream (as shown in box 45). This results in an adaptive threshold that automatically adjusts or compensates for the actual physical separation between the pair of people depending on the distance of that pair from the camera.
[0035] Figure 4 This is a flowchart illustrating an exemplary method 48 using a computing device (such as computing device 16, cloud-based server 26, and / or edge device 24) having one or more processors (such as processor 22) to estimate the number of objects currently present in a region of interest. A camera (such as camera 12) covers the region of interest and has a field of view (FOV) that captures the region of interest. Method 48 includes receiving a video stream from the camera, as shown in box 50. Analyzing the video stream to detect each of a plurality of objects within the region of interest and a measure of the size consumed by each of the plurality of objects detected in the FOV of the video stream, as shown in box 52. The plurality of objects may include one or more people. The plurality of objects may include any of a variety of inanimate objects, such as boxes, containers, etc. For example, the plurality of objects may include one or more vehicles. In some cases, the plurality of objects may include one or more of the following: people, animals, cars, boats, aircraft, robots, and drones, or some combination thereof.
[0036] Identify one or more neighbor pairs of objects from multiple possible pairs of objects within a region of interest, as shown in box 54. Cluster the detected multiple objects into one or more clusters, at least in part, based on the identified one or more neighbor pairs, as shown in box 56. In some cases, only the identified neighbor pairs of objects are clustered, without clustering those detected objects that were not identified as neighbor pairs.
[0037] Determine the estimated number of objects in each of one or more clusters, wherein the estimated number of objects in each of the one or more clusters is based at least in part on the quantity of one or more identified adjacent pairs of objects associated with the corresponding cluster, as shown in box 58. Display a representation of the estimated number of objects in each of the one or more clusters on a display, as shown in box 60. In some cases, different colors may be used to indicate clusters with different numbers of objects. As an example, clusters with relatively fewer objects may be displayed in yellow, while clusters with relatively more objects may be displayed in red. It should be understood that any of a variety of different colors may be used. In some cases, for example, a particular color may be used to display clusters of objects suspected of being potentially hazardous.
[0038] In some cases, identifying one or more adjacent pairs includes, for each of a plurality of possible pairs of objects within a region of interest, comparing a measure of the size consumed by each of the pairs of objects in the FOV of the video stream, and identifying the pair of objects as a possible adjacent pair when the measures of the size consumed by each of the pairs of objects in the FOV of the video stream do not differ by a predetermined amount. Identifying one or more adjacent pairs may include, for each pair or more pairs of objects, determining a measure of the distance between the pairs of objects in the FOV of the video stream, and identifying the pair of objects as one of the one or more adjacent pairs of objects when the measure of the distance between the pairs of objects in the FOV of the video stream is less than a threshold distance, wherein in some cases, the threshold distance depends on a measure of the size consumed by at least one pair of the pairs of objects in the FOV of the video stream.
[0039] Figure 5This is a flowchart illustrating an exemplary method 62 for estimating the presence of multiple people in a region of interest, the region of interest being covered by a camera (such as camera 12) having a field of view (FOV) that captures the region of interest. Method 62 includes receiving a video stream from the camera, as shown in box 64. The video stream is analyzed to find two or more people within the region of interest, as shown in box 66. It is determined, at least in part, based on the height in the FOV between the two people, whether two of the two or more people (e.g., a pair) are potentially eligible to be a neighbor pair, as shown in box 68. For people who are potentially eligible to be neighbors, at least in part based on the height in the FOV between the two people, method 62 includes determining several things, as shown in box 70. Box 70 includes determining the distance in the FOV between the two people, as shown in box 72. Box 70 includes determining an adaptive distance threshold relative to the two people, wherein the adaptive distance threshold is at least in part based on the height in the FOV of at least one of the two people (e.g., the smallest person), as shown in box 74. Box 70 includes determining, when the distance is less than an adaptive distance threshold relative to the two individuals, to assign an adjacent pair state to each of the two individuals, as shown in Box 76. Individuals found within the region of interest are clustered into one or more clusters based at least in part on the assigned adjacent pair states of the two or more individuals found within the region of interest, as shown in Box 78. In some cases, only the identified adjacent pairs of individuals are clustered, without clustering those individuals detected who were not identified as adjacent pairs.
[0040] In some cases, method 62 may further include determining an estimated number of people in each of one or more clusters, wherein the estimated number of people in each of one or more clusters is relative to the number of possible pairs between people in the corresponding cluster, and is based at least in part on the number of assigned adjacent pairs between people in the corresponding cluster, as shown in box 80.
[0041] Figure 6This is a flowchart illustrating an exemplary method 82 for determining which detected people are considered neighbor pairs. In some cases, every pair of detected people is considered (e.g., all possible pairs of people detected in a single frame of a video stream), as shown in box 86. A parameter R is calculated at box 88. Parameter R provides the ratio between the height of the shorter person and the height of the taller person in the considered pair. It should be understood that the difference in distance from the camera is reflected in the difference in relative height, particularly in video images that may be side-view or perspective images. At decision box 90, it is determined whether the calculated parameter R is greater than 0.5 or greater than another threshold. If the calculated parameter R is not greater than 0.5 or another threshold, it means that the two people in the considered pair may not be considered neighbors because they are too far apart (in the depth direction), and control returns to box 86, where another pair is considered. By determining the simple ratio (R) and whether the ratio (R) is greater than a threshold ratio, many possible pairs of people can be easily eliminated from the considered possible neighbor pairs using a simple algorithm that does not consume significant processing resources.
[0042] If the answer at decision box 90 is yes, and the calculated parameter R is greater than 0.5 or another threshold, this means that the two people may be considered neighbors because they are close enough together (in the depth direction). In this case, control proceeds to box 92 for the next level of analysis. The distance D between the pair of people in the FOV is calculated as shown in box 94, and the distance threshold D1 is calculated as shown in box 96. At decision box 98, it is determined whether D is less than D1. If not, the pair is determined not to be neighbors, and control returns to box 86, where another pair is considered. However, if the answer is yes, and D is less than D1, control proceeds to box 100, where it indicates that the pair is neighbors.
[0043] D1 represents an adaptive threshold. A pair of people closer to the camera appears larger in the field of view (FOV) of the video stream compared to a similarly sized pair farther away. Therefore, the distance measure in the FOV (e.g., in pixels) between a pair of people farther away from the camera will represent a greater physical separation between them compared to a similar distance measure in the FOV (e.g., in pixels) between a pair of people closer to the camera. Thus, a threshold distance (D1) is provided, which depends on the size measure consumed in the FOV of the video stream by at least one of the pairs of people (e.g., the smallest person) (as shown in box 46), and is sometimes influenced by the programmable parameter “e”. This results in an adaptive threshold that automatically adjusts or compensates for the actual physical separation between a pair of people depending on the distance of that pair from the camera.
[0044] Figure 7A graphical representation is provided of the difference between the actual number of relationships (e.g., adjacent pairs) and the theoretical maximum number of relationships. Case 1 shows five objects (possibly people, vehicles, etc.) labeled 102, 104, 106, 108, and 110. As shown, objects 102 and 104 can be considered neighbors. Objects 104 and 106 are neighbors, as are objects 106 and 108, and objects 108 and 110. Object 102 is also a neighbor of objects 106, 108, and 110. However, objects 104 and 108 are not neighbors. Objects 106 and 110 are not neighbors. This can be contrasted with Case 2, which shows the maximum number of relationships between people assigned to a cluster. In this case, each object 102, 104, 106, 108, and 110 is considered a neighbor of every other object 102, 104, 106, 108, and 110 in the cluster. For a total of "n" objects, the theoretical maximum number of relationships (such as being neighbors) can be given by the following formula:
[0045]
[0046] Figure 8 This is a flowchart illustrating an exemplary method 112 for calculating a correction factor used to correct for the number of people or objects in a cluster. In some cases, some people physically present in a monitored area may be obstructed from the camera's field of view by, for example, other people. This is especially true when a crowd gathers in the area. To correct for this, exemplary method 112 calculates a correction factor used to correct for the number of people (or objects) in a cluster. Method 112 begins at box 114, which represents the output of a clustering algorithm or method. At box 116, the cluster density is calculated as the ratio of the number of existing relationships (e.g., the number of adjacent pairs) to the number of maximum relationships (the number of possible pairs of cluster members), such as... Figure 7 As shown in box 118, the correction factor is calculated as the cluster density multiplied by the number of people in the cluster. As indicated in box 120, the estimated number of people in the cluster can be calculated as the sum of the number of people in the cluster and the correction factor. As shown in box 122, the output can include a range of people, ranging from the number reported by the clustering algorithm or method to the corrected number.
[0047] Although several illustrative embodiments of this disclosure have been described thus, those skilled in the art will readily understand that other embodiments can be made and used within the scope of the appended claims. However, it should be understood that this disclosure is illustrative in many respects only. Changes may be made to details, particularly those relating to shape, size, arrangement of parts, and exclusion and order of steps, without departing from the scope of this disclosure. The scope of this disclosure is, of course, defined by the language expressed in the appended claims.
Claims
1. A computer-implemented method for estimating the number of people currently present in a region of interest, said region of interest being covered by a camera having a field of view (FOV) that captures said region of interest, wherein people closer to the camera are larger in the FOV of a video stream than people of similar size farther from the camera, said method comprising: Receive video stream from the camera; Analyze the video stream to detect each of the multiple people in the region of interest and a measure of the size of each of the multiple people detected in the FOV of the video stream; Identifying one or more adjacent pairs from multiple possible pairs of the plurality of people within the region of interest, wherein identifying the one or more adjacent pairs includes: For each of the multiple possible pairs of people in the region of interest, compare the size measure of each person in the pair in the FOV of the video stream, and identify the pair of people as possible adjacent pairs when the size measure of each person in the pair in the FOV of the video stream does not differ by a predetermined amount or when the ratio of the size measure of the smaller person in the pair to the size measure of the larger person in the pair is greater than a threshold ratio; For each of the possible adjacent pairs of the person, a measure of the distance between the pairs of the person in the FOV of the video stream is determined, and when the measure of the distance between the pairs of the person in the FOV of the video stream is less than a threshold distance, the pairs of the person are identified as a pair of one or more adjacent pairs of the person, wherein the threshold distance depends on a measure of the size of at least one of the pairs of the person. Based on the assigned neighbor pair states, people in the identified pairs of people are clustered into one or more clusters, without clustering those people who are detected but not identified as part of any neighbor pair; Determine an estimated number of people in each of the one or more clusters, wherein the estimated number of people in each of the one or more clusters is determined based on the number of assigned adjacent pair state relations among the people in the corresponding cluster relative to the theoretical maximum number of pair relations in the cluster; and Display a representation of the estimated number of people in each of the one or more of the said clusters on the monitor.
2. The method of claim 1, wherein the measure of the size of each of the multiple people detected in the FOV of the video stream corresponds to one or more of the following: the height of each of the multiple people detected in the FOV of the video stream, the width of each of the multiple people detected in the FOV of the video stream, or the width of the head of each of the multiple people detected in the FOV of the video stream.
3. The method of claim 1, further comprising overlaying a representation of each of the one or more clusters on the video stream onto the display.
4. The method of claim 3, further comprising overlaying a representation of the estimated number of people in each of the one or more clusters on the display.
5. The method of claim 1, further comprising: An alert is sent to the operator when the estimated number of people in one or more of the clusters exceeds the cluster size threshold.
6. A computer-implemented method for estimating the number of objects currently present in a region of interest, said region of interest being captured by a camera having a field of view (FOV) covering said region of interest, wherein objects closer to the camera are larger in the FOV of the video stream than similarly sized objects farther from the camera, said method comprising: Receive video stream from the camera; The video stream is analyzed to detect each of a plurality of objects within the region of interest, and a measure of the size of each of the plurality of objects detected in the FOV of the video stream. Identify one or more adjacent pairs of objects from a plurality of possible pairs of objects within the region of interest, wherein identifying the one or more adjacent pairs includes, for each of the plurality of possible pairs of objects within the region of interest, comparing the size measure of each of the pairs of objects in the FOV of the video stream, and identifying the pairs of objects as possible adjacent pairs when the size measure of each of the pairs of objects in the FOV of the video stream differs by no more than a predetermined amount or when the ratio of the size measure of the smaller of the pairs of objects to the size measure of the larger of the pairs of objects is greater than a threshold ratio. Based on the assigned neighbor pair states, objects in the identified pairs of the objects are clustered into one or more clusters, without clustering those detected objects that are not identified as part of any neighbor pair. Determine an estimated number of objects in each of the one or more clusters, wherein the estimated number of objects in each of the one or more clusters is determined based on the number of adjacent pair state relations assigned to objects in the corresponding cluster relative to the theoretical maximum number of pair relations in the cluster. as well as Display a representation of the estimated number of objects in each of one or more of the said clusters on a monitor. Identifying the one or more adjacent pairs includes, for each pair or more objects, determining a measure of the distance between the pairs of objects in the FOV of the video stream, and identifying the pairs of objects as one of the one or more adjacent pairs of the objects when the measure of the distance between the pairs of objects in the FOV of the video stream is less than a threshold distance, wherein the threshold distance depends on a measure of the size of at least one of the pairs of objects in the FOV of the video stream.
7. The method of claim 6, wherein the plurality of objects includes one or more of the following: a person, an animal, a car, a ship, an aircraft, a robot, and a drone.