System, method and storage medium for judging personnel entry and exit of field
By setting up cameras at high places and defining event detection zones, and using head reference points to determine the entry and exit of people, the problem of misjudgment in crowd flow calculations in crowded situations is resolved, achieving higher judgment accuracy.
Patent Information
- Application Number
- CN202110710152.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-01-05
- Filing Date
- 2021-06-25
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2041-06-25
AI Technical Summary
Existing crowd flow calculation technology has difficulty accurately judging the entry and exit of people in crowded situations, especially the inability to effectively distinguish between multiple people entering and exiting at the same time or non-human objects, resulting in inaccurate calculations.
A camera is set up at a high altitude to capture entrances and exits at a low angle. The processor sets the event detection area and detects human images. The head reference point is used to determine whether a person has passed through the entrance or exit. The upper, lower, left, right and bottom boundaries are set to reduce false positives.
It improves the accuracy of crowd flow calculation, reduces the misjudgment rate in crowded situations, and avoids misjudgment caused by height differences and image obstruction.
Smart Images

Figure CN114202735B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image analysis technology, and more particularly to a system, method and storage medium for determining whether a person enters or exits a venue. Background Art
[0002] People counting technology has been widely used in many practical applications, such as counting people at entrances and exits of public places and managing access to restricted areas. Crowd-flow data from these scenarios can provide useful information for security, marketing operations, service quality, and resource allocation. Traditionally, common methods for counting people entering and exiting entrances and exits include manual counting or installing sensors such as infrared and radio frequency identification (RFID) at the entrances and exits. Manual counting requires manpower, infrared counting cannot distinguish between multiple people entering and exiting simultaneously or non-human objects, and RFID counting requires personnel to carry RFID tags. However, manual counting can result in undercounts due to fatigue, while infrared or RFID counting can also result in undercounts due to large numbers of people entering and exiting or people not carrying RFID tags. Consequently, these methods can lead to inaccurate calculations.
[0003] With technological advancements, the images captured by cameras are no longer limited to image storage. They can also be further analyzed to detect, track, and locate objects within the images. Therefore, existing crowd counting technologies have begun to employ automated detection methods based on analyzing camera images. Generally, before using camera images to detect people entering or exiting a venue, a boundary must be set in the camera image corresponding to the venue's entrance or exit. The number of people is then counted based on the number of feet or heads crossing the boundary. However, in crowded conditions, accurately detecting people entering or exiting a venue based on their feet is often impossible due to the lower body being obscured, preventing the foot image from appearing in the image. Furthermore, detecting people entering or exiting a venue based on their heads is less tolerant of height differences. Therefore, developing technologies that can more accurately count people entering or exiting a venue is a topic that those skilled in the art are striving to address. Summary of the Invention
[0004] The present invention provides a system, method and storage medium for judging whether a person enters or leaves a venue, which can reduce the misjudgment rate of detecting whether a person enters or leaves a venue and improve the judgment accuracy.
[0005] One embodiment of the present invention provides a system for determining whether a person enters or exits an area. The system includes a camera and a processor. The camera is mounted at a height to capture an entrance or exit at a low angle and simultaneously output a video stream. The processor is coupled to the camera. The processor is configured to receive the video stream and set an event detection zone corresponding to the entrance or exit, wherein the event detection zone includes an upper boundary, a lower boundary, and an inner area, and the lower boundary includes a left boundary, a right boundary, and a bottom boundary. The processor is configured to detect and track a person image corresponding to a person in the video stream, wherein the person image can be a full-body image or a partial image. The processor is configured to determine whether the person has passed or not passed through the entrance or exit based on a first detection result and a second detection result. When the first detection result indicates that the coordinate position of the candidate area corresponding to the person image first appears within the inner area, and the second detection result indicates that the candidate area leaves the event detection zone through the lower boundary, the processor determines that the person has passed through the entrance or exit. When the first detection result indicates that the candidate area enters the inner area from outside the event detection area through the lower boundary, and the second detection result indicates that the coordinate position where the candidate area disappears is located in the inner area, the processor determines that the person passes through the entrance and exit.
[0006] One embodiment of the present invention provides a method for determining whether a person has entered or exited an area, applicable to a system comprising a camera and a processor, wherein the camera is mounted at a high location to capture an entrance or exit at a low angle and simultaneously output a video stream, wherein the processor receives the video stream. The method comprises the following steps: setting an event detection zone corresponding to the entrance or exit, wherein the event detection zone includes an upper boundary, a lower boundary, and an inner region, wherein the lower boundary includes a left boundary, a right boundary, and a bottom boundary; detecting and tracking a person image corresponding to a person in the video stream, wherein the person image may be a full-body image or a partial image; and determining whether the person has passed or not passed through the entrance or exit based on a first detection result and a second detection result. The processor determines that the person has passed through the entrance or exit when the first detection result indicates that the coordinate position of the candidate region corresponding to the person image first appears within the inner region, and the second detection result indicates that the candidate region has left the event detection zone through the lower boundary. When the first detection result indicates that the candidate area enters the inner area from outside the event detection area through the lower boundary, and the second detection result indicates that the coordinate position where the candidate area disappears is located in the inner area, the processor determines that the person passes through the entrance and exit.
[0007] One embodiment of the present invention provides a non-volatile computer storage medium. The non-volatile computer storage medium includes at least one program instruction. When an electronic device loads and executes the at least one program instruction, the above-mentioned method for determining whether a person enters or exits a zone can be performed.
[0008] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments, but this does not limit the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 This is a block diagram of a system for determining whether a person enters or exits a venue according to an embodiment of the present invention;
[0010] Figure 2 A schematic diagram of a scenario according to an embodiment of the present invention;
[0011] Figure 3 This is a flow chart of a method for determining whether a person enters or exits a venue according to an embodiment of the present invention;
[0012] Figure 4 A schematic diagram of an event detection area according to an embodiment of the present invention;
[0013] Figure 5 A schematic diagram of a candidate region according to an embodiment of the present invention;
[0014] Figures 6A-6C is a schematic diagram of an event based on a first entry method according to an embodiment of the present invention;
[0015] Figures 6D-6F is a schematic diagram of an event based on the second entry method according to an embodiment of the present invention;
[0016] Figures 6G-6I is a schematic diagram of an event based on the third entry method according to an embodiment of the present invention;
[0017] Figure 7 The flowchart of the method for determining whether a person enters or exits an area according to an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0018] Reference will now be made in detail to exemplary embodiments of the present invention, examples of which are illustrated in the accompanying drawings. Whenever possible, the same reference numerals are used in the drawings and the description to refer to the same or like parts.
[0019] The foregoing and other technical aspects, features, and benefits of the present invention will be more clearly understood in the following detailed description of a preferred embodiment with reference to the accompanying drawings. Directional terms such as up, down, left, right, front, and back, used in the following embodiments, are intended solely to refer to the directions in the accompanying drawings. Therefore, these directional terms are intended to illustrate and not to limit the present invention.
[0020] Figure 1 This is a block diagram of a system for determining whether a person enters or exits a field according to an embodiment of the present invention. Figure 1 The system 100 for determining whether a person enters or exits a venue includes a camera 110 , a memory 120 , and a processor 130 .
[0021] Camera 110 is used to capture images and may be a camera using a charge coupled device (CCD), a complementary metal-oxide semiconductor (CMOS) component, or other component lens. Alternatively, camera 110 may be an image capture device capable of capturing depth information, such as a depth camera or a stereoscopic camera. Camera 110 may be implemented using any model and brand of camera, and the present invention is not limited thereto.
[0022] The memory 120 is used to store various program codes and data required for execution of the system 100. The memory 120 may be implemented, for example but not limited to, any form of fixed or removable random access memory (RAM), read-only memory (ROM), flash memory, hard disk drive (HDD), solid state drive (SSD), or similar devices or combinations thereof, and the present invention is not limited thereto.
[0023] The processor 130 is, for example, a central processing unit (CPU), or other programmable general purpose or special purpose microprocessors (Microprocessor), digital signal processors (DSP), programmable controllers, application specific integrated circuits (ASIC), programmable logic devices (PLD), or other similar devices or combinations of these devices, and the present application is not limited thereto. The processor 130 is connected to the camera 110 and the memory 120 to receive images from the camera 110, access program codes and data of the memory 120, operate and process data, etc., to complete various operations required by the system 100. In addition, the processor 130 can drive a display or a network interface to display image and data operation processing results on the display or transmit to the network according to specific requirements, and the present application is not limited thereto.
[0024] In an embodiment of the present application, the camera 110 of the system 100 is connected to the processor 130 in an external manner by wire or wirelessly, or the memory 120 and the processor 130 are configured in the camera 110 and connected to various electrical elements of the camera 110, and the present application is not limited thereto.
[0025] Figure 2 A schematic diagram of an embodiment of the present application. Please refer to Figure 2 The camera 110 can be arranged at a high place to shoot the entrance 5 at a downward angle and simultaneously output a video stream corresponding to a field of view (FOV) 20 of the camera 110 to the processor 130. At this time, the optical axis of the camera 110 is directed below the horizontal line and forms an angle with the horizontal line to shoot the entrance 5. This arrangement is commonly used in general field entrance monitors. In this embodiment, the left and right sides of the entrance 5 are separated from the inside and outside of the field by an opaque wall. The camera 110 can be arranged inside or outside the field. The opaque wall refers to an opaque object that can block the line of sight on the left and right sides of the entrance 5, and is not limited to a vertical wall.
[0026] Figure 3 A flowchart of a method for determining whether a person enters or exits a field according to an embodiment of the present application. The method of this embodiment is applicable to at least the above-mentioned system 100, but the present application is not limited thereto. Please refer to Figures 1 to 3 The detailed steps of the method for determining whether a person enters or exits a field according to this embodiment are described below in combination with various devices and elements of the system 100.
[0027] At step S302, the processor 130 is configured to receive the video stream and set an event detection zone corresponding to the doorway. Specifically, the processor 130 sets the event detection zone corresponding to the image frame in advance. The event detection zone includes a boundary and an inner region, which is, for example, a rectangle, a trapezoid, a polygon, an ellipse, or other geometric shapes, and the present application is not limited thereto. In addition, the processor 130 sets the event detection zone according to the body parameters of the person in the real world and the boundary of the doorway and the wall. For example, the processor 130 can receive the body parameters of the person input by the user through an input device (not shown in the figure) and automatically identify the doorway in the image frame to calculate the event detection zone in the image frame. Alternatively, the processor 130 can read the coordinate setting file of the event detection zone from the memory 120, and the present application is not limited thereto.
[0028] Figure 4 A schematic diagram of the event detection zone of an embodiment of the present application. Please refer to Figure 4 When the camera is set at a high place to shoot the doorway 5 at a downward angle, the image frame 401 included in the video stream can cover the doorway 5 and the wall surface 4011, the floor 4012, and other scenes. The wall surface 4011 is an opaque wall that can block the person outside the doorway 5. Since the camera is shooting at a downward angle, if the doorway 5 is rectangular in the real world, it will appear as a trapezoid in the image frame 401, which is wider at the top and narrower at the bottom, as shown in Figure 4 In this embodiment, the processor 130 sets the event detection zone 40 corresponding to the doorway 5. For example, the event detection zone 40 includes an upper boundary 41, a lower boundary, and an inner region 45. The inner region 45 is a closed region enclosed by the upper boundary 41 and the lower boundary, and the lower boundary includes a left side boundary 42, a right side boundary 43, and a bottom boundary 44. In this embodiment, the upper left corner coordinate of the event detection zone 40 is (x1, y1), the upper right corner coordinate is (x4, y1), the lower left corner coordinate is (x2, y2), and the lower right corner coordinate is (x3, y2). It should be noted that the lower boundary can also be set as a semicircular arc or any line that can enclose a closed region with an area with the upper boundary, and the present application is not limited thereto.
[0029] In addition, the processor 130 sets the event detection zone 40 based on the physical parameters of people in the real world and the boundary between the entrance and exit 5 and the wall. Specifically, the processor 130 sets the height position of the upper boundary 41 in the image frame 401 based on a first preset height, and sets the height position of the bottom boundary 44 in the image frame 401 based on a second preset height. In this embodiment, the first preset height is the height of the person with the tallest height expected to be detected (e.g., 200 cm). The second preset height is the height of the person with the shortest height expected to be detected (e.g., 100 cm). The processor 130 sets the height position of the upper boundary 41 in the image frame 401 to be higher than the height corresponding to the first preset height in the image frame 401, so that anyone (even the tallest person) standing in front of the entrance or exit or inside the wall (i.e., on the same side as the camera) will have their corresponding head reference point lower than the upper boundary 41. Furthermore, the processor 130 sets the height of the bottom boundary 44 in the image frame 401 to be lower than the height corresponding to the second predetermined height in the image frame 401, so that anyone (even the shortest person) standing behind the entrance or outside the wall (i.e., on a different side from the camera) will have their head reference point higher than the bottom boundary 44. In other words, the processor 130 determines the first predetermined height and the second predetermined height based on the height range of people in the real world.
[0030] In addition, refer to Figure 4 , the entrance and exit 5 may include a left wall and a right wall, wherein the left wall corresponds to a left edge line, and the right wall corresponds to a right edge line. The processor 130 may set the left boundary 42 and the right boundary 43 based on the left edge line and the right edge line. In this embodiment, the processor 130 sets the left boundary 42 to the left of the left edge line and the distance from the left edge line to the preset head width range d, and sets the right boundary 43 to the right of the right edge line and the distance from the right edge line to the preset head width range d. Here, the preset head width range d is between the width w (for example, 30 pixels) corresponding to the head of the person in the image frame 401 multiplied by the maximum multiple and the minimum multiple. In one embodiment, the preset head width range is: 0.5w<d<4w.
[0031] It should be noted that the height or head width of the person's image, the height corresponding to the first preset height and the second preset height in the image frame, and the preset head width range can be expressed in distance units (e.g., centimeters or millimeters) or in image pixels, but the present invention is not limited thereto. Processor 130 automatically converts all parameters and variables into the same unit. For example, in one embodiment, 1 centimeter is equivalent to 37.795275591 pixels. Processing unit 130 can thus convert all parameters and variables between centimeters and pixels.
[0032] In one embodiment, the processor 130 marks the event detection zone 40 in the image frame and displays it on the display. In another embodiment, the processor 130 also marks the first predetermined height and the second predetermined height in the image frame and displays them on the display. In other words, the processor 130 can mark any geometric shape and parameters associated with the event detection zone 40 in the image frame, and the present invention is not limited thereto.
[0033] In step S304, processor 130 is configured to detect and track human images corresponding to people in the video stream. Specifically, processor 130 reads consecutive image frames from the video stream and detects and tracks candidate regions corresponding to human images in the consecutive image frames. Processor 130 may perform human figure detection on the image frames to define the candidate regions, and define the top center point of the candidate regions as the head reference point. In this embodiment, processor 130 may detect full-body images or partial images (e.g., heads) of people, and the number of detected people may be one or more.
[0034] Specifically, processor 130 may perform human detection using computer vision or deep learning models to detect human images in image frames. For example, a deep learning model may be implemented using a learning network such as a convolutional neural network (CNN), but the present invention is not limited thereto. A convolutional neural network is composed of at least one convolution layer, at least one pooling layer, and at least one fully connected layer. The front end of a convolutional neural network typically consists of a convolutional layer and a pooling layer connected in series or in parallel, and is used to obtain feature values of the image. These feature values can be multidimensional arrays and can be considered feature vectors representing the image. The back end of the convolutional neural network includes a fully connected layer. The fully connected layer classifies objects in the image based on the feature values generated by the convolutional and pooling layers, and can obtain object information corresponding to the identified objects. The object information includes a bounding box used to enclose the identified object, and the object information also includes the type of the identified object. In this embodiment, the convolution operation method can be implemented using any convolution operation steps known in the art, and the present invention is not limited thereto. The detailed steps and implementation methods can be adequately taught, suggested, and implemented based on common knowledge in the art, and thus will not be elaborated upon.
[0035] Then, after detecting the image of the person in the image frame, the processor 130 may define a candidate area corresponding to the image of the person. For example, the processor 130 may define a bounding box as the candidate area. The processor 130 sets the candidate area to be associated with a specific person image, and the size of the candidate area is at least sufficient to surround the corresponding person image. Among them, the candidate area may include a head reference point. If the candidate area corresponds to a full-body image of the person, the position of the head reference point may be the center point of the upper boundary of the candidate area or any point on the upper boundary. If the candidate area corresponds to a head image of the person, the position of the head reference point may be the center point of the upper boundary of the candidate area, any point on the upper boundary or the center point of the candidate area, but the present invention is not limited to this.
[0036] Figure 5 This is a schematic diagram of a candidate area according to an embodiment of the present invention. Figure 5 The image frame 501 includes a person image 50 and a candidate region 51 corresponding to the person image 50 , and the processor 130 sets the center point of the upper boundary of the candidate region 51 as the head reference point P1 .
[0037] In this embodiment, processor 130 also tracks the movement of a person. For example, processor 130 may employ various known object tracking techniques to track candidate regions associated with the same person, or may track candidate regions by analyzing the correlation between the positions of candidate regions in previous and current image frames to determine the movement of a person, but the present invention is not limited thereto. The detailed steps and implementation methods for tracking the movement of a person can be adequately taught, suggested, and implemented based on common knowledge in the art, and are therefore not further elaborated upon.
[0038] Back to Figure 3 In step S306, the processor 130 is configured to determine whether the person has passed through the entrance or not based on the first detection result and the second detection result. In this embodiment, the processor 130 tracks the image of the person to generate a tracking trajectory of the corresponding candidate area, and the processor 130 determines at least two detection results based on the tracking trajectory. Specifically, the processor 130 tracks the image of the person corresponding to the same person, and generates a first detection result and a second detection result based on the tracking trajectory of the candidate area. Furthermore, the processor 130 determines whether the person has passed through the entrance or not based on the first detection result and the second detection result. Here, the processor 130 may track a full-body image or a partial image corresponding to the same person to generate a trajectory of the head reference point, but the present invention is not limited to this.
[0039] This embodiment takes the head reference point as an example for detailed description. Figure 4 Based on the event detection area 40 set by the present invention, the candidate area of each person image can enter or leave the event detection area 40 in the following six ways:
[0040] Method 1: The head reference point enters the inner area 45 from outside the event detection area 40 through the upper boundary 41 .
[0041] Mode 2: The head reference point leaves the event detection area 40 from the inner area 45 through the upper boundary 41 .
[0042] Mode 3: The head reference point enters the inner area 45 from outside the event detection area 40 through the lower boundary (the left boundary 42 , the right boundary 43 , or the bottom boundary 44 ).
[0043] Mode 4: The head reference point leaves the event detection area 40 from the inner area 45 through the lower boundary (the left boundary 42 , the right boundary 43 , or the bottom boundary 44 ).
[0044] Mode 5: The coordinate position where the head reference point first appears is located in the inner area 45.
[0045] Mode 6: The coordinate position where the head reference point disappears is located in the inner area 45.
[0046] Accordingly, the first detection result generated by the processor 130 based on the tracking trajectory can be one of the six methods of entering the event detection area 40, method 1, method 3, or method 5, and the second detection result can be one of the six methods of leaving the event detection area 40, method 2, method 4, or method 6. Based on the tracking trajectory, the processor 130 can determine whether the head reference point has sequentially occurred in any of the three methods of entering the event detection area 40 and the three methods of leaving the event detection area 40, thereby determining whether the person has passed through the entrance or exit. Based on the three entry methods and three exit methods corresponding to the six methods of entering and exiting the event detection area 40, the following nine events can be generated by combining the three methods:
[0047] It should be noted that in the following events 1 through 9, the disappearance of the head reference point indicates that the processor 130 no longer tracks a person image associated with the same person in the image frame and, therefore, is unable to provide candidate regions and head reference points corresponding to the person image. Furthermore, the initial appearance of the head reference point in the image frame indicates that the processor 130 first detects a person image associated with the person in the image frame and, therefore, generates candidate regions and head reference points corresponding to the person image.
[0048] Event 1: Please refer to Figure 6A , Figure 6AThis is a schematic diagram of an event based on the first entry method according to an embodiment of the present invention. In image frame 601, head reference point P1 and head reference point P2 represent the previous head reference point and current head reference point, respectively, of a person image corresponding to the same person detected by processor 130. In fact, head reference point P1 and head reference point P2 represent the coordinate positions of the head reference point of the person image at two time points, when the corresponding candidate area enters event detection area 40 and when the corresponding candidate area leaves event detection area 40, respectively, during the process of processor 130 tracking the person image. In other words, head reference point P1 represents the previous coordinate position of the person image, and head reference point P2 represents the current coordinate position of the person image. The same applies to events 2 through 9 below. In event 1, the first detection result is that head reference point P1 enters the inner area 45 from outside the event detection area 40 through the upper boundary 41, and the second detection result is that head reference point P2 leaves the event detection area 40 through the upper boundary 41.
[0049] Event 2: Please refer to Figure 6B , Figure 6B This is a schematic diagram illustrating an event based on the first entry method according to an embodiment of the present invention. In image frame 602, the first detection result is that head reference point P1 enters interior region 45 from outside event detection area 40 through upper boundary 41, and the second detection result is that the coordinate position at which head reference point P2 disappears is located within interior region 45.
[0050] Event 3: Please refer to Figure 6C , Figure 6C Figure 6 is a schematic diagram illustrating an event based on the first entry method according to an embodiment of the present invention. In image frame 603, the first detection result is that head reference point P1 enters interior region 45 from outside event detection region 40 through upper boundary 41, and the second detection result is that head reference point P2 leaves event detection region 40 through the lower boundary (left boundary 42, right boundary 43, or bottom boundary 44).
[0051] Event 4: Please refer to Figure 6D , Figure 6D FIG6 is a schematic diagram illustrating an event based on the second entry method according to an embodiment of the present invention. In image frame 604 , the first detection result is that the coordinate position of the head reference point P1 first appears within the inner region 45 , and the second detection result is that the head reference point P2 leaves the event detection region 40 through the upper boundary 41 .
[0052] Event 5: Please refer to Figure 6E , Figure 6E FIG6 is a schematic diagram illustrating an event based on the second entry method according to an embodiment of the present invention. In image frame 605 , the first detection result is that the coordinate position where head reference point P1 first appears is located within inner region 45 , and the second detection result is that the coordinate position where head reference point P2 disappears is located within inner region 45 .
[0053] Event 6: Please refer to Figure 6F , Figure 6F This is a schematic diagram illustrating an event based on the second entry method according to an embodiment of the present invention. In image frame 606, the first detection result is that the coordinate position of head reference point P1 first appears within interior region 45, and the second detection result is that head reference point P2 leaves event detection region 40 through the lower boundary (left boundary 42, right boundary 43, or bottom boundary 44).
[0054] Event 7: Please refer to Figure 6G , Figure 6G This is a schematic diagram illustrating an event based on the third entry method according to an embodiment of the present invention. In image frame 607, the first detection result is that head reference point P1 enters interior area 45 from outside event detection area 40 through the lower boundary (left boundary 42, right boundary 43, or bottom boundary 44), and the second detection result is that head reference point P2 leaves event detection area 40 through upper boundary 41.
[0055] Event 8: Please refer to Figure 6H , Figure 6H This is a schematic diagram illustrating an event based on the third entry method according to an embodiment of the present invention. In image frame 608, the first detection result is that head reference point P1 enters interior region 45 from outside event detection area 40 through the lower boundary (left boundary 42, right boundary 43, or bottom boundary 44), and the second detection result is that the coordinate position at which head reference point P2 disappears is located within interior region 45.
[0056] Event 9: Please refer to Figure 6I , Figure 6I This is a schematic diagram illustrating an event based on the third entry method according to an embodiment of the present invention. In image frame 609, the first detection result is that head reference point P1 enters interior area 45 from outside event detection area 40 through the lower boundary (left boundary 42, right boundary 43, or bottom boundary 44), and the second detection result is that head reference point P2 leaves event detection area 40 through the lower boundary (left boundary 42, right boundary 43, or bottom boundary 44).
[0057] Back to Figure 3 In step S306, the processor 130 determines that the person has not passed through the entrance and exit in the above events 1, 2, 4, 5, and 9. Furthermore, the processor 130 determines that the person has passed through the entrance and exit in the above events 3, 6, 7, and 8.
[0058] Combined with the above events 1, 2, 4, and 5, that is, when the first detection result is that the head reference point P1 enters the internal area 45 from outside the event detection area 40 through the upper boundary 41 or the coordinate position where the head reference point P1 first appears is located in the internal area 45, and the second detection result is that the head reference point P2 leaves the event detection area 40 through the upper boundary 41 or the coordinate position where the head reference point P2 disappears is located in the internal area 45, the processor 130 determines that the person has not passed through the entrance and exit.
[0059] The following uses different embodiments to illustrate the specific application of the result of determining whether a person has passed through an entrance or exit in step S306 in a real-world setting.
[0060] [First embodiment]
[0061] Figure 7 This is a flow chart of a method for determining whether a person enters or exits a field in accordance with an embodiment of the present invention. The method of this embodiment is at least applicable to the above-mentioned system 100, but the present invention is not limited thereto. Figure 1 、 Figure 2 The following describes the detailed steps of the method for determining whether a person enters or exits the venue according to this embodiment, using the various devices and components of system 100. In this embodiment, camera 110 is located within the venue. That is, if a person moves from the same side of the venue as camera 110 through the entrance and exit to a different side (the other side of the entrance and exit) than camera 110, processor 130 will determine that the person has exited the venue. Conversely, if a person moves from a different side of the venue as camera 110 through the entrance and exit to the same side as camera 110, processor 130 will determine that the person has entered the venue.
[0062] Please refer to Figure 7 In step S3061, the processor 130 determines whether the head reference point enters the event detection area through the lower boundary when the candidate area enters the event detection area. For example, if the determination result of step S3061 is the first detection result, reference may be made to the aforementioned entry and exit methods 1, 3, and 5. If the processor 130 determines that the head reference point does not enter the event detection area through the lower boundary when the candidate area enters the event detection area (step S3061, determination is no), then in step S3062, the processor 130 records the candidate area as an entry candidate. Next, in step S3063, the processor 130 determines whether the head reference point leaves the event detection area through the lower boundary when the candidate area leaves the event detection area. For example, if the determination result of step S3063 is the second detection result, reference may be made to the aforementioned entry and exit methods 2, 4, and 6.
[0063] If processor 130 determines that the head reference point has not left the event detection area through the lower boundary when the candidate area leaves the event detection area (step S3063, judgment is negative), then in step S3064, processor 130 determines that the person's movement status is staying outside the event detection area. Specifically, the judgment result corresponding to step S3064 can correspond to the aforementioned events 1, 2, 4, and 5.
[0064] like Figure 6A As shown, if head reference point P1 enters inner region 45 from outside event detection zone 40 through upper boundary 41, the candidate region corresponding to head reference point P1 is recorded as an entry candidate. Subsequently, if head reference point P2 leaves event detection zone 40 from inner region 45 through upper boundary 41, it is determined to be outside the zone.
[0065] like Figure 6B As shown, if head reference point P1 passes through upper boundary 41 from outside event detection area 40 and enters inner area 45, the candidate area corresponding to head reference point P1 is recorded as an entry candidate. Subsequently, if the coordinate position of head reference point P2 disappears within inner area 45, it is determined to be outside the event detection area.
[0066] like Figure 6D As shown, when the coordinate position of the head reference point P1 first appears in the inner area 45, the candidate area corresponding to the head reference point P1 is recorded as the entry candidate state. Afterwards, if the head reference point P2 leaves the event detection area 40 through the upper boundary 41, it is determined to be staying outside the event detection area.
[0067] like Figure 6E As shown, when the coordinate position of the head reference point P1 first appears in the inner area 45, the candidate area corresponding to the head reference point P1 is recorded as the entry candidate state. If the coordinate position of the head reference point P2 disappears in the inner area 45, it is determined that the person is staying outside the venue.
[0068] Back to Figure 7 If processor 130 determines that the head reference point exits the event detection area through the lower boundary when the candidate area exits the event detection area (step S3063, judgment is yes), then in step S3065, processor 130 determines that the person's movement status is entry. Specifically, the judgment result corresponding to step S3065 can correspond to the aforementioned events 3 and 6.
[0069] like Figure 6C As shown, if head reference point P1 enters inner region 45 from outside event detection region 40 through upper boundary 41, the candidate region corresponding to head reference point P1 is recorded as an entry candidate. Subsequently, if head reference point P2 leaves event detection region 40 through a lower boundary (left boundary 42, right boundary 43, or bottom boundary 44), it is considered an entry.
[0070] like Figure 6F As shown, when the coordinate position of head reference point P1 first appears in inner region 45, the candidate region corresponding to head reference point P1 is recorded as an entry candidate. Subsequently, if head reference point P2 leaves event detection region 40 through the lower boundary (left boundary 42, right boundary 43, or bottom boundary 44), it is determined to be an entry.
[0071] Back to Figure 7 If the processor 130 determines that the head reference point entered the event detection area through the lower boundary when the candidate area entered the event detection area (step S3061, judgment is yes), then in step S3066, the processor 130 records the candidate area as an exit candidate. Next, in step S3067, the processor 130 determines whether the head reference point left the event detection area through the lower boundary when the candidate area left the event detection area. For example, if the judgment result in step S3067 is the second detection result, refer to the aforementioned entry and exit methods 2, 4, and 6.
[0072] If processor 130 determines that the head reference point has not left the event detection area through the lower boundary when the candidate area leaves the event detection area (step S3067, judgment is negative), then in step S3068, processor 130 determines that the person's movement status is exit. Specifically, the judgment result corresponding to step S3068 can correspond to the aforementioned events 7 and 8.
[0073] like Figure 6G As shown, if head reference point P1 enters interior region 45 from outside event detection region 40 through the lower boundary (left boundary 42, right boundary 43, or bottom boundary 44), the candidate region corresponding to head reference point P1 is recorded as an exit candidate. Subsequently, if head reference point P2 leaves event detection region 40 through upper boundary 41, it is considered an exit candidate.
[0074] like Figure 6H As shown, if head reference point P1 enters inner region 45 from outside event detection area 40 through the lower boundary (left boundary 42, right boundary 43, or bottom boundary 44), the candidate region corresponding to head reference point P1 is recorded as an exit candidate. Subsequently, if the coordinate position at which head reference point P2 disappears is within inner region 45, it is considered an exit.
[0075] Back to Figure 7 If processor 130 determines that the head reference point exited the event detection area through the lower boundary when the candidate area exited the event detection area (step S3067, judgment is yes), then in step S3069, processor 130 determines that the person's movement status is staying within the event detection area. Specifically, the judgment result corresponding to step S3069 can correspond to the aforementioned event 9.
[0076] like Figure 6I As shown, if head reference point P1 enters interior region 45 from outside event detection region 40 through the lower boundary (left boundary 42, right boundary 43, or bottom boundary 44), the candidate region corresponding to head reference point P1 is recorded as an exit candidate. Subsequently, if head reference point P2 leaves event detection region 40 through the lower boundary (left boundary 42, right boundary 43, or bottom boundary 44), it is determined to be within the region.
[0077] [Second embodiment]
[0078] In this embodiment, the camera 110 is set outside the venue. That is, corresponding to the two sides of the venue connected by the entrance and exit, if a person moves from the same side as the camera 110 through the entrance and exit to the other side of the camera 110 (the other side of the entrance and exit), the processor 130 will determine that the person has entered. Conversely, if a person moves from the other side of the camera 110 through the entrance and exit to the same side as the camera 110, the processor 130 will determine that the person has left. The details of this embodiment can be referred to the detailed description of the aforementioned first embodiment, replacing the entry candidate state with the exit candidate state, the exit candidate state with the entry candidate state, entry with exit, and exit with entry, and replacing staying outside with staying inside, and staying inside with staying outside. No further details will be given here.
[0079] In summary, the system, method, and storage medium provided by the present invention for determining whether a person enters or exits a venue can be configured with event detection zones corresponding to entrances and exits. Based on these zones, the system determines whether a person has passed through an entrance or exit, further determining whether the person's movement status is entry, exit, or lingering. By placing a camera at a high altitude and determining entry and exit based on the head reference point of a person's image, the present invention avoids the situation where the lower body image of a person is obscured in crowded conditions, making it impossible to determine entry and exit based on the position of the foot image. Consequently, the present invention can reduce the false positive rate when detecting entry and exit in crowded situations.
[0080] Furthermore, the present invention sets an upper boundary for taller people and a lower boundary for shorter people. This prevents a taller person's image from accidentally touching the upper boundary, causing a false positive, when standing in front of the entrance or inside the wall (on the same side as the camera). It also prevents a shorter person's image from falling below the lower boundary, causing a false negative, when standing behind the entrance or outside the wall (on a different side from the camera). Furthermore, the present invention sets left and right boundaries to distinguish whether a person is on the same side or different side of the camera before entering the entrance from either side (the corresponding person's image enters the event detection zone), and whether a person is on the same side or different side of the camera when leaving the entrance from either side (the corresponding person's image leaves the event detection zone). This judgment is not affected by the person's height. Therefore, the present invention can reduce the error rate in detecting people entering and leaving the entrance in situations where there are significant differences in people's height.
[0081] Finally, the present invention sets the lower boundary (including the left boundary, the right boundary, and the bottom boundary) as the basis for judging whether a person in the image enters or exits the venue. In fact, the judgment position of the person entering or exiting the venue has been moved to the same side as the camera (the backlight mitigation zone). For example, in the first embodiment where the camera is set up inside the venue, when a person enters the venue, all events in which the head reference point of the tracked person's image disappears in the inner area (the backlight severity zone) are considered to be staying or passing outside the venue ( Figure 6B 、 Figure 6E ), is not listed as an entry / exit event, so it does not affect the accuracy of entry / exit judgment. Only when the head reference point crosses the lower boundary (backlight mitigation area) is it judged that the person enters the venue ( Figure 6C 、 Figure 6F When a person exits the scene, the head reference point of the tracked person's image crosses the lower boundary (backlight mitigation area) and enters the inner area (backlight severity area). Once the head reference point disappears (including tracking interruption), it is determined that the person has exited the scene ( Figure 6H ). Therefore, the present invention can reduce the misjudgment rate of detecting people entering and leaving the venue in the situation of backlight interference.
[0082] Therefore, the system, method, and storage medium provided by this invention for determining the entry and exit of personnel can be integrated with existing surveillance systems, facilitating the integration of applications and services related to personnel attribute analysis. Furthermore, this system can reduce the false positive rate in detecting entry and exit in crowded environments and with significant height differences, thereby improving the accuracy of such detection.
[0083] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A system for judging whether a person enters or exits a venue, characterized in that: include: A camera is installed at a high position to shoot the entrance and exit from a low angle and output a video stream simultaneously; as well as a processor coupled to the camera, wherein the processor is configured to receive the video stream and set an event detection zone corresponding to the entrance, wherein the event detection zone includes an upper boundary, a lower boundary, and an inner area, and the lower boundary includes a left boundary, a right boundary, and a bottom boundary; The processor is configured to detect and track a person image corresponding to a person in the video stream, wherein the person image may be a full-body image or a partial image. The processor is configured to determine whether the person has passed through the entrance or not based on the first detection result and the second detection result. When the first detection result indicates that the coordinate position of the candidate area corresponding to the person image first appears in the inner area, and the second detection result indicates that the candidate area leaves the event detection area through the lower boundary, the processor determines that the person passes through the entrance and exit. When the first detection result indicates that the candidate area passes through the lower boundary from outside the event detection area and enters the inner area, and the second detection result indicates that the coordinate position where the candidate area disappears is located in the inner area, the processor determines that the person passes through the entrance; The entrance and exit includes a left wall and a right wall, the left wall corresponds to a left edge line, the left boundary is to the left of the left edge line, the right wall corresponds to a right edge line, and the right boundary is to the right of the right edge line.
2. The system for determining whether a person enters or exits a field as claimed in claim 1, wherein: The candidate area corresponding to the person image includes a head reference point, and the processor is configured to determine whether the person has passed through the entrance or not according to the head reference point.
3. The system for determining whether a person enters or exits a field as claimed in claim 2, wherein: The processor is configured to perform human detection on the image frame of the video stream to define the candidate area, and The processor is configured to define a top center point of the candidate area as the head reference point.
4. The system for determining whether a person enters or exits a venue as claimed in claim 1, wherein: The processor sets the height position of the upper boundary in the image frame according to a first preset height, and sets the height position of the bottom boundary in the image frame according to a second preset height, wherein the first preset height and the second preset height are determined based on a real-world height range of people.
5. The system for determining whether a person enters or exits a venue as claimed in claim 4, wherein: The distance between the left boundary and the left edge line is a preset head width range, and the distance between the right boundary and the right edge line is the preset head width range.
6. The system for determining whether a person enters or exits a venue as claimed in claim 1, wherein: When the first detection result indicates that the candidate area enters the inner area from outside the event detection area through the upper boundary, and the second detection result indicates that the candidate area leaves the event detection area through the lower boundary, the processor determines that the person passes through the entrance and exit. When the first detection result indicates that the candidate area enters the inner area from outside the event detection area through the lower boundary, and the second detection result indicates that the candidate area leaves the event detection area through the upper boundary, the processor determines that the person passes through the entrance and exit.
7. The system for determining whether a person enters or exits a venue as claimed in claim 1, wherein: When the first detection result indicates that the candidate area enters the inner area from outside the event detection area through the upper boundary or the coordinate position where the candidate area first appears is located in the inner area, and the second detection result indicates that the candidate area leaves the event detection area through the upper boundary or the coordinate position where the candidate area disappears is located in the inner area, the processor determines that the person has not passed through the entrance and exit. When the first detection result indicates that the candidate area enters the inner area from outside the event detection area through the lower boundary, and the second detection result indicates that the candidate area leaves the event detection area through the lower boundary, the processor determines that the person has not passed through the entrance and exit.
8. The system for determining whether a person enters or exits a field as claimed in claim 2, wherein: The entrance and exit connect the inside and outside of the venue, the camera is set in the venue, and the processor further determines whether the person is entering or leaving the venue based on the first detection result and the second detection result. The first detection result indicates that the head reference point does not pass through the lower boundary when the candidate area enters the event detection area, and the second detection result indicates that the head reference point passes through the lower boundary when the candidate area leaves the event detection area, and the processor determines that the person is entering. The first detection result indicates that the head reference point passes through the lower boundary when the candidate area enters the event detection area, and the second detection result indicates that the head reference point does not pass through the lower boundary when the candidate area leaves the event detection area, and the processor determines that the person is leaving.
9. The system for determining whether a person enters or exits a venue as claimed in claim 2, wherein: The entrance and exit connect the inside and outside of the venue, the camera is set outside the venue, and the processor further determines whether the person is entering or leaving the venue based on the first detection result and the second detection result. The first detection result indicates that the head reference point does not pass through the lower boundary when the candidate area enters the event detection area, and the second detection result indicates that the head reference point passes through the lower boundary when the candidate area leaves the event detection area, and the processor determines that the person is exiting. The first detection result indicates that the head reference point passes through the lower boundary when the candidate area enters the event detection area, and the second detection result indicates that the head reference point does not pass through the lower boundary when the candidate area leaves the event detection area, and the processor determines that the person is entering.
10. A method for determining whether a person enters or exits an area, for use in a system comprising a camera and a processor, wherein the camera is positioned at a height to capture the entrance or exit at a low angle and synchronously output a video stream, and the processor receives the video stream, characterized in that: The method comprises: Setting an event detection zone corresponding to the entrance and exit, wherein the event detection zone includes an upper boundary, a lower boundary, and an inner area, and the lower boundary includes a left boundary, a right boundary, and a bottom boundary; Detecting and tracking a person image corresponding to a person in the video stream, where the person image may be a full-body image or a partial image; and determining whether the person has passed through the entrance or exit based on a first detection result and a second detection result, wherein the processor determines that the person has passed through the entrance or exit when the first detection result indicates that the coordinate position of the candidate area corresponding to the image of the person first appears in the inner area, and the second detection result indicates that the candidate area leaves the event detection area through the lower boundary. When the first detection result indicates that the candidate area passes through the lower boundary from outside the event detection area and enters the inner area, and the second detection result indicates that the coordinate position where the candidate area disappears is located in the inner area, the processor determines that the person passes through the entrance; The entrance and exit includes a left wall and a right wall, the left wall corresponds to the left edge line, and the right wall corresponds to the right edge line. The step of setting the event detection area corresponding to the entrance and exit further includes: setting the left boundary to the left of the left edge line; and setting the right boundary to the right of the right edge line.
11. The method for determining whether a person enters or exits a field as claimed in claim 10, wherein: The candidate region corresponding to the person image includes a head reference point. The method further includes: It is determined based on the head reference point whether the person has passed through the entrance or exit.
12. The method for determining whether a person enters or exits a venue according to claim 11, wherein: The method further comprises: Performing human figure detection on the image frames of the video stream to define the candidate area; and The top center point of the candidate region is defined as the head reference point.
13. The method for determining whether a person enters or exits a venue according to claim 10, wherein: The step of setting the event detection zone corresponding to the entrance or exit further includes: The height position of the upper boundary in the image frame is set according to a first preset height, and the height position of the bottom boundary in the image frame is set according to a second preset height, wherein the first preset height and the second preset height are determined according to a height range of people in the real world.
14. The method for determining whether a person enters or exits a field as claimed in claim 10, wherein: The step of setting the event detection zone corresponding to the entrance or exit further includes: Setting the distance between the left boundary and the left edge line as a preset head width range; and The distance between the right boundary and the right edge line is set as the preset head width range.
15. The method for determining whether a person enters or exits a venue according to claim 10, wherein: The method further comprises: When the first detection result indicates that the candidate area enters the inner area from outside the event detection area through the upper boundary, and the second detection result indicates that the candidate area leaves the event detection area through the lower boundary, it is determined that the person passes through the entrance; and When the first detection result indicates that the candidate area enters the inner area from outside the event detection area through the lower boundary, and the second detection result indicates that the candidate area leaves the event detection area through the upper boundary, it is determined that the person passes through the entrance and exit.
16. The method for determining whether a person enters or exits a venue according to claim 10, wherein: The method further comprises: If the first detection result indicates that the candidate area enters the inner area from outside the event detection area through the upper boundary or the coordinate position where the candidate area first appears is located in the inner area, and the second detection result indicates that the candidate area leaves the event detection area through the upper boundary or the coordinate position where the candidate area disappears is located in the inner area, it is determined that the person has not passed through the entrance or exit; as well as When the first detection result indicates that the candidate area enters the inner area from outside the event detection area through the lower boundary, and the second detection result indicates that the candidate area leaves the event detection area through the lower boundary, it is determined that the person has not passed through the entrance and exit.
17. The method for determining whether a person enters or exits a field as claimed in claim 11, wherein the entrance and exit connect the field and the outside of the field, and the camera is set inside the field, characterized in that: The method further comprises: Determine whether the person is entering or exiting based on the first detection result and the second detection result. The first detection result indicates that the head reference point does not pass through the lower boundary when the candidate area enters the event detection area, and the second detection result indicates that the head reference point passes through the lower boundary when the candidate area leaves the event detection area, and the person is determined to be entering. The first detection result indicates that the head reference point passes through the lower boundary when the candidate area enters the event detection area, and the second detection result indicates that the head reference point does not pass through the lower boundary when the candidate area leaves the event detection area, and the person is determined to have left the event.
18. The method for determining whether a person enters or exits a site as claimed in claim 10, wherein the entrance and exit connect the inside of the site and the outside of the site, and the camera is set outside the site, characterized in that: The method further comprises: Determine whether the person is entering or exiting based on the first detection result and the second detection result. The first detection result indicates that the head reference point of the candidate area does not pass through the lower boundary when the candidate area enters the event detection area, and the second detection result indicates that the head reference point passes through the lower boundary when the candidate area leaves the event detection area, and the person is determined to have left the event detection area. The first detection result indicates that the head reference point passes through the lower boundary when the candidate area enters the event detection area, and the second detection result indicates that the head reference point does not pass through the lower boundary when the candidate area leaves the event detection area, and the person is determined to be entering.
19. A non-volatile computer storage medium, characterized in that The non-volatile computer storage medium includes at least one program instruction. When the electronic device loads and executes the at least one program instruction, the method of claim 10 can be performed.
Citation Information
Patent Citations
Apparatus, method and system for monitoring presence of persons in an area
CN104137155A
In-out event detection method and system
CN104899574A
Head detection method and apparatus
CN106845383A