Personnel gathering identification method and electronic equipment
Through the improved YOLOv7 object detection model and height distance adjustment interchange ratio (HDMIoU) clustering algorithm, the limitations of personnel aggregation behavior detection in the existing technology are solved, and the identification and alarm of the global image area is realized, and the powerful scene generalization ability is achieved.
Patent Information
- Application Number
- CN202311499836.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-13
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art has limitations in the detection of personnel gathering behavior. The target detection method can only detect a single area, and cannot perform global image detection and alarm, and the target clustering method requires specific front-end devices to collect data.
The improved YOLOv7 object detection model is adopted, and targeted model enhancement is carried out by adding Biformer module and small object detection layer, combined with data enhancement methods such as Mosaic and MixUp. Then, the clustering calculation of the detection box is performed using the height distance adjustment intersect ratio (HDMIoU) and clustering threshold, and then the recognition of personnel gathering behavior is performed.
It realizes the recognition of the gathering behavior of the global image area, has strong scene generalization ability, and can adaptively solve the projection transformation problem of the near, large and small far, avoiding detection defects limited to a single area.
Smart Images

Figure CN119992438A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image recognition, and in particular to a method and electronic equipment for recognizing a gathering of people. Background Art
[0002] Crowds of people are a precursor to major safety accidents, so crowd gathering behavior recognition is a very important and commonly used technology in the security field. In public safety management, public video surveillance is used to obtain images or videos of designated points, which are automatically processed by computers to analyze the number of people and determine whether crowds have occurred. At the same time, alarm screenshots and videos are saved to the database to form reports and pushed to relevant managers, which can greatly improve the management and control efficiency of the monitored area and avoid accidents.
[0003] With the development of technology, currently known public patents for pedestrian gathering behavior detection are mainly divided into target clustering-based and target detection-based methods.
[0004] The crowd gathering detection algorithm based on target clustering is to perform structured calibration on the coordinate position of the personnel, and traverse the personnel information to divide the gathering circle for clustering judgment, and then determine whether the personnel gathering phenomenon occurs. For example, the invention patent application CN 116595400 A obtains the longitude and latitude data of the target node from the data platform, and then finds the non-chain gathering group by drawing circles to cluster, and then determines the gathering behavior of the personnel. This patent often requires specific front-end equipment to collect data before the back-end algorithm can be started. The crowd gathering judgment method based on target detection performs human target detection by deep feature extraction or traditional image feature extraction, and counts the number of people in the pre-set area in the image, and judges whether the crowd gathering behavior occurs based on this number. For example, the invention patent application CN 114529860 A counts the density of people by thermal energy images, and then determines the gathering of people. The method based on target detection is fast and simple in logic, but it can often only detect a single area, and cannot perform global image detection and alarm, which has great limitations in actual use.
[0005] Glossary:
[0006] Aggregation and direct access: The spatial position relationship between two detection frames is defined. If frame A is the core frame and the HDMIoU between detection frame B and A is less than the minimum aggregation ratio MinHDMIoU, frame A to frame B is considered to be aggregated and directly reachable.
[0007] Attention Module Biformer: A two-layer routing attention module that achieves efficient allocation of computation in a dynamic, query-aware manner to improve the accuracy of deep learning networks.
[0008] Backbone: Backbone network is a network used for feature extraction. It represents a part of the network and is generally used by the front end to extract image information and generate feature maps for use by subsequent networks.
[0009] SPPCSPC: A spatial pyramid pooling layer that can convert multi-scale feature maps of any size into fixed-size feature vectors. Summary of the invention
[0010] In order to solve the above technical problems, the present invention proposes a method and electronic device for identifying a crowd. The purpose of the present invention is achieved through the following technical solutions:
[0011] A method for identifying a gathering of people comprises the following steps:
[0012] S1. Train to obtain a trained target detection model;
[0013] S2. Input the monitoring image into the trained target detection model, detect the head detection frame in the monitoring image, and enlarge the head detection frame to a half-body detection frame that includes the upper body of the detected person:
[0014] S3. Calculate the height distance between the half-body detection frames and adjust the intersection-over-union ratio HDMIoU:
[0015]
[0016] ρ and c are the introduced distance information, ρ(box1, box2) represents the distance between the center points of the two half-body detection boxes, and c represents the diagonal distance of the smallest box that can include the two half-body detection boxes; box1 and box2 represent the two half-body detection boxes respectively; HIoU represents the height intersection over union, and IoU represents the intersection over union;
[0017] S4. Clustering is performed by adjusting the intersection-over-union ratio HDMIoU based on the height distance between the half-body detection frames. If the number of samples in the base cluster obtained by clustering is greater than the preset minimum number of clusters MinClusters, it is determined that people are gathering and an alarm is issued. Otherwise, monitoring continues.
[0018] For further improvement, the calculation method of the intersection over union (IoU) is as follows:
[0019]
[0020] S box1 and S box2 Represent the areas of the two half-body detection boxes respectively.
[0021] For further improvement, the calculation method of the high intersection-over-union ratio HIoU is as follows:
[0022]
[0023] HIoU stands for Height Intersection Over Union, and IoU stands for Intersection Over Union.
[0024] and represents the ordinate of the upper left corner of the first half-body detection frame, Indicates the ordinate of the lower right corner of the first half-body detection box; Indicates the ordinate of the upper left corner of the second half-body detection frame. Indicates the vertical coordinate of the lower right corner of the second half-body detection box; min() means taking the minimum value in the brackets, and max() means taking the maximum value in the brackets.
[0025] In a further improvement, the target detection model is a YOLO target detection model; the YOLO target detection model includes YOLOv5, YOLOv6 and YOLOv7.
[0026] As a further improvement, the YOLO target detection model is a YOLOv7 target detection model, and a dynamic sparse attention module Biformer is added after the SPPCSPC at the Backbone end of the YOLOv7 target detection model; the YOLOv7 target detection model also adds a small target detection layer with a feature map size of 160×160.
[0027] As a further improvement, in step S1, when training the target detection model, data enhancement is performed on the training data set, and the data enhancement methods include Mosaic, MixUp and Random affine.
[0028] As a further improvement, in step S2, the half-body detection frame is obtained by magnifying the head detection frame by two times in horizontal direction and three times in vertical direction.
[0029] As a further improvement, the value of the minimum number of clusters MinClusters is not less than 3.
[0030] As a further improvement, the monitoring image is obtained by sampling image frames at intervals in a decoded offline or online monitoring video.
[0031] An electronic device comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the above method.
[0032] The beneficial effects of the present invention are:
[0033] First, the improved YOLOv7 model with added Biformer module and small target detection layer is used as the head detection algorithm, and targeted model enhancement is performed based on a series of data enhancement methods such as Mosaic and MixUp to enhance the robustness of the detection model. After the detected head frame information is scaled proportionally, the height distance adjustment intersection over union (HDMIoU) value between each other is traversed, and then the height distance adjustment intersection over union (HDMIoU) value and clustering threshold are used to perform clustering calculation of the detection frame, and then the gathering behavior of people is identified. The height distance adjustment intersection over union (HDMIoU) value is calculated by the area, height, and distance ratio between the detection frames, and can well adapt to the projection transformation of near large and far small. Therefore, the present invention can identify the gathering behavior of people in the entire area of the image only through ordinary surveillance video, and has a strong scene generalization ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] The present invention is further described with reference to the accompanying drawings, but the contents in the accompanying drawings do not constitute any limitation to the present invention.
[0035] Figure 1 It is a schematic diagram of the process of the present invention;
[0036] Figure 2 Illustration of the head detection frame Figure 1
[0037] Figure 3 Illustration of the head detection frame Figure 2 ;
[0038] Figure 4 This is an illustration of proportional scaling of the detection frame. Figure 1 ;
[0039] Figure 5 This is an illustration of proportional scaling of the detection frame. Figure 2 ;
[0040] Figure 6 It is a schematic diagram of high intersection-to-union ratio;
[0041] Figure 7 This is a schematic diagram of distance intersection ratio;
[0042] Figure 8 For warning Figure 1 ;
[0043] Fig. 9 For warning Figure 2 ;
[0044] Fig.10 For warning Figure 3 . DETAILED DESCRIPTION
[0045] In order to make the purpose, technical solutions and advantages of the invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and examples.
[0046] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.
[0047] In the description of the present application, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate positions or positional relationships based on the positions or positional relationships shown in the accompanying drawings, which are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present application. In addition, the terms "first" and "second" are used only for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more features. In the description of the present application, the meaning of "multiple" is two or more, unless otherwise clearly and specifically defined.
[0048] In this application, the word "exemplary" is used to mean "serving as an example, illustration, or description". Any embodiment described in this application as "exemplary" is not necessarily to be construed as being preferred or advantageous over other embodiments. The following description is given to enable any technician in the field to implement and use the present application. In the following description, details are listed for the purpose of explanation. It should be understood that a person of ordinary skill in the art can recognize that the present application can be implemented without using these specific details. In other instances, well-known structures and processes will not be elaborated in detail to avoid obscuring the description of the present application with unnecessary details. Therefore, the present application is not intended to be limited to the embodiments shown, but is consistent with the widest scope consistent with the principles and features disclosed in the present application.
[0049] like Figure 1 The present invention proposes a method for identifying crowd gathering behavior based on rectangular detection frame clustering:
[0050] (1) Optimization of target detection model for crowd gathering scenes
[0051] The present invention conducts targeted data enhancement and model optimization on the YOLOv7 detection model for passenger flow statistics scenarios.
[0052] Yolov7 is an algorithm model in the field of target detection. It is based on a deep learning method. By analyzing and processing images, it can recognize and locate targets in images. Yolov7 has high accuracy and efficiency in target detection tasks, so it has been widely used. The core idea of Yolov7 is to transform the target detection problem into a regression problem. It divides the image into multiple grids, each of which is responsible for detecting the target. Each grid predicts the category, location and confidence of the target. Yolov7 can detect multiple targets in one forward propagation by jointly predicting the target's location information with the features of the image. This gives Yolov7 a clear advantage in speed. In Yolov7, the target detection process can be divided into three stages: feature extraction, feature fusion and target prediction.
[0053] On the data side, we use a combination of data enhancement methods such as Mosaic, MixUp, and Random affine to improve model accuracy.
[0054] Mosaic is a new method proposed in YOLOV4, which is suitable for target detection. The main idea is to stitch four pictures into one picture as a training sample. Since Mosaic is used for target detection, the coordinates of the target box must also be changed accordingly when stitching. The advantage of this is that the background of the picture is enriched, and the stitching of four pictures together increases the batch_size in disguise.
[0055] Mixup is an algorithm used in computer vision to perform mixed-class enhancement on images. It fuses samples and labels in the same way to obtain a new training sample. It can mix images between different classes to expand the training data set.
[0056] RandomAffine is an affine transformation, which is achieved through the combination of a series of atomic transformations, including translation, scaling, rotation, and shear.
[0057] On the model side, structural optimization is made in the Backbone and Head parts of the YOLOv7 detection model to improve algorithm performance. The present invention adds a dynamic sparse attention mechanism Biformer after the SPPCSPC at the end of the Backbone to help the network capture long-distance contextual dependencies and improve the network's feature extraction capabilities. The Biformer module achieves flexible computing power allocation through double-layer routing, allowing each Query in the attention mechanism to process a small portion of the most semantically relevant Key-Value pairs. Specifically, for the feature Figure X ∈α H*W*C , divided into S*S different regions, each of which contains feature vectors, that is Through linear mapping, we can get Q = X r W q , K = X r W k 、V=X r W v , where W q , W k , W v are the projection weights of Query, Key, and Value respectively. By constructing a directed graph, we find the top k most relevant region indexes of a region, i.e., I r =topkIndex(Q r (K r ) T ), where Q r and K r To calculate the mean of Query and Key in the current region and compare the region levels, the attention operation can be expressed as O = Attention(Q,K g ,V g )+LCE(V), where LCE(V) is the local context enhancement term, K g =gather(K,I r ), V g =gather(V,I r ) is the tensor value after Key and Value are aggregated.
[0058] The original YOLOv7 model has only three detection layers in the Head part, corresponding to three sets of initialized Anchor values. When the input image size is 640×640, the original detection layer feature map size is 20×20, 40×40, and 80×80. Mapped to a 1920×1080 size image, the smallest width and height of the target size that can be detected are 24 pixels and 13.5 pixels respectively. If the target is smaller than this size, the network is difficult to recognize. Therefore, adding a small target detection layer with a feature map size of 160×160, the minimum detection target will become 12 pixels and 6.75 pixels, which increases the target detection range and is more suitable for head detection in video surveillance.
[0059] (2) Detection frame scaling
[0060] The patent of the present invention uses an improved target detection model based on YOLOv7 for target detection. In the task of identifying crowds, the camera is installed at a high point and often encounters crowded situations. If upper body detection or full body detection is performed, occlusion is prone to occur, resulting in missed detection. Therefore, the target of the detection model in this patent is the head. The head detection model obtained by the data processing and training method in 1) can obtain the rectangular box selection coordinates, center point coordinates and confidence of the head in the image frame. In order to obtain a height distance adjusted intersection-over-union (HDMIoU) that is more suitable for clustering calculations, the head detection frame is scaled proportionally according to the human body proportions to half the size of the body. The scaling ratio is 1:2 horizontally and 1:3 vertically. Figure 2 shown.
[0061] (3) Height distance adjustment cross-over ratio HDMIoU
[0062] The intersection over union (IoU) can effectively provide strong clues between detection frames. For example, the height and distance information of the frames can also provide useful weak clues, which helps to make up for the strong recognition ability of clues about the spatial position relationship between detection frames. For example, the height of the detection frame reflects the depth information to a certain extent, which mainly depends on the distance between the target and the camera. Since the surveillance camera has a certain pitch angle, the detection frames between targets at a certain distance may have intersections. Therefore, the height information is introduced as a penalty term, making the intersection over union calculation information a more effective clue. On this basis, the distance information between the frames is introduced, and the spatial information between each other can be extracted even when the detection frames do not intersect. Based on this, this patent proposes a height distance adjustment intersection over union (HDMIoU), introduces the height and distance information of the detection frame, and calculates and obtains the spatial position clues between the detection frames. The ratio calculation method is adopted, so that the HDMIoU value range is between -1 and 1, which has scale invariance, and can thus well adaptively solve the projection transformation problem caused by the near large and far small images.
[0063] The traditional intersection over union (IoU) can provide strong clues for box quality inspection and is calculated as follows: Where S box1 , S box2 Represents the area of each of the two detection boxes. The high intersection-over-union ratio HIoU can be used as an effective weak clue and is calculated as follows: Figure 3 As shown, the formula is The two boxes are defined as (x1, y1) represents the upper left corner, and (x2, y2) represents the lower right corner. Based on the height intersection-and-union ratio, distance information is further introduced, and the height distance adjustment intersection-and-union ratio calculation formula is: Among them, ρ and c are the introduced distance information, which represent the distance between the center points of the two boxes and the diagonal distance of the smallest box that can include the two boxes, respectively, as follows Figure 4 shown.
[0064] (4) Rectangular clustering algorithm
[0065] The invention adjusts the intersection-and-union ratio according to the height distance and uses a rectangular clustering algorithm to detect crowd gathering. The definition of clustering and the specific algorithm flow are as follows.
[0066] First, traverse the half-body detection frames, calculate the height distance between each other and adjust the intersection and union ratio HDMIoU, and then perform clustering calculation of the target based on this. When there are n detection frames near a detection frame whose distance intersection and union ratio is less than the minimum aggregation ratio MinHDMIoU, and n is greater than the minimum number of aggregation frames MinRecs, it is determined to be clustered, and such detection frames are called core frames. When a detection frame does not belong to the core frame, but there is a core frame near it whose distance intersection and union ratio is less than the minimum aggregation ratio MinHDMIoU, such a frame is identified as a bounding box. When a detection frame does not belong to a core frame or a bounding frame, it is identified as a noise frame.
[0067] If box A is a core box, and the HDMIoU between detection box B and A is less than the minimum clustering ratio MinHDMIoU, it is determined that box A and box B are directly clustered. If there are core boxes A2, ...., An, and core box A is directly clustered to core box A2, ...., core box A(n-1) is directly clustered to core box An, and core box An is directly clustered to any box B, then it is determined that core box A is clustered to any box B. If there is a core box C, and its relationship with any box A and B is directly clustered, then it is determined that box A and box B are connected to each other. If the two detection boxes are not in a clustered connection relationship, the two boxes are determined to be non-clustered connected.
[0068] The definition of cluster based on rectangular detection box clustering algorithm is: the set of rectangular box samples connected by the maximum cluster derived from the cluster reachable relationship, the samples in the set are clustered into a base cluster, if the number of samples in the base cluster is greater than or equal to the minimum number of clusters MinClusters, then this base cluster is considered to be a cluster. The rectangular clustering algorithm proposed in this patent is mainly limited by the input threshold minimum clustering ratio MinHDMIoU, the minimum number of clustering boxes MinRecs, and the minimum number of clusters MinClusters. The specific algorithm flow is as follows:
[0069] Step 1: Input the detection box dataset InputSet, the minimum aggregation ratio MinHDMIoU, the minimum number of clustered boxes MinRecs, and the minimum number of clusters MinClusters.
[0070] Step 2: If there is data in InputSet, randomly select a rectangular box R; if it is empty, jump to step 7.
[0071] Step 3: Traverse and compare the HDMIoU of the rectangular box R with all the rectangular boxes in InputSet and DealSet, and determine whether it is a core box based on the thresholds MinHDMIoU and MinRecs.
[0072] Step 4: If the rectangular box R is a core box, take all the data object boxes that are directly connected to the rectangular box R to form a base cluster, and go to step 5; if the rectangular box R is not a core box, move the rectangular box R from InputSet to DealSet and return to step 2.
[0073] Step 5: Traverse all unprocessed data rectangles in the base cluster, find the data object frame directly reached by the cluster in InpuSet and DealSet, and take out the data and include it in this base cluster.
[0074] Step 6: After step 5, if the number of object frames in this base cluster increases, repeat step 5; if the number of object frames in this base cluster does not increase, return to step 2.
[0075] Step 7: Traverse the base clusters obtained in the above steps. If the number of samples in the base cluster is greater than or equal to MinClusters, a cluster is formed and the cluster is used as the algorithm output.
[0076] (5) Parameter configuration and alarm
[0077] The present invention is mainly used in various scenarios where people may gather in cities. Users need to set up online monitoring video streams or offline videos, configure video frame extraction frequency, define detection areas, and set parameter thresholds in the algorithm. The parameter thresholds include the minimum aggregation ratio MinHDMIoU, the minimum number of aggregation frames MinRecs, and the minimum number of clusters MinClusters. The sensitivity of the algorithm can be controlled by setting MinDIoU and MinRecs, and MinClusters can be set to control the minimum number of people gathering. It should be noted that the value of the minimum number of aggregation frames MinRecs should be less than or equal to the value of the minimum number of clusters MinClusters.
[0078] According to the result of rectangular clustering in (4), it is judged whether a gathering of people occurs. If the detection information in the image is clustered, it is judged that a gathering of people occurs and an alarm is generated. When the parameter thresholds of minimum aggregation ratio MinDIoU = -0.25, minimum number of clustered frames MinRecs = 3, and minimum number of clusters MinClusters = 5, that is, the minimum number of people for gathering alarm is 5, the alarm picture is as follows: Figure 5 shown.
[0079] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for identifying a gathering of people, characterized by: The steps include: S1. Train to obtain a trained target detection model; S2. Input the monitoring image into the trained target detection model, detect the head detection frame in the monitoring image, and enlarge the head detection frame to a half-body detection frame that includes the upper body of the detected person: S3. Calculate the height distance between the half-body detection frames and adjust the intersection-over-union ratio HDMIoU: ρ and c are the introduced distance information, ρ(box1, box2) represents the distance between the center points of the two half-body detection boxes, and c represents the diagonal distance of the smallest box that can include the two half-body detection boxes; box1 and box2 represent the two half-body detection boxes respectively; HIoU represents the height intersection over union, and IoU represents the intersection over union; S4. Clustering is performed by adjusting the intersection-over-union ratio HDMIoU based on the height distance between the half-body detection frames. If the number of samples in the base cluster obtained by clustering is greater than the preset minimum number of clusters MinClusters, it is determined that people are gathering and an alarm is issued. Otherwise, monitoring continues.
2. The method for identifying a gathering of people according to claim 1, wherein: The calculation method of the intersection over union (IoU) is as follows: S box1 and S box2 Respectively represent the areas of the two half-body detection boxes.
3. The method for identifying a gathering of people according to claim 1, wherein: The calculation method of the high intersection-over-union ratio HIoU is as follows: HIoU means Height Intersection over Union, IoU means Intersection over Union; and represents the ordinate of the upper left corner of the first half-body detection frame, Indicates the ordinate of the lower right corner of the first half-body detection box; Indicates the ordinate of the upper left corner of the second half-body detection frame. Indicates the ordinate of the lower right corner of the second half-body detection box; min() means taking the minimum value of the values in the brackets, and max() means taking the maximum value of the values in the brackets.
4. The method for identifying a gathering of people according to claim 1, wherein: The target detection model is a YOLO target detection model; the YOLO target detection model includes YOLOv5, YOLOv6 and YOLOv7.
5. The method for identifying a gathering of people according to claim 4, characterized in that: The YOLO target detection model is a YOLOv7 target detection model, and a dynamic sparse attention module Biformer is added after the SPPCSPC at the Backbone end of the YOLOv7 target detection model; the YOLOv7 target detection model also adds a small target detection layer with a feature map size of 160×160.
6. The method for identifying a gathering of people according to claim 1, wherein: In step S1, when training the target detection model, data enhancement is performed on the training data set, and the data enhancement methods include Mosaic, MixUp and Randomaffine.
7. The method for identifying a gathering of people according to claim 1, wherein: In step S2, the half-body detection frame is obtained by magnifying the head detection frame by two times in horizontal direction and three times in vertical direction.
8. The method for identifying a gathering of people according to claim 1, wherein: The value of the minimum number of clusters MinClusters is not less than 3.
9. The method for identifying a gathering of people according to claim 1, wherein: The monitoring image is obtained by sampling image frames at intervals in a decoded offline or online monitoring video.
10. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Personnel gathering situation analysis method and analysis system
CN114529860A
Personnel aggregation discovery algorithm
CN116595400A