A method, system, device and medium for identifying the rapid gathering behavior of a crowd

Through head detection model and cluster analysis, combined with congestion coefficient calculation method, the accuracy problem of traditional methods in identifying rapid crowd aggregation in complex environments is solved, and efficient crowd aggregation recognition in diverse scenarios is achieved.

CN114627406BActive Publication Date: 2025-07-08GRG BANKING EQUIPMENT CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202210147617.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-17
Publication Date
2025-07-08
Estimated Expiration
2042-02-17

AI Technical Summary

Technical Problem

The existing traditional methods of crowd gathering detection are difficult to adapt to diverse environmental conditions, especially in the case of high background noise and complex environment, and it is impossible to accurately identify the rapid crowd gathering behavior.

Method used

The head detection model is used to detect the continuous frame images by head detection, and the head position is represented by rectangular frames. Cluster analysis and congestion coefficient calculation methods are used to identify the crowd gathering area and judge the congestion situation. Combined with the rectangular area growth processing, it can adapt to the needs of multiple scenarios.

Benefits of technology

It improves the accuracy and adaptability of rapid crowd aggregation detection, and can accurately identify rapid crowd aggregation behavior in complex and changing environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114627406B_ABST
    Figure CN114627406B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, system, device and medium for identifying rapid crowd gathering behavior. The identification method includes: acquiring consecutive frame images, and sequentially performing head detection on each frame image based on a head detection model; performing clustering analysis on the gathering areas of the entire frame image according to the head detection results, calculating the number of people's heads in each clustered gathering area, and calculating the congestion coefficient corresponding to each frame image; when there is a target frame image with a congestion coefficient greater than a preset sparse coefficient, calculating and updating the congestion coefficients and the maximum number of people in the gathering area of all frame images within a preset time period starting from the target frame image; determining whether there is a frame image within the preset time period whose congestion coefficient and maximum number of people in the gathering area both exceed their corresponding preset thresholds. If so, output the result of rapid crowd gathering. The present invention solves the problem of difficult feature expression caused by serious lack of human body contours when the crowd is dense, and at the same time improves the adaptability of the identification method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video surveillance, and particularly to a method, system, device, and computer-readable storage medium for identifying rapid crowd gathering behavior. Background Art

[0002] At present, with the improvement of public safety awareness and the rapid development of security technology, video surveillance systems have been gradually applied to the field of safe cities. The traditional monitoring method of manually retrieving and viewing real-time videos can no longer meet the needs of the rapid development of cities. With the continuous improvement of computer performance and the continuous development of computer vision technology, image processing-based methods are increasingly used in intelligent video surveillance.

[0003] Traditional crowd gathering detection methods mainly include the optical flow method (refer to patent document CN107330372B), the inter-frame difference method (refer to patent document CN112232316B), and the mathematical statistics method (refer to patent document CN105117683B); among them, the optical flow method and the inter-frame difference method estimate the crowd gathering situation by extracting foreground features, but the optical flow method takes a long time, and the inter-frame difference method has poor noise adaptability. The mathematical statistics method is mainly designed for specific scenarios, and it is difficult for the above three methods to adapt to diverse environmental conditions. Summary of the Invention

[0004] In order to overcome the deficiencies of the prior art, one of the purposes of the present invention is to provide an identification method for rapid crowd gathering behavior with strong adaptability, which can be applied to multiple scenario requirements.

[0005] Another purpose of the present invention is to provide an identification system for rapid crowd gathering behavior.

[0006] Another purpose of the present invention is to provide an electronic device.

[0007] Another purpose of the present invention is to provide a computer-readable storage medium.

[0008] One of the purposes of the present invention is achieved by adopting the following technical solutions:

[0009] An identification method for rapid crowd gathering behavior includes:

[0010] Obtaining consecutive frame images, and sequentially performing head detection on each frame image based on a head detection model;

[0011] Performing clustering analysis on the gathering areas of the entire frame image according to the head detection results, calculating the number of human heads in each clustered gathering area, and calculating the congestion coefficient corresponding to each frame image;

[0012] When there is a target frame image with the congestion coefficient greater than the preset sparse coefficient, calculate and update the congestion coefficient and the number of people in the largest aggregation area of all frame images starting from the target frame image within a preset time period; determine whether there is a frame image within the preset time period in which both the congestion coefficient and the number of people in the largest aggregation area exceed their corresponding preset thresholds. If so, output the result of rapid crowd aggregation.

[0013] Further, the head detection model is obtained by training a neural network with an image sample set labeled with pedestrian heads; and when the head detection model detects a head in the image, a corresponding rectangular box is generated and displayed at the position of each head.

[0014] Further, the method of cluster analysis is as follows:

[0015] Calculate the distance between each rectangular box corresponding to a head and the rest of the rectangular boxes, and store all the rectangular boxes with a distance less than the first fixed distance in the first neighboring rectangular array;

[0016] Calculate the number of adjacent rectangular boxes of each rectangular box, determine whether the number of adjacent rectangular boxes of each rectangular box is less than the preset value, and delete the rectangular boxes with the number of adjacent rectangular boxes less than the preset value and their corresponding first neighboring rectangular arrays to obtain the preliminary aggregation area;

[0017] Perform rectangular region growing processing on each preliminary aggregation area to obtain the final aggregation area.

[0018] Further, the method of rectangular region growing processing is as follows:

[0019] Traverse each aggregation area, calculate the target rectangular box with the most neighboring rectangular boxes in each aggregation area, obtain the set of rectangular boxes in the area expanded outward with the midpoint of the target rectangular box in the current aggregation area as the center and the first fixed distance as the radius, traverse the set of rectangular boxes to search for rectangular boxes with a distance less than or equal to the second fixed distance and store them in the second neighboring rectangular array; where the second fixed distance is less than the first fixed distance;

[0020] Perform a union operation on the first neighboring rectangular array and the second neighboring rectangular array corresponding to each aggregation area to obtain the region growing result of each aggregation area.

[0021] Further, the first fixed distance = 2.5 * (square of the width value of the rectangular box + square of the height value of the rectangular box), and the second fixed distance = 1.5 * (square of the width value of the rectangular box + square of the height value).

[0022] Further, the calculation method of the congestion coefficient for the aggregation area in the current frame image is as follows:

[0023]

[0024] Among them, len[L(n)] represents the size of the current detected rectangular box array, and len(G p ) represents the total number of people in all aggregation areas, and len(G b ) represents the total number of aggregation areas. β ∈ (0.4, 0.5, 0.6), and the value of β is related to the number of people in the largest aggregation area, corresponding to the number ranges of ([3, 9), [9, 12), [12, +∞)) respectively.

[0025] Furthermore, the calculation method of the congestion coefficient for the non-aggregation areas in the current frame image is as follows:

[0026] Crd = α * len[L(n)];

[0027] Among them, the scene coefficient α satisfies 0.01 ≤ α ≤ 0.04, and len[L(n)] represents the size of the current detected rectangular box array; when Crd exceeds 1.0, the value is taken as 1.0.

[0028] The second object of the present invention is realized by adopting the following technical solution:

[0029] An identification system for rapid crowd aggregation behavior, which executes the identification method for rapid crowd aggregation behavior as described above, includes:

[0030] A head detection module, configured to obtain consecutive frame images and perform head detection on each frame image in sequence based on a head detection model;

[0031] A crowd analysis module, configured to perform clustering analysis on the aggregation areas of the entire frame image according to the head detection results, calculate the number of people's heads in each clustered aggregation area, and calculate the congestion coefficient corresponding to each frame image; when there is a target frame image whose congestion coefficient is greater than a preset sparse coefficient, calculate and update the congestion coefficient and the number of people in the largest aggregation area of all frame images within a preset time period starting from the target frame image;

[0032] A judgment and output module, configured to judge whether there is a frame image within the preset time period whose congestion coefficient and the number of people in the largest aggregation area both exceed their corresponding preset thresholds. If so, output the result of rapid crowd aggregation.

[0033] The third object of the present invention is realized by adopting the following technical solution:

[0034] An electronic device, which includes a processor, a memory, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it realizes the above-mentioned identification method for rapid crowd aggregation behavior.

[0035] The fourth object of the present invention is achieved by the following technical solution:

[0036] A computer-readable storage medium stores a computer program thereon, and when the computer program is executed, the above-mentioned method for identifying the rapid crowd gathering behavior is implemented.

[0037] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0038] The present invention proposes a method for describing crowd characteristics, which uses a head recognition method to solve the problem of difficult feature expression caused by serious lack of human body contours when the crowd is densely distributed; the present invention accurately divides the crowd-dense areas in the image according to the crowd distribution by clustering each aggregation area, improving the accuracy of rapid crowd gathering detection; at the same time, the combination of the congestion coefficient calculation method makes the method of the present invention have stronger scene adaptability, can adapt to multiple scene requirements at the same time, and has high feasibility. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 It is a schematic diagram of the principle of the method for identifying the rapid crowd gathering behavior of the present invention;

[0040] Figure 2 It is a schematic flowchart of the method for identifying the rapid crowd gathering behavior of the present invention;

[0041] Figure 3 It is a schematic diagram of the neural network structure of the head detection model of the present invention;

[0042] Figure 4 It is a schematic diagram of the modules of the system for identifying the rapid crowd gathering behavior of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] Next, in combination with the accompanying drawings and specific embodiments, the present invention will be further described. It should be noted that, on the premise of no conflict, the following-described embodiments or technical features can be arbitrarily combined to form new embodiments.

[0044] Embodiment 1

[0045] This embodiment provides a method for identifying the rapid crowd gathering behavior. This method can be applied to actual scenarios with high background noise and complex and changeable environments, can adapt to multiple scene requirements at the same time, and has stronger adaptability.

[0046] The rapid crowd gathering behavior in this embodiment specifically refers to the situation where the crowd quickly changes from a sparse and non-congested state to a congested state within a specified time, that is, the sudden crowd gathering describes that within a certain period of time, the crowd changes from the initial sparse state, or even no one, to quickly gather at one or several positions, and the congestion degree and the number of people in the gathering area exceed the preset threshold.

[0047] According to the characteristics of the rapid aggregation of the above-mentioned crowd, referring to Figure 1 、 Figure 2 as shown, this embodiment proposes a method for identifying rapid crowd aggregation:

[0048] Step S1: Obtain consecutive frame images, and sequentially perform head detection on each frame image based on the head detection model;

[0049] Step S2: Perform clustering analysis on the aggregation area of the entire frame image according to the head detection results, calculate the number of people's heads in each clustered aggregation area, and calculate the congestion coefficient corresponding to each frame image;

[0050] Step S3: Compare the congestion coefficient with a preset value. When there is a target frame image with the congestion coefficient greater than the preset sparse coefficient, calculate and update the congestion coefficient and the maximum number of people in the aggregation area of all frame images within a preset time period starting from the target frame image; determine whether there is a frame image within the preset time period whose congestion coefficient and the maximum number of people in the aggregation area both exceed their corresponding preset thresholds. If so, output the result of rapid crowd aggregation.

[0051] This embodiment considers that when the crowd is aggregating, it is in a dense state, the human body contour information is severely missing, while the head information is roughly complete. The head feature is the main crowd information description feature in the dense crowd; therefore, this embodiment uses a pedestrian head detection method to describe the general distribution of the crowd in a certain frame image. Before analyzing the current frame image, this embodiment needs to pre-establish a head detection model.

[0052] Specifically, this embodiment pre-obtains an image sample set, which includes a large number of images containing head features. Each sample image is based on the scale of 1920*1080. If the sample image exceeds 1920*1080 pixels, the image sample needs to be cropped and normalized, and then the target with a head width greater than 20 pixels in each sample image is marked. The image samples after the above processing are used as the training samples of the convolutional neural network for learning and training to construct a head detection model.

[0053] Such as Figure 3As shown in the figure, the convolutional neural network of the head detection model in this embodiment includes a total of 19 convolutional layers from conv0 / relu to conv18 / relu and 5 Pooling layers. The regression layers are respectively composed of the outputs after relu of conv6, conv9, conv12, conv14, conv16, and conv18 layers. Each regression layer has 6 PriorBoxes. The size of the input image of the network is 640*640. After training and learning the input sample images, a head detection model containing background and head information can be obtained.

[0054] In this embodiment, after obtaining consecutive frame images through actual scene shooting, each frame is input into the already built head detection model. When the head detection model detects that there is a head in the image, a corresponding rectangular box is generated and displayed at the position of each head.

[0055] In this embodiment, rectangular boxes are used to represent the positions of pedestrian heads. When people gather, there will be a phenomenon that multiple rectangular boxes are closely attached to each other. Therefore, only by considering the relationship between each rectangular box can the crowd distribution be described. Generally, when more than 3 people gather in the same area, there is a greater possibility of congestion. Therefore, this embodiment takes at least 3 people forming an aggregation area as the basic point, and calculates the number of aggregation areas in the entire frame image and the number of heads in each aggregation area.

[0056] The principle of obtaining the aggregation area G(Y) in this embodiment is as follows:

[0057] Generally, it is considered that the target person with the largest number of adjacent people in the crowd is distributed in the densest central area. If, when calculating the aggregation area, the target rectangular box with the largest number of adjacent rectangular boxes can be found, this target rectangular box is closest to the center of an aggregation area. Combining rectangular region growing processing, all the rectangular boxes in this aggregation area can be roughly found.

[0058] The method for calculating the target rectangular box with the largest number of adjacent rectangular boxes in this embodiment is as follows:

[0059] Let an array composed of a group of detected rectangular boxes be L(n), and the current rectangular box be represented as L(k) (k∈[0,n]). Calculate the distance between the current rectangular box L(k) and the other rectangular boxes except L(k). All the rectangular boxes with a distance less than the first fixed distance are stored in the first adjacent rectangular box array. The calculation method can be expressed as:

[0060]

[0061] Among them, |L(k)→L(k+m) represents the distance between the central points of two rectangular frames L(k) and L(k+m), where m ∈ [-k, 0) ∪ (0, n-k); Vec[L(k)] represents the distribution of adjacent rectangular frames corresponding to all rectangular frames.

[0062] Sort Vec[L(k)], and the target rectangular frame with the most adjacent rectangular frames can be found. The target rectangular frame represents the center of the aggregation area.

[0063] The first fixed distance is the basis for judging whether the "adjacent rectangular frame" is satisfied. Specifically: in an array composed of a group of detected rectangular frames, the distance between the central points of two rectangular frames is used as the distance S(x1, x2) between different rectangular frames. When obtaining the number of other rectangular frames adjacent to a certain rectangular frame x1, it is necessary to calculate in combination with the first fixed distance. If the distance between a rectangular frame x1 and another rectangular frame x2 is less than the first fixed distance, the two are in an adjacent relationship; if the distance between the two is greater than the first fixed distance, the two are not in an adjacent relationship. In this embodiment, considering the perspective distortion factor due to the distance from the camera in the image, the first fixed distance is fixed to 2.5 times the sum of the squares of the width and height of the rectangular frame x1, that is, the first fixed distance = (the square of the width value of the rectangular frame + the square of the height value of the rectangular frame) * 2.5, denoted as dis(L). The calculation of adjacent rectangular frames is to calculate all rectangular frames that satisfy the distance from the center point of x1 to be less than 2.5 times the sum of the squares of the width and height of x1.

[0064] During the calculation of adjacent rectangular frames, 2.5 times the sum of the squares of the width and height is used, that is, the distance between the target rectangular frame and other rectangular frames within the aggregation area is kept within the range of the first fixed distance. If there are still other rectangular frames distributed at the edge of the aggregation area of a certain rectangular frame, it should belong to this aggregation area. To further optimize the aggregation area, calculate the number of adjacent rectangular frames of each rectangular frame, judge whether the number of adjacent rectangular frames of each rectangular frame is less than the preset value, and delete the rectangular frames with the number of adjacent rectangular frames less than the preset value and their corresponding first adjacent rectangular arrays to obtain a preliminary aggregation area; in this embodiment, the rectangular frames with the number of adjacent rectangular frames less than 2 and the corresponding first adjacent rectangular arrays are deleted.

[0065] After that, traverse each aggregation area in the set of preliminary aggregation areas, and then perform rectangular area growth processing on each aggregation area to obtain the final aggregation area G(Y). Specifically, the method of the rectangular area growth processing is as follows:

[0066] Step S21: Traverse each aggregation area. After calculating the target rectangle with the most neighboring rectangles in each aggregation area, obtain the set of rectangles within the area expanded outward with the midpoint of the target rectangle in the current aggregation area as the center and a first fixed distance as the radius. For example, assume that one of the aggregation areas is Vec[L(b)], and the target rectangle with the most neighboring rectangle frames is L(b). Calculate the set of rectangles distributed outside the circle with the midpoint of the target rectangle L(b) in Vec[L(b)] as the center and a first fixed distance as the radius, denoted as Vec0[L(b)].

[0067] Step S22: Traverse each element of the rectangle set Vec0[L(b)], and search for rectangle frames in L(n) with a distance less than or equal to a second fixed distance, and store them in the second neighboring rectangle array Vec1[L(b)]. The second fixed distance is less than the first fixed distance. In this embodiment, the second fixed distance is 1.5 times the sum of the squares of the width and height of the rectangle frame, that is, the second fixed distance = (the square of the width value of the rectangle frame + the square of the height value) * 1.5.

[0068] Step S23: Obtain the rectangle frame Vec last [L(b)] formed by the union of the first neighboring rectangle array Vec[L(b)] and the second neighboring rectangle array Vec1[L(b)], which is the result of the rectangle region growth of a certain aggregation area.

[0069] Record the results of all aggregation areas after the rectangle region growth process as the final aggregation area G(Y), and the number of heads in the final aggregation area is Y(X), and the number of heads can be directly reflected by the number of head rectangle frames detected by the head detection model.

[0070] In this example, the calculation of the congestion coefficient is divided into two cases, that is, when there is an aggregation area and when there is no aggregation area. Specifically:

[0071] When the crowd distribution is discrete and there is no aggregation area, the probability of the crowd appearing at each position in the environment is approximately the same, that is, the crowd conforms to a uniform distribution. Based on this feature, this embodiment introduces a scene coefficient α, which represents the ratio of the number of people appearing at a certain position in the scene to the full occupancy of the entire scene. When the scene is relatively open, the value of α is small, and when the scene has a narrow field of view, the value of α is large. After multiple experiments in this embodiment, it is shown that 0.01 ≤ α ≤ 0.04 is the best. Therefore, the calculation method of the congestion coefficient when there is no aggregation area in the current frame image is as follows:

[0072] Crd = α * len[L(n)];

[0073] Among them, the scene coefficient α satisfies 0.01 ≤ α ≤ 0.04, and len[L(n)] represents the size of the current detected rectangle array; when Crd exceeds 1.0, it takes the value of 1.0.

[0074] When there is an obvious crowd gathering area in the scene, then the environment gradually shows congestion. The congested location will become more important compared to the location where the discrete crowd is distributed. Therefore, the calculation method of the congestion coefficient when there is an aggregation area in the current frame image of this embodiment is as follows:

[0075]

[0076] Among them, len[L(n)] represents the size of the current detected rectangle array, len(G p ) represents the total number of people in all aggregation areas, len(G b ) represents the total number of aggregation areas, β ∈ (0.4, 0.5, 0.6), and the value of β is related to the number of people in the largest aggregation area, corresponding to the number ranges ([3, 9), [9, 12), [12, +∞)) respectively.

[0077] After calculating the congestion coefficient in this embodiment, if the congestion coefficient increases significantly within a preset time period and the number of people in the aggregation area exceeds the preset threshold, it can be considered as a rapid crowd gathering behavior. Therefore, through a large number of scene tests in this embodiment, a suitable coefficient value representing crowd sparsity is found. It is more reasonable to take the sparsity coefficient as 0.3, that is, when it is less than this value, it is not included in the time of rapid aggregation. Refer to Figure 2 As shown, when the congestion coefficient is greater than the sparsity coefficient, the timing starts, and the congestion coefficient and the number of people in the largest aggregation area of all frames within the preset time period are calculated and updated; if neither the congestion coefficient nor the number of people in the largest aggregation area exceeds the preset threshold set in advance within the preset time period, then it is considered that there is no rapid crowd gathering during this preset time period, and the head detection can continue for the next frame image to analyze the crowd gathering behavior of the next frame image. If the congestion coefficient and the number of people in the largest aggregation area both exceed the preset threshold in a certain frame within the preset time period, it represents the existence of a rapid crowd gathering behavior. Since the crowd cannot disperse immediately in the next frame image after there is a crowd gathering behavior in a certain frame image, therefore, when judging the rapid crowd gathering behavior, it can also be considered whether the conditions that both the congestion coefficient and the number of people in the largest aggregation area exceed the preset threshold are met in the remaining frames. If so, it can be considered that the crowd rapidly gathers within the time corresponding to the frame image where both the congestion coefficient and the number of people in the largest aggregation area exceed the preset threshold.

[0078] Embodiment 2

[0079] As Figure 4As shown in the figure, this embodiment provides a recognition system for the rapid crowd gathering behavior, which executes the recognition method for the rapid crowd gathering behavior described in Embodiment 1, including:

[0080] A head detection module, configured to obtain consecutive frame images and sequentially perform head detection on each frame image based on a head detection model;

[0081] A crowd analysis module, configured to perform clustering analysis on the gathering areas of the entire frame image according to the head detection results, calculate the number of human heads in each clustered gathering area, and calculate the congestion coefficient corresponding to each frame image; when there is a target frame image whose congestion coefficient is greater than a preset sparse coefficient, calculate and update the congestion coefficient and the maximum number of people in the gathering area of all frame images starting from the target frame image within a preset time period;

[0082] A judgment and output module, configured to judge whether there is a frame image in the preset time period whose congestion coefficient and the maximum number of people in the gathering area both exceed their corresponding preset thresholds. If so, output the result of rapid crowd gathering.

[0083] This embodiment uses the Box-Gather (detection box aggregation) method to process the head detection results, obtain the gathering areas and the number of human heads in each area; and adopts a congestion coefficient calculation method to calculate the congestion degree (0-1) of the entire image using the detected number of heads; within a specified period of time, when the congestion degree exceeds a certain threshold and the number of people in the gathering area exceeds a preset threshold, it is considered that the crowd gathers rapidly. This embodiment can adapt to scenarios with large background noise and complex and changeable environments, and greatly improve the adaptability of crowd gathering behavior recognition.

[0084] This embodiment provides an electronic device, which includes a processor, a memory, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the recognition method for the rapid crowd gathering behavior in Embodiment 1; in addition, this embodiment also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed, it implements the above-mentioned recognition method for the rapid crowd gathering behavior.

[0085] The system, device, and storage medium in this embodiment and the method in the foregoing embodiment are multiple aspects based on the same inventive concept. The implementation process of the method has been described in detail before. Therefore, those skilled in the art can clearly understand the structure and implementation process of the system, device, and storage medium in this embodiment according to the foregoing description. For the sake of brevity of the specification, it will not be repeated here.

[0086] The above embodiments are only preferred embodiments of the present invention, and the scope of protection of the present invention cannot be limited thereby. Any non-substantial changes and substitutions made by those skilled in the art based on the present invention fall within the scope of protection required by the present invention.

Claims

1. A method for identifying the rapid aggregation behavior of a crowd, characterized in that, Including: Obtain consecutive frame images, and sequentially perform head detection on each frame image based on a head detection model; Perform clustering analysis on the aggregation regions of the entire frame image according to the head detection results, calculate the number of people's heads in each clustered aggregation region, and calculate the congestion coefficient corresponding to each frame image; wherein, the method of the clustering analysis is: calculate the distance between each rectangular box corresponding to a head and the remaining rectangular boxes according to the positions of the rectangular boxes corresponding to each head, and store all the rectangular boxes with a distance less than a first fixed distance in a first neighboring rectangular array; calculate the number of adjacent rectangular boxes of each rectangular box, determine whether the number of adjacent rectangular boxes of each rectangular box is less than a preset value, and delete the rectangular boxes with the number of adjacent rectangular boxes less than the preset value and their corresponding first neighboring rectangular arrays to obtain a preliminary aggregation region; perform rectangular region growing processing on each preliminary aggregation region to obtain a final aggregation region; The calculation method for the congestion coefficient of the aggregation region in the current frame image is: Among them, len[L(n)] represents the size of the current detected rectangle array, and len(G p ) represents the total number of people in all gathering areas, and len(G b ) represents the total number of gathering areas. β ∈ (0.4, 0.5, 0.6), and the value of β is related to the number of people in the largest gathering area, corresponding to the number ranges of [3, 9), [9, 12), and [12, +∞) respectively; The calculation method for the congestion coefficient of the non-aggregation region in the current frame image is: Crd = α *len[L(n)]; Among them, the scene coefficient α is 0.01 ≤ α ≤ 0.04, and len[L(n)] represents the size of the current detected rectangle array; when Crd exceeds 1.0, the value is 1.0; When there is a target frame image with a congestion coefficient greater than a preset sparse coefficient, calculate and update the congestion coefficient and the maximum number of people in the aggregation region of all frame images within a preset time period starting from the target frame image; determine whether there is a frame image within the preset time period with both the congestion coefficient and the maximum number of people in the aggregation region exceeding their corresponding preset thresholds, and if so, output the result of rapid crowd aggregation.

2. The method for identifying the rapid crowd gathering behavior according to claim 1, wherein The head detection model is obtained by training a neural network with an image sample set labeled with pedestrian heads; and when the head detection model detects that there is a head in the image, a corresponding rectangular box is generated and displayed at the position of each head.

3. The method for identifying the rapid crowd gathering behavior according to claim 1, wherein The method of the rectangular region growing processing is: traverse each aggregation region, calculate the target rectangular box with the most neighboring rectangular boxes in each aggregation region, obtain the set of rectangular boxes within the region expanded outward with the midpoint of the target rectangular box in the current aggregation region as the center and the first fixed distance as the radius, traverse the set of rectangular boxes to search for rectangular boxes with a distance less than or equal to a second fixed distance and store them in a second neighboring rectangular array; wherein the second fixed distance is less than the first fixed distance; Perform a union operation on the first neighboring rectangular array and the second neighboring rectangular array corresponding to each aggregation region to obtain the region growing result of each aggregation region.

4. The method for identifying the rapid crowd gathering behavior according to claim 3, wherein The first fixed distance = (the square of the width value of the rectangular box + the square of the height value of the rectangular box) * 2.5, and the second fixed distance = (the square of the width value of the rectangular box + the square of the height value) * 1.

5.

5. An identification system for the rapid aggregation behavior of a crowd, characterized in that, Execute the method for identifying rapid crowd aggregation behavior as described in any one of claims 1 to 4, including: a head detection module, configured to obtain consecutive frame images and sequentially perform head detection on each frame image based on a head detection model; A crowd analysis module, which is used to perform clustering analysis on the aggregation areas of the entire frame image according to the head detection results, calculate the number of human heads in each clustered aggregation area, and calculate the congestion coefficient corresponding to each frame image; when there is a target frame image whose congestion coefficient is greater than the preset sparse coefficient, calculate and update the congestion coefficient and the maximum number of people in the aggregation area of all frame images within a preset time period starting from the target frame image. A judgment output module, which is used to judge whether there is a frame image within the preset time period whose congestion coefficient and the maximum number of people in the aggregation area both exceed their corresponding preset thresholds. If so, output the result of rapid crowd aggregation.

6. An electronic device, characterized in that, It includes a processor, a memory, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the method for identifying the rapid crowd aggregation behavior according to any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the computer program is executed, it implements the method for identifying the rapid crowd aggregation behavior according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • A method for detecting and warning of dense crowds in public places

    CN105117683B

  • An analytical method for a video-based crowd density and abnormal behavior detection system.

    CN107330372B

  • Crowd gathering detection methods, devices, electronic equipment and storage media

    CN112232316B

  • Extraction method for dynamic crowd gathering characteristics

    CN103839065A

  • Crowd aggregation detection method based on human head detection

    CN112560807A