Method and apparatus for identifying crowd gathering, and computer readable storage medium

By using head recognition model and clustering model in the video surveillance system, the clustering of people is automatically identified, which solves the problem of inefficiency in the existing technology and realizes efficient and accurate video surveillance crowd gathering recognition.

WO2025148631A1PCT designated stage expired Publication Date: 2025-07-17CETC BIGDATA RES INST CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/139786
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-08
Filing Date
2024-12-17
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

In the prior art, video surveillance systems require a large number of manpower to monitor crowd gatherings, which is inefficient and difficult to accurately identify crowd gatherings in distant and near areas.

Method used

The head recognition model is used to label the image frames in the video stream data, obtain the coordinates and area of the center point of the head frame, and cluster them through the clustering model to determine whether crowd gathering occurs based on the clustering results.

Benefits of technology

Automatically identifying crowd gatherings through artificial intelligence improves efficiency, reduces the demand for manpower, and improves the accuracy of identifying crowd gatherings in distance and near areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024139786_17072025_PF_FP_ABST
    Figure CN2024139786_17072025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in embodiments of the present application are a method and apparatus for identifying crowd gathering, and a computer readable storage medium, which are used for improving the efficiency. The method of the embodiments of the present application comprises: acquiring video stream data photographed by a target camera; inputting the video stream data into a pre-trained human head recognition model to obtain an image sequence to be detected in which human head bounding boxes are labeled, wherein the human head recognition model is used for labeling human heads in images; acquiring center point coordinates of each human head bounding box in said image sequence; calculating the area of each human head bounding box in said image sequence; inputting the center point coordinates and the area of each human head bounding box into a pre-trained clustering model, and using the clustering model to cluster center points of the human head bounding boxes of the images in said image sequence one by one, so as to obtain a clustering result; and on the basis of the clustering result, determining whether crowd gathering occurs.
Need to check novelty before this filing date? Find Prior Art

Description

Method, device and computer-readable storage medium for identifying crowd gatherings Technical Field

[0001] The embodiments of the present application relate to the field of image processing technology, and in particular to a method, device, and computer-readable storage medium for identifying crowd gatherings. Background Art

[0002] Crowds in cities often present numerous security risks, such as terrorist attacks, stampedes, and violent incidents, which constantly threaten people's lives and property. Video surveillance systems are now widely deployed throughout cities. Using video surveillance data to identify crowds is an effective way to monitor crowd gatherings and take appropriate measures to mitigate potential risks.

[0003] In the prior art, video surveillance images are usually observed manually, which requires a large amount of manpower to monitor whether a crowd has gathered, and the efficiency is very low. Summary of the Invention

[0004] The embodiments of the present application provide a method, device, and computer-readable storage medium for identifying crowd gatherings, which can improve efficiency.

[0005] A first aspect of an embodiment of the present application provides a method for identifying a crowd gathering, comprising:

[0006] Get the video stream data captured by the target camera;

[0007] Inputting the video stream data into a pre-trained head recognition model to obtain a sequence of images to be detected with head frames marked, wherein the head recognition model is used to mark heads in the images;

[0008] Obtaining the center coordinates of the head frame in the image sequence to be detected;

[0009] Calculating the area of ​​each head frame in the image sequence to be detected;

[0010] Inputting the center point coordinates and the areas of the head frames into a pre-trained clustering model, and using the clustering model to cluster the image sequence to be detected one by one about the center points of the head frames to obtain a clustering result;

[0011] Determine whether crowd gathering occurs based on the clustering result.

[0012] Optionally, clustering the image sequence to be detected one by one about the center point of the head frame using the clustering model to obtain the clustering result includes:

[0013] Acquire an image in the sequence of images to be detected as a target image according to a chronological order;

[0014] Obtaining the center point of any head frame in the target image as the target point;

[0015] Determining clustered target points of the target points using the areas of the head frames of the target images;

[0016] Calculate the distance between the target point and each cluster target point respectively;

[0017] Based on the distance between the target point and each cluster target point, clustering is performed using the clustering model to obtain a clustering result of the target point;

[0018] Repeating the steps of acquiring target points to clustering using the clustering model to obtain a clustering result of the target image;

[0019] Repeat the steps of acquiring the target image to clustering using the clustering model to obtain a clustering result of the image sequence to be detected.

[0020] Optionally, before inputting the center point coordinates and the areas of the head frames into a pre-trained clustering model, the method further includes:

[0021] Obtain the first set of images with labeled head frames and clustering results;

[0022] The initialized clustering model is trained using the first image set, and the clustering model after each training is evaluated using the silhouette coefficient. When the training is completed, a trained clustering model is obtained.

[0023] Optionally, determining whether crowd gathering occurs according to the clustering result includes:

[0024] Determine whether there is a class in the clustering result that is greater than or equal to the clustering threshold;

[0025] If so, it is determined that a crowd gathering has occurred.

[0026] Optionally, after determining that a crowd gathering occurs, the method further includes:

[0027] The clustering degree is determined by comparing the clusters that are greater than or equal to the clustering threshold with the clustering degree threshold.

[0028] Optionally, after determining that a crowd gathering occurs, the method further includes:

[0029] Cutting out a crowd gathering segment from the video stream data according to an image in which a crowd gathering occurs in the image sequence to be detected;

[0030] The crowd gathering segment is saved.

[0031] Optionally, before inputting the video stream data into a pre-trained head recognition model, the method further includes:

[0032] Obtain video surveillance data of the target detection area;

[0033] Preprocessing the video surveillance data to obtain a set of labeled head frame images;

[0034] Get the industry's head-labeled image set;

[0035] Mixing the head frame image set with the head mark image set to obtain a second image set;

[0036] The initialized head recognition model is iteratively trained using the second image set, and when the training is completed, a trained head recognition model is obtained.

[0037] A second aspect of an embodiment of the present application provides a device for identifying a crowd gathering, including:

[0038] A first acquisition unit is used to acquire video stream data captured by a target camera;

[0039] a labeling unit, configured to input the video stream data into a pre-trained head recognition model to obtain a sequence of images to be detected with head frames annotated, wherein the head recognition model is used to label heads in images;

[0040] A second acquiring unit is used to acquire the coordinates of the center point of the head frame in the image sequence to be detected;

[0041] a calculation unit, configured to calculate the area of ​​each head frame in the sequence of images to be detected;

[0042] a clustering unit, configured to input the center point coordinates and the areas of the head frames into a pre-trained clustering model, and cluster the image sequence to be detected one by one with respect to the center points of the head frames using the clustering model to obtain a clustering result;

[0043] A determination unit is used to determine whether a crowd gathering occurs based on the clustering result.

[0044] A third aspect of the embodiments of the present application provides a device for identifying a crowd gathering, including:

[0045] processor, memory, input and output units, and buses;

[0046] The processor is connected to the memory, the input and output unit, and the bus;

[0047] A program is stored in the memory, and the processor calls the program to execute the method in the first aspect and any possible implementation of the first aspect.

[0048] A fourth aspect of an embodiment of the present application provides a computer-readable storage medium, on which a program is stored. When the program is executed on a computer, the computer executes the method in the first aspect and any possible implementation of the first aspect.

[0049] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:

[0050] The method provided in the embodiment of the present application uses a head recognition model to mark the heads of image frames in video stream data to obtain a sequence of images to be detected, and then uses a clustering model to perform clustering to obtain clustering results. The clustering results are used to determine whether a crowd has gathered. By using artificial intelligence to detect whether a crowd has gathered in video stream data, it is not necessary to spend a lot of manpower for monitoring, thereby improving efficiency. In addition, the clustering model performs clustering based on the coordinates of the center point and the area of ​​the head frame, which can reduce the situation where people far away and near are mistakenly judged as crowds, which is conducive to improving accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] FIG1 is a flow chart of an embodiment of a method for identifying a crowd gathering according to an embodiment of the present application;

[0052] FIG2 is a flow chart of an embodiment of clustering the center points of the head frames of each image sequence to be detected in accordance with the present application;

[0053] FIG3 is a schematic diagram of a process for training a clustering model according to an embodiment of the present application;

[0054] FIG4 is a flow chart of an embodiment of determining whether a crowd gathering occurs based on clustering results according to an embodiment of the present application;

[0055] FIG5 is a schematic diagram of a process for training a head recognition model according to an embodiment of the present application;

[0056] FIG6 is a schematic structural diagram of an embodiment of a device for identifying crowd gatherings in an embodiment of the present application;

[0057] FIG7 is a schematic structural diagram of another embodiment of the device for identifying crowd gatherings in the embodiment of the present application. DETAILED DESCRIPTION

[0058] The embodiments of the present application provide a method, device, and computer-readable storage medium for identifying crowd gatherings, which can improve efficiency.

[0059] The method of the present application can be applied to a server, a terminal or other device with logic processing capabilities, and the present application does not limit this. For the convenience of description, the following description is based on an example in which the execution subject is a server.

[0060] The embodiments of the present application will be described below with reference to the accompanying drawings.

[0061] Referring to FIG. 1 , an exemplary method for identifying a crowd gathering in an embodiment of the present application includes the following steps:

[0062] 101. Obtain video stream data captured by the target camera;

[0063] The target camera is used to capture images of areas where crowds need to be monitored, such as train stations, bus stations, city squares, and tourist attractions. The target camera captures the image, generates video stream data, and transmits it to a server, allowing the server to access the video stream data. It should be noted that the video stream data captured by the target camera can be real-time video stream data or previously captured and stored video stream data, and this application does not limit this.

[0064] 102. Input the video stream data into a pre-trained head recognition model to obtain a sequence of images to be detected with head frames marked. The head recognition model is used to mark the heads in the images.

[0065] The head recognition model can identify heads in images and annotate them with rectangular boxes. Therefore, the server feeds the video stream data into a pre-trained head recognition model. The server then uses the head recognition model to identify and annotate heads in the video stream data, generating a sequence of images to be detected with the head boxes annotated.

[0066] 103. Obtain the center coordinates of the head frame in the image sequence to be detected;

[0067] The head recognition model generates a coordinate system based on the input image. Therefore, when the head recognition model recognizes and annotates video stream data, the coordinates of the four corners and the center point of the head frame on the image are clear and can be obtained by the server.

[0068] 104. Calculate the area of ​​each head frame in the image sequence to be detected;

[0069] The head recognition model generates a coordinate system based on the input image. Therefore, when the head recognition model recognizes and annotates video stream data, the coordinates of the four corner points and the center point of the head frame on the image are clear. Therefore, for any head frame, the server can calculate the corresponding length and width, and then calculate the area.

[0070] 105. Input the center point coordinates and the areas of each head frame into a pre-trained clustering model, and use the clustering model to cluster the center points of the head frames of each image sequence to be detected to obtain a clustering result.

[0071] After obtaining the center point coordinates and the area of ​​each head frame, the server can input the center point coordinates and area of ​​each head frame of each image in the image sequence to be detected into a pre-trained clustering model. The clustering model is used to cluster the center points of the head frames of each image in the image sequence to be detected to obtain the clustering results. It should be noted that the clustering results include many classes, each class contains some points. It should be noted that not all points are clustered. Some points are isolated and do not form clusters with other points. During clustering, two center points with too large an area difference will not be clustered.

[0072] 106. Determine whether crowd gathering occurs based on the clustering results.

[0073] After obtaining the clustering results, the server can use the clustering results to determine whether a crowd gathering occurs.

[0074] In this embodiment, after acquiring the video stream data captured by the target camera, the server processes the video stream data using a head recognition model to obtain a sequence of images to be detected with marked head frames. The server then obtains the center point coordinates of the head frames in the sequence of images to be detected, calculates the area of ​​each head frame, and inputs the center point coordinates and the area of ​​each head frame into a clustering model. The clustering model is used to cluster the center point coordinates and the area of ​​each head frame, thereby obtaining a clustering result. Finally, the server determines whether a crowd has gathered based on the clustering result. The server detects whether a crowd has gathered in the video stream data using artificial intelligence, eliminating the need for extensive manpower monitoring and thus improving efficiency. Furthermore, the clustering model performs clustering based on the center point coordinates and the area of ​​the head frame, which can reduce the misjudgment of distant and nearby people as a crowd, thereby improving accuracy.

[0075] Referring to FIG. 2 , an embodiment of the present application includes clustering the center points of the head frames of each image sequence to be detected. The clustering includes:

[0076] 201. Acquire an image in the sequence of images to be detected as a target image according to a chronological order;

[0077] 202. Obtain the center point of any head frame in the target image as the target point;

[0078] 203. Determine the clustered target points of the target points using the area of ​​each head frame in the target image;

[0079] 204. Calculate the distance between the target point and each cluster target point respectively;

[0080] 205. Based on the distance between the target point and each cluster target point, clustering is performed using a clustering model to obtain a clustering result of the target point.

[0081] When the server performs clustering of the center points of the head frames of each image sequence to be detected, it first obtains the earliest image in the image sequence to be detected as the target image, then uses the center point of any head frame in the target image as the target point, and determines the cluster target point from all other points in the target image based on the area of ​​the head frame corresponding to the target point. Specifically, the area of ​​the head frame corresponding to the target point is compared with the area of ​​the head frame corresponding to each other point in turn. For the area of ​​the head frame corresponding to any other point, the difference between the area and the area of ​​the head frame corresponding to the target point is first calculated, and then the ratio of the difference to the area of ​​the head frame corresponding to the target point is calculated. Only when the ratio falls within the preset error range can it be considered that the point and the target point are at a similar distance (that is, it can be considered that the actual distances of the people corresponding to the two points from the target camera are similar in reality). At this time, the point is determined to be the cluster target point of the target point. Otherwise, the point is determined not to be the cluster target point of the target point.

[0082] After determining the cluster target points for the target point, the server calculates the distance between the target point and the cluster target points. Specifically, the distance calculation method is not limited and can be used, such as Euclidean distance, Manhattan distance, Minkowski distance, etc. After calculating the distance, it can be used for clustering. Points whose distance from the target point is less than or equal to a preset distance are determined to belong to the same cluster as the target point. The server clusters each point in the target image in turn, then merges similar points to ultimately obtain the clustering results for the target image.

[0083] After obtaining the clustering result of an image, the server repeats the above steps to obtain the clustering result of the image sequence to be detected.

[0084] In this embodiment, the server first uses the area of ​​each head frame in the target image to determine the cluster target points for the target points, and then performs the clustering operation. Because in the video captured by the camera, people near the camera are usually larger than those far away, the labeled head frames are also affected. Before clustering, the server filters out head frames that clearly do not belong to the same group based on the difference in head frame area. This can reduce the influence of people far away when clustering people near the camera, and also reduce the interference caused by people near the camera when clustering people far away, thereby improving the accuracy of the clustering results.

[0085] Referring to FIG3 , an embodiment of training a clustering model in the present application includes:

[0086] 301. Obtain a first image set with labeled head frames and clustering results;

[0087] 302. The initialized clustering model is trained using the first image set, and the clustering model after each training is evaluated using the silhouette coefficient. When the training is completed, a trained clustering model is obtained.

[0088] During the training phase, after the server trains the clustering model and obtains the clustering results, for a specific sample (the sample is the center point of the head frame) in the sample space (i.e., an image), the server can calculate the average distance a between it and other samples in its class, as well as the average distance b between the sample and all samples in the closest class. The silhouette coefficient of the sample is calculated according to the silhouette coefficient calculation formula, and the arithmetic mean of the silhouette coefficients of all samples in the entire sample space is taken as the performance indicator of clustering division. The effect of the clustering model is evaluated in combination with the silhouette coefficient. Among them, the value range of the silhouette coefficient is [-1,1], -1 represents poor clustering effect; 1 represents good clustering effect; 0 represents cluster overlap and no good cluster division. The silhouette coefficient calculation formula is as follows: c = (ba) / max (a, b) Formula (1)

[0089] Here, c represents the silhouette coefficient, a represents the average distance between a sample and other samples in its class, b represents the average distance between a sample and all samples in the closest class, and max(a,b) represents the maximum of a and b. The closer c is to 1, the better the clustering effect, and the closer c is to -1, the worse the clustering effect.

[0090] The server can obtain the first image set with labeled head frames and clustering results, and use the first image set to train the initialized clustering model. At the end of each training, the silhouette coefficient is used to evaluate the effect of the clustering model. When the staff determines that the clustering model can achieve good clustering effect, the training can be ended to obtain a trained clustering model.

[0091] In this embodiment, the clustering model is evaluated by the silhouette coefficient, which can improve the robustness.

[0092] Referring to FIG. 4 , an embodiment of the present application for determining whether a crowd gathering occurs based on clustering results includes:

[0093] 401. Determine whether there is a class in the clustering result that is greater than or equal to the clustering threshold. If so, execute step 402; if not, execute step 403;

[0094] After obtaining the clustering results, the server can use the clustering threshold to determine the clustering of each class in the clustering results. When the number of points in a class is greater than or equal to the clustering threshold, step 402 is executed; when the number of points in all classes is less than the clustering threshold, step 403 is executed. It should be noted that the staff can flexibly set the clustering threshold according to actual conditions.

[0095] 402. Confirmation of a crowd gathering;

[0096] When there is any cluster in the clustering result that is greater than or equal to the clustering threshold, the server determines that crowd clustering occurs.

[0097] 403. Ensure that no crowds gather.

[0098] When there is no cluster in the clustering result that is greater than or equal to the clustering threshold, the server determines that crowd clustering occurs.

[0099] In this embodiment, the server uses the gathering threshold to determine whether a crowd gathering occurs. Different gathering thresholds can be set for different scenarios, thereby improving flexibility.

[0100] Furthermore, after determining that a crowd has gathered, the server can use the gathering degree threshold to subdivide the gathering degree. For different gathering degrees, the staff can take different measures to handle them.

[0101] Furthermore, after determining that a crowd has gathered, the server can extract a crowd gathering segment from the video stream data based on the image of the crowd gathering, and then save the crowd gathering segment, which is conducive to trace management.

[0102] Referring to FIG5 , an embodiment of training a head recognition model in the present application includes:

[0103] 501. Obtain video surveillance data of the target detection area;

[0104] 502. Preprocess the video surveillance data to obtain a set of labeled head frame images;

[0105] 503. Obtain a set of head-labeled images of the industry;

[0106] 504. Mix the head frame image set and the head mark image set to obtain a second image set;

[0107] 505. Perform iterative training on the initialized head recognition model using the second image set. When the training is completed, a trained head recognition model is obtained.

[0108] In this embodiment, the video surveillance data of the target detection area and the industry's head marking image set are added to the second image set, which can improve the recognition rate of the target detection area detection without causing overfitting.

[0109] Referring to FIG6 , an embodiment of a device for identifying a crowd gathering in an embodiment of the present application includes:

[0110] The first acquisition unit 601 is used to acquire video stream data captured by a target camera;

[0111] The labeling unit 602 is used to input the video stream data into a pre-trained head recognition model to obtain a sequence of images to be detected with head frames marked. The head recognition model is used to label the heads in the images.

[0112] The second acquisition unit 603 is used to obtain the center coordinates of the head frame in the image sequence to be detected;

[0113] A calculation unit 604 is used to calculate the area of ​​each head frame in the image sequence to be detected;

[0114] The clustering unit 605 is configured to input the center point coordinates and the areas of the head frames into a pre-trained clustering model, and use the clustering model to cluster the center points of the head frames of each image sequence to be detected to obtain a clustering result.

[0115] The determination unit 606 is used to determine whether a crowd gathering occurs according to the clustering result.

[0116] In this embodiment, after the first acquisition unit 601 acquires the video stream data captured by the target camera, the annotation unit 602 processes the video stream data using a head recognition model to obtain a sequence of images to be detected with annotated head frames. The second acquisition unit 603 then obtains the center point coordinates of the head frames in the sequence of images to be detected. Simultaneously, the calculation unit 604 calculates the area of ​​each head frame and inputs the center point coordinates and the area of ​​each head frame into a clustering model. The clustering unit 605 uses the clustering model to cluster the images based on the center point coordinates and the area of ​​each head frame, thereby obtaining a clustering result. Finally, the determination unit 606 determines whether a crowd gathering has occurred based on the clustering result. Detecting whether a crowd gathering has occurred in video stream data using artificial intelligence eliminates the need for extensive human monitoring and thus improves efficiency. Furthermore, the clustering model performs clustering based on the center point coordinates and the area of ​​the head frames, which can reduce the misclassification of distant and nearby people as crowd gatherings, thereby improving accuracy.

[0117] Optionally, the clustering unit 605 may be specifically configured to:

[0118] Acquire an image in the image sequence to be detected as the target image according to the time sequence;

[0119] Get the center point of any head frame in the target image as the target point;

[0120] The clustering target points of the target points are determined by using the area of ​​each head frame in the target image;

[0121] Calculate the distance between the target point and each cluster target point respectively;

[0122] Based on the distance between the target point and each cluster target point, clustering is performed using the clustering model to obtain the clustering results of the target point;

[0123] Repeat the steps from acquiring the target points to clustering using the clustering model to obtain the clustering result of the target image;

[0124] Repeat the steps from acquiring the target image to clustering using the clustering model to obtain the clustering result of the image sequence to be detected.

[0125] Optionally, the apparatus for identifying a crowd gathering may further include a training unit, which may be used to:

[0126] Obtain the first set of images with labeled head frames and clustering results;

[0127] The initialized clustering model is trained using the first image set, and the clustering model after each training is evaluated using the silhouette coefficient. When the training is completed, a trained clustering model is obtained.

[0128] Optionally, the determining unit 606 may be specifically configured to:

[0129] Determine whether there are classes in the clustering results that are greater than or equal to the clustering threshold;

[0130] If so, it is determined that a crowd gathering has occurred.

[0131] Optionally, the device for identifying a crowd gathering may further include:

[0132] The comparison unit is used to compare the clusters greater than or equal to the clustering threshold with the clustering degree threshold to determine the clustering degree.

[0133] Optionally, the apparatus for identifying a crowd gathering may further include a storage unit, which may be used to:

[0134] Cutting out crowd gathering segments from the video stream data according to the images in the image sequence to be detected where crowd gathering occurs;

[0135] Save the crowd gathering clip.

[0136] Optionally, the apparatus for identifying a crowd gathering may further include a second training unit, which may be used to:

[0137] Obtain video surveillance data of the target detection area;

[0138] Preprocess the video surveillance data to obtain a set of labeled head frame images;

[0139] Get the industry's head-labeled image set;

[0140] Mixing the head frame image set with the head labeled image set to obtain a second image set;

[0141] The initialized head recognition model is iteratively trained using the second image set, and when the training is completed, a trained head recognition model is obtained.

[0142] In this embodiment, the device for identifying crowd gatherings can also execute the steps in the embodiments shown in Figures 1 to 5 above and achieve the same effect, which will not be repeated here.

[0143] Referring to FIG. 7 , another embodiment of the apparatus for identifying a crowd gathering in the embodiment of the present application includes:

[0144] Processor 701, memory 702, input and output unit 703 and bus 704;

[0145] The processor 701 is connected to the memory 702, the input and output unit 703 and the bus 704;

[0146] The memory 702 stores a program, and the processor 701 calls the program to execute the steps in the embodiments shown in FIG. 1 and FIG. 2 .

[0147] In this embodiment, the functions of the processor 701 correspond to the steps in the embodiments shown in Figures 1 to 2 above, and will not be repeated here.

[0148] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0149] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0150] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0151] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0152] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

Claims

1. A method for identifying crowd gathering, characterized in that, Including: Obtain the video stream data captured by the target camera; Input the video stream data into a pre-trained human head recognition model to obtain a sequence of images to be detected with labeled human head frames, where the human head recognition model is used to label the human heads in the images; Obtain the center point coordinates of the human head frames in the sequence of images to be detected; Calculate the areas of each human head frame in the sequence of images to be detected; Input the center point coordinates and the areas of each human head frame into a pre-trained clustering model, and use the clustering model to perform clustering on the sequence of images to be detected one by one with respect to the center points of the human head frames to obtain a clustering result; Determine whether a crowd gathering has occurred according to the clustering result.

2. The method according to claim 1, characterized in that, The performing clustering on the sequence of images to be detected one by one with respect to the center points of the human head frames by using the clustering model to obtain a clustering result includes: Obtain an image in the sequence of images to be detected as a target image according to the chronological order; Obtain the center point of any human head frame in the target image as a target point; Determine the clustering target points of the target point by using the areas of each human head frame in the target image; Calculate the distances between the target point and each clustering target point respectively; Based on the distances between the target point and each clustering target point, perform clustering by using the clustering model to obtain the clustering result of the target point; Repeat the steps from obtaining the target point to performing clustering by using the clustering model to obtain the clustering result of the target image; Repeat the steps from obtaining the target image to performing clustering by using the clustering model to obtain the clustering result of the sequence of images to be detected.

3. The method according to claim 1, wherein Before inputting the center point coordinates and the areas of each human head frame into a pre-trained clustering model, the method further includes: Obtain a first image set with labeled human head frames and clustering results; Train the initialized clustering model by using the first image set, and evaluate the clustering model after each training by using the silhouette coefficient. When the training is completed, obtain the trained clustering model.

4. The method according to claim 1, wherein The determining whether a crowd gathering has occurred according to the clustering result includes: Judge whether there is a class greater than or equal to the gathering threshold in the clustering result; If so, determine that a crowd gathering has occurred.

5. The method according to claim 4, wherein After determining that a crowd gathering has occurred, the method further includes: Compare the class greater than or equal to the gathering threshold with the gathering degree threshold to determine the gathering degree.

6. The method according to claim 4, wherein After determining that a crowd gathering has occurred, the method further includes: Extract the crowd gathering segment from the video stream data according to the images with crowd gathering in the sequence of images to be detected; Save the crowd gathering segment.

7. The method according to any one of claims 1 to 6, characterized in that, Before inputting the video stream data into a pre-trained human head recognition model, the method further includes: Obtain the video surveillance data of the target detection area; Preprocess the video surveillance data to obtain a set of labeled human head frame images; Obtain the industry's human head marked image set; Mix the set of human head frame images with the human head marked image set to obtain a second image set; Iteratively train the initialized human head recognition model by using the second image set. When the training is completed, obtain the trained human head recognition model.

8. A device for identifying crowd gathering, characterized in that, Including: A first acquisition unit, configured to acquire video stream data captured by a target camera; A labeling unit, configured to input the video stream data into a pre-trained human head recognition model to obtain a sequence of to-be-detected images with labeled human head frames, where the human head recognition model is used to label human heads in images; A second acquisition unit, configured to acquire the central point coordinates of the human head frames in the sequence of to-be-detected images; A calculation unit, configured to calculate the areas of the human head frames in the sequence of to-be-detected images; A clustering unit, configured to input the central point coordinates and the areas of the human head frames into a pre-trained clustering model, and use the clustering model to perform clustering on the sequence of to-be-detected images one by one with respect to the central points of the human head frames to obtain a clustering result; A determination unit, configured to determine whether crowd gathering occurs according to the clustering result.

9. A device for identifying crowd gathering, characterized in that, Including: A processor, a memory, an input / output unit, and a bus; The processor is connected to the memory, the input / output unit, and the bus; The memory stores a program, and the processor calls the program to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, where a program is stored on the computer-readable storage medium, and when the program is executed on a computer, the computer is caused to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Crowd density determination method and device, equipment and storage medium

    CN113537172A

  • Method, system and equipment for identifying rapid crowd gathering behavior and medium

    CN114627406A

  • Personnel gathering behavior identification method and electronic equipment

    CN116363597A

  • Method and device for identifying crowd gathering and computer readable storage medium

    CN117994719A

  • Identification method and system for quick gathering behavior of crowd, and device and medium

    WO2023155482A1