Crowd gathering identification method and device, electronic equipment and storage medium

By performing identification processing, expansion processing and fusion processing on the recognized image, the problem of low accuracy in crowd gathering recognition in the prior art is solved, and higher recognition accuracy and wider applicable scenarios are achieved.

CN119992444APending Publication Date: 2025-05-13SHENZHEN INTELLIFUSION TECHNOLOGIES CO LTD +2
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411979485.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing crowd gathering recognition method has low accuracy in crowd gathering recognition, is susceptible to environmental factors, and is limited in applicable scenarios.

Method used

By identifying the image to be recognized, the number of people and the identification box are obtained, and the expansion process is performed when the number of people is greater than the threshold, the identification error caused by posture or occlusion is reduced, and a more accurate expansion box is obtained. The spatial relationship between the personnel is identified through the fusion process to determine whether there is crowd gathering in the target area.

Benefits of technology

It improves the accuracy of crowd gathering identification, reduces the influence of environmental factors, and expands applicable scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992444A_ABST
    Figure CN119992444A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a crowd gathering recognition method. The method comprises the following steps: acquiring a to-be-recognized image of a target area; performing identification processing based on the to-be-identified image to obtain the number of personnel in the to-be-identified image and a personnel identification frame; when the personnel number is greater than a preset first personnel number threshold value, performing first expansion processing based on the personnel identification frames to obtain a personnel expansion frame corresponding to each personnel identification frame; performing fusion processing based on the personnel expansion frame to obtain at least one fusion region; and based on the fusion region, determining whether crowd gathering exists in the target region. The first expansion processing can reduce identification errors caused by personnel postures or partial shielding, a more accurate personnel expansion frame is obtained, fusion processing is carried out based on the personnel expansion frame, the spatial relation between personnel can be effectively identified, a fusion area is obtained, and the fusion efficiency is improved. Therefore, whether crowd gathering exists in the target area can be accurately judged according to the fusion area, and the accuracy of crowd gathering recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and in particular to a method, device, electronic device and storage medium for identifying crowd gatherings. Background Art

[0002] With the continuous development of information technology, many related methods for public safety have been derived, for example, crowd gathering methods for avoiding stampedes and other incidents. Traditional crowd gathering recognition usually uses video frame extraction as the target image detection area, detects the number of heads in the area, and obtains the number of heads in the detection area in the target image; thereby judging whether a crowd gathers in the area based on the number of heads. However, the above method only makes judgments based on the number of heads, and the judgment method is relatively simple, easily affected by environmental factors, and is limited in applicable scenarios, which leads to a decrease in the accuracy of crowd gathering recognition. Therefore, how to provide a crowd gathering recognition method that can improve the accuracy of crowd gathering recognition has become an urgent problem to be solved. Summary of the invention

[0003] The embodiment of the present invention provides a crowd gathering recognition method, which aims to solve the problem of low crowd gathering recognition accuracy of existing crowd gathering recognition methods. By performing recognition processing on the image to be recognized, the number of people and the person recognition frame in the image to be recognized can be obtained. When the number of people is greater than the corresponding threshold, the first expansion processing is performed, which can reduce the recognition error caused by the posture of the person or partial occlusion, and obtain a more accurate person expansion frame. Based on the person expansion frame, fusion processing is performed, and the spatial relationship between people can be effectively identified to obtain a fusion area, so that it can be accurately judged whether there is a crowd gathering in the target area according to the fusion area, so as to avoid the influence of environmental factors and improve the accuracy of crowd gathering recognition.

[0004] In a first aspect, an embodiment of the present invention provides a method for identifying a crowd gathering, the method comprising the following steps:

[0005] Acquire the image to be identified in the target area;

[0006] Performing recognition processing based on the image to be recognized, obtaining the number of people and a person recognition frame in the image to be recognized;

[0007] When the number of personnel is greater than a preset first personnel number threshold, performing a first expansion process based on the personnel identification frame to obtain a personnel expansion frame corresponding to each of the personnel identification frames;

[0008] Perform fusion processing based on the personnel expansion frame to obtain at least one fusion area;

[0009] Based on the fusion area, determine whether there is a crowd gathering in the target area.

[0010] Optionally, performing a first expansion process based on the person identification frame to obtain a person expansion frame corresponding to each person identification frame includes:

[0011] When the person recognition frame is a head recognition frame, performing a second expansion process based on the head recognition frame to obtain a human body expansion frame corresponding to the head recognition frame;

[0012] The first expansion process is performed based on the human body expansion frame to obtain a person expansion frame corresponding to the person identification frame.

[0013] Optionally, performing a second expansion process based on the head recognition frame to obtain a human body expansion frame corresponding to the head recognition frame includes:

[0014] Determining an expansion ratio corresponding to the human body expansion frame based on the scene type of the image to be recognized;

[0015] Based on the expansion ratio, a second expansion process is performed on the head recognition frame in a preset direction to obtain a human body expansion frame corresponding to the head recognition frame.

[0016] Optionally, the performing fusion processing based on the personnel expansion frame to obtain at least one fusion area includes:

[0017] Performing human body detection processing based on the person identification frame corresponding to the person expansion frame, and obtaining the human body confidence corresponding to each person expansion frame;

[0018] Filtering the personnel expansion frame based on the human body confidence to obtain a target personnel expansion frame;

[0019] Fusion processing is performed based on the target person expansion frame to obtain at least one fusion area.

[0020] Optionally, the performing fusion processing based on the target person expansion frame to obtain at least one fusion area includes:

[0021] Calculate the intersection-and-union ratio between adjacent target person expansion frames;

[0022] Based on the intersection-and-union ratio, the person identification frames corresponding to the adjacent target person expansion frames are fused to obtain at least one fused area.

[0023] Optionally, determining whether there is a crowd gathering in the target area based on the fusion area includes:

[0024] determining the number of persons in said fusion area;

[0025] Comparing the number of personnel with a preset second number threshold of personnel to obtain a number comparison result;

[0026] Based on the quantity comparison result, determine whether there is a crowd gathering in the target area.

[0027] Optionally, the fusion area corresponds to area coordinates, and the area coordinates are determined according to image coordinates of the person identification frame in the image to be identified, and determining whether there is a crowd gathering in the target area based on the fusion area includes:

[0028] determining the number of persons in said fusion area;

[0029] Based on the region coordinates, the region area of ​​the fusion region is calculated;

[0030] Based on the number of people in the fusion area and the area of ​​the area, a personnel aggregation coefficient of the fusion area is calculated;

[0031] Based on the personnel gathering coefficient and a preset coefficient threshold, determine whether there is a crowd gathering in the target area.

[0032] In a second aspect, an embodiment of the present invention further provides a crowd gathering recognition device, the crowd gathering recognition device comprising:

[0033] An acquisition module, used for acquiring an image to be identified in a target area;

[0034] A recognition module, used to perform recognition processing based on the image to be recognized, and obtain the number of people and a person recognition frame in the image to be recognized;

[0035] An expansion module, configured to perform a first expansion process based on the personnel identification frames to obtain a personnel expansion frame corresponding to each of the personnel identification frames when the number of the personnel is greater than a preset first personnel number threshold;

[0036] A fusion module, used for performing fusion processing based on the personnel expansion frame to obtain at least one fusion area;

[0037] A determination module is used to determine whether there is a crowd gathering in the target area based on the fusion area.

[0038] In a third aspect, an embodiment of the present invention provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the crowd gathering identification method provided in an embodiment of the present invention when executing the computer program.

[0039] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the crowd gathering identification method provided in the embodiment of the invention are implemented.

[0040] In an embodiment of the present invention, an image to be identified in a target area is obtained; identification processing is performed based on the image to be identified to obtain the number of people and a person identification frame in the image to be identified; when the number of people is greater than a preset first threshold of the number of people, a first expansion processing is performed based on the person identification frame to obtain a person expansion frame corresponding to each person identification frame; a fusion processing is performed based on the person expansion frame to obtain at least one fusion area; based on the fusion area, it is determined whether there is a crowd gathering in the target area. By performing identification processing on the image to be identified, the number of people and a person identification frame in the image to be identified can be obtained. When the number of people is greater than the corresponding threshold, the first expansion processing is performed, which can reduce the identification error caused by the posture of the person or partial occlusion, and obtain a more accurate person expansion frame. Fusion processing is performed based on the person expansion frame, which can effectively identify the spatial relationship between people and obtain a fusion area, so that it can accurately judge whether there is a crowd gathering in the target area based on the fusion area, so as to avoid the influence of environmental factors and improve the accuracy of crowd gathering identification. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0042] Figure 1 is a flow chart of a crowd gathering identification method provided by an embodiment of the present invention;

[0043] Figure 2 is a structural schematic diagram of a crowd gathering identification device provided in an embodiment of the present invention;

[0044] Figure 3 It is a structural schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0045] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0046] like Figure 1 As shown, Figure 1 : is a flow chart of a crowd gathering recognition method provided by an embodiment of the present invention, including:

[0047] 101. Obtain an image to be identified in a target area.

[0048] In an embodiment of the present invention, the crowd gathering recognition method can be applied to a security management platform. The security management platform can be constructed by a server or a server cluster. The server or server cluster can be any electronic device with image processing, image recognition, data storage, data transmission and other functions. The target area can be any public place where an image acquisition device is deployed, such as a park, a railway station, a commercial street, etc. Alternatively, the target area can be an area corresponding to the shooting range of an image acquisition device deployed in a public place, that is, a real area corresponding to the image area of ​​the image to be recognized.

[0049] The security management platform can acquire the image to be identified through an image acquisition device deployed in the target area, and implement the crowd gathering identification method according to the image to be identified to determine whether there is a crowd gathering in the target area.

[0050] It should be noted that after the above-mentioned image to be identified is acquired, the above-mentioned image to be identified can also be stored as a historically acquired image in the above-mentioned security management platform. When the data volume of the above-mentioned historically acquired image is sufficient (for example, a time threshold is satisfied), crowd gathering prediction processing can also be performed based on the above-mentioned historically acquired image and the image to be identified acquired at the current moment, to predict whether a crowd will gather in the target area at a second moment, and the above-mentioned second moment is later than the above-mentioned current moment in time.

[0051] Specifically, the above prediction processing can be implemented through machine learning or deep learning models. By setting a time threshold, analyzing the speed of the growth or shrinkage trend of the person identification frame within the time threshold (or understood as a short period of time), the human body trend can be analyzed. When the person identification frame shows a shrinking trend, it can be judged that the person is moving away from the deployment position of the above-mentioned image acquisition device. When the person identification frame shows an increasing trend, it can be judged that the person is moving close to the deployment position of the image acquisition device. Alternatively, the probability of crowd gathering in a specified direction or position can be judged based on the displacement direction of the person identification frame. Through the human body trend (i.e., the growth or shrinkage trend of the person identification frame), it is possible to analyze the target area and predict whether crowd gathering is occurring or is about to occur.

[0052] Furthermore, the image acquisition device deployed in the target area may be single or multiple in any number. When multiple image acquisition devices are deployed in the target area, multiple images to be identified may be acquired at the same time, and each image to be identified corresponds to one image acquisition device.

[0053] 102. Perform recognition processing based on the image to be recognized to obtain the number of people and the person recognition frame in the image to be recognized.

[0054] In the embodiment of the present invention, the above recognition process can be implemented by any personnel recognition algorithm, for example, the above personnel recognition algorithm can be a YOLO (You Only Look Once) algorithm, a DensePose algorithm, a mono2d-body-detection algorithm, etc. Specifically, a corresponding personnel recognition model can be constructed and trained by the above personnel recognition algorithm, and the above image to be recognized is provided to the above personnel recognition model, and the above personnel recognition model can perform recognition processing on the above image to be recognized to obtain the personnel recognition frame and the number of personnel in the above image to be recognized.

[0055] It should be noted that the above-mentioned personnel identification frame can be a partial personnel identification frame or a complete personnel identification frame. The above-mentioned partial personnel identification frame can be a head recognition frame or a human body recognition frame, and the above-mentioned human body recognition frame can be a complete human body recognition frame or a partial human body recognition frame. The above-mentioned number of personnel corresponds to the above-mentioned personnel identification frame. Specifically, the number of the above-mentioned personnel identification frames can be counted (that is, the number of complete human body recognition frames, partial human body recognition frames, head recognition frames and complete personnel recognition frames), and the above-mentioned number of personnel can be obtained after the statistics are completed. It can be simply understood that when only the head is detected or the head is blocked and only a complete human body or a partial human body is detected, it is also counted as one person.

[0056] 103. When the number of personnel is greater than a preset first personnel number threshold, a first expansion process is performed based on the personnel identification frame to obtain a personnel expansion frame corresponding to each personnel identification frame.

[0057] In an embodiment of the present invention, the first threshold of the number of people can be set based on historical crowd gathering recognition experience, or obtained based on a limited number of tests, or can be set based on environmental changes in the target area, and the environmental changes can be changes in parameters such as contrast, brightness, and weather. For example, when the light intensity is lower than the light intensity threshold, the first threshold of the number of people can be reduced by 10%, and when the weather is rainy or foggy, the threshold can be reduced by 15%.

[0058] The above-mentioned first expansion processing can be based on the center of the personnel identification frame and perform expansion processing according to a preset expansion ratio. The above-mentioned expansion ratio can be set according to historical personnel gathering identification experience, or obtained according to a limited number of experiments, for example, it can be 1 times, 1.5 times, etc.

[0059] After the first expansion processing, each person identification box can correspond to a person expansion box. In order to distinguish the person identification box from the person expansion box in the image to be identified, the person identification box and the person expansion box can be distinguished by using different border colors, or by labeling, etc.

[0060] It should be noted that if the number of people is less than or equal to the first number threshold, there is no need to perform the first expansion process, and the process can return to step 101 to obtain a new image to be identified (ie, the image to be identified at the next moment).

[0061] 104. Perform fusion processing based on the personnel expansion frame to obtain at least one fusion area.

[0062] In an embodiment of the present invention, the above-mentioned fusion processing can be to use the image area occupied by each two adjacent expansion frames as a basis for judging whether the personnel identification frames corresponding to the two expansion frames need to be fused, thereby fusion processing is performed on the two personnel identification frames. After judging whether the personnel identification frames corresponding to each two personnel expansion frames need to be fused, and fusing the personnel identification frames that need to be fused, a regular or irregular area can be obtained as the above-mentioned fusion area.

[0063] Specifically, if there is an intersection between the personnel expansion frames or the minimum distance between two similar frames is within a threshold range, the personnel identification frames corresponding to the two personnel expansion frames are fused into one region. This can be achieved specifically through an adaptive frame fusion algorithm, which can be adjusted based on multiple conditions (such as the similarity between frames, the distance between frames, etc.). The adaptive frame fusion algorithm can be a machine learning algorithm, a clustering algorithm, etc.

[0064] More specifically, when judging whether the two personnel identification frames corresponding to the two personnel expansion frames need to be merged by judging whether the inter-frame distance (i.e., the distance between two adjacent personnel expansion frames) is less than a set distance threshold, the above distance threshold can be dynamically adjusted according to the degree of crowding of the actual scene in the above image to be identified, the angle and resolution of the above image acquisition device, so as to enhance adaptability in different environments.

[0065] 105. Based on the fusion area, determine whether there is crowd gathering in the target area.

[0066] In an embodiment of the present invention, the number of personnel identification frames in the fusion area may be counted to obtain the number of personnel in the fusion area, or the number of personnel in the fusion area may be obtained by performing personnel identification processing on the fusion area again.

[0067] After determining the number of people in the above-mentioned fusion area, the above-mentioned number of people can be compared with a preset second threshold of the number of people to obtain a comparison result, and based on the above-mentioned comparison result, it can be judged whether there is a crowd gathering in the above-mentioned target area.

[0068] Alternatively, the above-mentioned fusion area corresponds to area coordinates, and the above-mentioned area coordinates can be determined according to the image coordinates of the above-mentioned personnel identification box in the above-mentioned image to be identified. The image area of ​​the above-mentioned fusion area is calculated through the area coordinates corresponding to the above-mentioned fusion area, and the personnel aggregation coefficient is calculated according to the above-mentioned number of personnel and the above-mentioned image area, and it is judged whether there is a crowd gathering in the above-mentioned target area according to the above-mentioned personnel aggregation coefficient.

[0069] In an embodiment of the present invention, an image to be identified in a target area is obtained; identification processing is performed based on the image to be identified to obtain the number of people and a person identification frame in the image to be identified; when the number of people is greater than a preset first threshold of the number of people, a first expansion processing is performed based on the person identification frame to obtain a person expansion frame corresponding to each person identification frame; a fusion processing is performed based on the person expansion frame to obtain at least one fusion area; based on the fusion area, it is determined whether there is a crowd gathering in the target area. By performing identification processing on the image to be identified, the number of people and a person identification frame in the image to be identified can be obtained. When the number of people is greater than the corresponding threshold, the first expansion processing is performed, which can reduce the identification error caused by the posture of the person or partial occlusion, and obtain a more accurate person expansion frame. Fusion processing is performed based on the person expansion frame, which can effectively identify the spatial relationship between people and obtain a fusion area, so that it can accurately judge whether there is a crowd gathering in the target area based on the fusion area, so as to avoid the influence of environmental factors and improve the accuracy of crowd gathering identification.

[0070] It is understandable that in the specific implementation of this application, data related to images to be identified, scene types, etc. are involved. When the embodiments in this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data and the construction, training and use of related models need to comply with relevant laws, regulations and standards of relevant countries and regions.

[0071] It should be noted that the crowd gathering identification method provided in the embodiment of the present invention can be applied to computers, servers and other devices that can perform crowd gathering identification.

[0072] Optionally, in the step of performing a first expansion processing based on the personnel identification frame to obtain a personnel expansion frame corresponding to each personnel identification frame, when the personnel identification frame is a head recognition frame, a second expansion processing can be performed based on the head recognition frame to obtain a human body expansion frame corresponding to the head recognition frame; and a first expansion processing can be performed based on the human body expansion frame to obtain a personnel expansion frame corresponding to the personnel identification frame.

[0073] In the embodiment of the present invention, when the above-mentioned person recognition frame is a head recognition frame, if the above-mentioned head recognition frame is directly expanded, the obtained person expansion frame is usually not able to directly select the whole person well, so the above-mentioned head recognition frame can be firstly expanded to obtain the human body expansion frame corresponding to the head recognition frame. Thus, the first expansion process can be performed on the human body expansion frame to obtain the corresponding person expansion frame, which can well select the whole person.

[0074] Specifically, usually, the head is above the human body and smaller than the human body, so the expansion direction of the second expansion process can be determined according to the positional relationship between the head and the human body, and the expansion ratio of the second expansion process can be determined according to the proportional relationship between the head and the human body. Thus, the second expansion process is performed according to the expansion direction and the expansion ratio to obtain the human body expansion frame.

[0075] It should be noted that the head recognition frame can be a head recognition frame obtained by detecting the head, or a face recognition frame obtained by detecting only the face, and the face recognition frame is used as the head recognition frame. Generally, the detection error rate of the face recognition frame is lower than the detection error rate of the head recognition frame. Therefore, if the head recognition frame is obtained by using the face as the detection target, its detection accuracy will be improved accordingly. The specific settings can be made according to actual needs.

[0076] Optionally, in the step of performing a second expansion processing based on the head recognition frame to obtain a human body expansion frame corresponding to the head recognition frame, the expansion ratio corresponding to the human body extension frame can also be determined based on the scene type of the image to be recognized; based on the expansion ratio, the head recognition frame is subjected to a second expansion processing in a preset direction to obtain a human body extension frame corresponding to the head recognition frame.

[0077] In an embodiment of the present invention, the above-mentioned scene type may be a scene type for long-distance shooting or a scene type for close-up shooting. Since the human body and objects in the scene are far away from the image acquisition device during long-distance shooting, the image resolution is low and the details may be blurred, while the human body and objects in the scene are close to the image acquisition device during close-up shooting, the image details are richer and the resolution is higher. Therefore, a basic expansion ratio can be set according to the proportional relationship between the head and the human body (i.e., the proportional relationship between the head and the human body in the actual scene), and the above-mentioned basic expansion ratio is adjusted accordingly according to the above-mentioned scene type to obtain the expansion ratio of the above-mentioned second expansion processing. For example, when the above-mentioned scene type is long-distance shooting, since its image details may be blurred and the resolution is low, the above-mentioned expansion ratio can be correspondingly enlarged (the farther the distance, the larger the above-mentioned expansion ratio can be adjusted accordingly). On the contrary, when the above-mentioned scene type is close-up shooting, since its image details are richer and the resolution is higher, the above-mentioned expansion ratio can be correspondingly reduced (the closer the distance, the smaller the above-mentioned expansion ratio can be adjusted accordingly).

[0078] For example, the above expansion ratio may include a horizontal expansion ratio and a vertical expansion ratio. According to the position of the head recognition frame, the frame may be expanded to both sides (i.e., horizontally) and downwards at a certain ratio, with both sides expanding by 1 times (i.e., the horizontal expansion ratio) and downwards expanding by 6 times (i.e., the vertical expansion ratio). This frame is used as a human body expansion frame (the detected position of the head recognition frame is expanded to both sides and downwards at a specific ratio to infer the human body area.

[0079] In a possible embodiment, the image acquisition device may further include a depth perception sensor, and the expansion ratio may be adaptively adjusted in combination with the depth perception sensor.

[0080] Optionally, in the step of performing fusion processing based on the personnel expansion frame to obtain at least one fusion area, human body detection processing can also be performed based on the personnel recognition frame corresponding to the personnel expansion frame to obtain the human body confidence corresponding to each personnel expansion frame; the personnel expansion frame is filtered based on the human body confidence to obtain the target personnel expansion frame; and fusion processing is performed based on the target personnel expansion frame to obtain at least one fusion area.

[0081] In an embodiment of the present invention, the above-mentioned human body detection process may be live body detection. For example, assuming that there are a person identification frame A1 and a person identification frame A2, the person in the person identification frame A1 is actually a person in the billboard, and the person in the person identification frame A2 is a live person. At this time, the human body confidence of the person identification frame A1 is less than the human body confidence of the person identification frame A2. The above-mentioned human body detection process may be implemented by any human body detection algorithm based on deep learning, such as the FasterR-CNN algorithm, Detectron2, etc.

[0082] After determining the human confidence corresponding to each personnel expansion frame, the personnel expansion frame whose human confidence is greater than the preset human confidence threshold can be used as the target personnel expansion frame, and the personnel expansion frame whose human confidence is less than the preset human confidence threshold can be eliminated.

[0083] Therefore, after the fusion processing is performed based on the target person expansion frame, the influence of billboards, projections or other objects on the fusion area can be avoided, and a more accurate fusion area can be obtained.

[0084] Optionally, in the step of performing fusion processing based on the target personnel expansion frame to obtain at least one fusion area, the intersection-and-union ratio between adjacent target personnel expansion frames can also be calculated; based on the intersection-and-union ratio, the personnel identification frames corresponding to the adjacent target personnel expansion frames are fused to obtain at least one fusion area.

[0085] In the embodiment of the present invention, it can be understood that since the personnel expansion frame is obtained by expanding based on the personnel identification frame, and the target personnel expansion frame is selected in the personnel expansion frame, one target personnel expansion frame can correspond to one personnel identification frame, and the above intersection-union ratio can be understood as the ratio between the intersection area and the union area between the two target personnel expansion frames. Whether the two target personnel expansion frames are adjacent is understood as whether the image distance of the two target personnel expansion frames in the image to be identified is less than a preset image distance threshold. If it is less than, it means that the two target personnel expansion frames are adjacent, and if it is greater, it means that the two target personnel expansion frames are not adjacent.

[0086] Specifically, the above intersection-and-union ratio can be further explained by the following formula:

[0087]

[0088] Among them, the above IOU is expressed as the above intersection-overlap ratio, the above Area of ​​Overlap is expressed as the intersection area between the two target person expansion boxes, and the above Area of ​​Union is expressed as the union area between the two target person expansion boxes.

[0089] After calculating the intersection-and-union ratio between two adjacent target personnel expansion frames, the above intersection-and-union ratio can be compared with a preset intersection-and-union ratio coefficient (for example, 0.4). It is then considered that there is an association between the two adjacent target personnel expansion frames. All associated target personnel expansion frames are recorded, and the personnel identification frames corresponding to all associated target personnel expansion frames are fused to obtain a fusion area.

[0090] Optionally, in the step of determining whether there is a crowd gathering in the target area based on the fusion area, the number of people in the fusion area can also be determined; the number of people is compared with a preset second number of people threshold to obtain a quantity comparison result; based on the quantity comparison result, it is determined whether there is a crowd gathering in the target area.

[0091] In an embodiment of the present invention, the number of people in the above-mentioned fusion area can be obtained by counting the person identification boxes in the above-mentioned fusion area, or can also be obtained by re-detection through a person detection algorithm. The above-mentioned preset second number of people threshold can be set by historical crowd gathering recognition experience, or obtained through a limited number of experiments. By comparing the number of people with the preset second number of people threshold, a quantity comparison result is obtained. When the quantity comparison result is that the number of people is greater than or equal to the second number of people threshold, it is determined that there is a crowd gathering in the target area. On the contrary, when the quantity comparison result is that the number of people is less than the second number of people threshold, it is determined that there is no crowd gathering in the target area.

[0092] Optionally, in the step of determining whether there is a crowd gathering in the target area based on the fusion area, the number of people in the fusion area can also be determined; the area of ​​the fusion area can be calculated based on the area coordinates; the people gathering coefficient of the fusion area can be calculated based on the number of people and the area of ​​the fusion area; based on the people gathering coefficient and a preset coefficient threshold, it is determined whether there is a crowd gathering in the target area.

[0093] In an embodiment of the present invention, the above-mentioned fusion area corresponds to regional coordinates, and the above-mentioned regional coordinates are determined according to the image coordinates of the above-mentioned person identification frame in the above-mentioned image to be identified. Specifically, since the above-mentioned fusion area is obtained by frame fusion of the above-mentioned person identification frames, after the person identification frames are frame fused with each other, the coordinates of the corner points (i.e., the upper left corner, lower left corner, upper right corner, or lower right corner) of the person identification frame at the outermost edge of the fusion area are the regional coordinates of the fusion area. Specifically, the person identification frame at the outermost edge of the fusion area can be determined in the fusion area, and the target corner point of the outermost person identification frame can be determined. Specifically, the target corner point can be determined according to the orientation of the outermost person identification frame in the fusion area (for example, assuming that the outermost person identification frame is in the upper left corner area of ​​the fusion area, the target corner point can include at least one of the upper left corner, lower left corner, and upper right corner of the outermost person identification frame, but not the lower right corner). By using the coordinates of the target corner points as the above-mentioned area coordinates and connecting each adjacent area coordinate, the shape of the fusion area can be determined. Based on the shape of the fusion area, the calculation formula of the corresponding geometric shape can be determined, and the area of ​​the above-mentioned area can be determined based on the calculation formula and the above-mentioned area coordinates.

[0094] It can be understood that the above-mentioned fusion area can be an irregular shape formed by the fusion of multiple personnel identification frames (i.e., the fusion of rectangular frames). Therefore, the above-mentioned fusion area can be divided into multiple different geometric figures, such as triangles, rectangles, etc., so as to calculate the area of ​​the above-mentioned area according to the calculation formula of the geometric shape corresponding to each geometric figure and the corresponding area coordinates.

[0095] The above personnel aggregation coefficient can be further explained by the following formula:

[0096]

[0097] The CDC is represented by the population cluster coefficient, the Population Count is represented by the number of people in the fusion area, and the Area of ​​Region is represented by the area of ​​the fusion area. The preset coefficient threshold can also be set based on historical crowd cluster recognition experience or obtained through a limited number of experiments.

[0098] When determining the above-mentioned personnel gathering coefficient and the above-mentioned preset coefficient threshold, the above-mentioned personnel gathering coefficient can be compared with the above-mentioned preset coefficient threshold. If the above-mentioned personnel gathering coefficient is greater than or equal to the above-mentioned preset coefficient threshold, it is determined that there is crowd gathering in the target area. Conversely, if the above-mentioned personnel gathering coefficient is less than the above-mentioned preset coefficient threshold, it is determined that there is no crowd gathering in the target area.

[0099] like Figure 2As shown, an embodiment of the present invention further provides a crowd gathering recognition device, comprising:

[0100] An acquisition module 201 is used to acquire an image to be identified in a target area;

[0101] The recognition module 202 is used to perform recognition processing based on the image to be recognized, and obtain the number of people and the person recognition frame in the image to be recognized;

[0102] An expansion module 203 is used to perform a first expansion process based on the personnel identification frame to obtain a personnel expansion frame corresponding to each personnel identification frame when the number of personnel is greater than a preset first personnel number threshold;

[0103] A fusion module 204, configured to perform fusion processing based on the personnel expansion frame to obtain at least one fusion area;

[0104] The determination module 205 is used to determine whether there is a crowd gathering in the target area based on the fusion area.

[0105] Optionally, the expansion module 203 includes:

[0106] A first expansion submodule is used for, when the person recognition frame is a head recognition frame, performing a second expansion process based on the head recognition frame to obtain a human body expansion frame corresponding to the head recognition frame;

[0107] The second expansion submodule is used to perform the first expansion process based on the human body expansion frame to obtain a person expansion frame corresponding to the person identification frame.

[0108] Optionally, the first expansion submodule includes:

[0109] A first determining unit, configured to determine an expansion ratio corresponding to the human body expansion frame based on a scene type of the image to be recognized;

[0110] The first expansion unit is used to perform a second expansion process on the head recognition frame according to a preset direction based on the expansion ratio to obtain a human body expansion frame corresponding to the head recognition frame.

[0111] Optionally, the fusion module 204 includes:

[0112] A first detection submodule is used to perform human body detection processing based on the person identification frame corresponding to the person expansion frame, and obtain the human body confidence corresponding to each person expansion frame;

[0113] A first filtering submodule, configured to filter the person expansion frame based on the human body confidence level to obtain a target person expansion frame;

[0114] The first fusion submodule is used to perform fusion processing based on the target person expansion frame to obtain at least one fusion area.

[0115] Optionally, the first fusion submodule includes:

[0116] A first calculation unit is used to calculate the intersection-and-union ratio between adjacent target person expansion frames;

[0117] The first fusion unit is used to fuse the person identification frames corresponding to the adjacent target person expansion frames based on the intersection-over-union ratio to obtain at least one fusion area.

[0118] Optionally, the determining module 205 includes:

[0119] A first determination submodule is used to determine the number of people in the fusion area;

[0120] A first comparison submodule, used to compare the number of personnel with a preset second number threshold of personnel to obtain a quantity comparison result;

[0121] The second determination submodule is used to determine whether there is a crowd gathering in the target area based on the quantity comparison result.

[0122] Optionally, the fusion area corresponds to area coordinates, and the area coordinates are determined according to the image coordinates of the person identification frame in the image to be identified. The determination module 205 includes:

[0123] A third determination submodule is used to determine the number of people in the fusion area;

[0124] A first calculation submodule, configured to calculate the area of ​​the fusion region based on the region coordinates;

[0125] A second calculation submodule is used to calculate the personnel aggregation coefficient of the fusion area based on the number of personnel in the fusion area and the area of ​​the area;

[0126] The fourth determination submodule is used to determine whether there is crowd gathering in the target area based on the crowd gathering coefficient and a preset coefficient threshold.

[0127] like Figure 3 As shown, an embodiment of the present invention further provides an electronic device, characterized in that it includes a processor, and the processor can execute any one of the above-mentioned crowd gathering identification methods.

[0128] Specifically, it includes a processor 301 and a memory 302, and a computer program for executing a crowd gathering recognition method stored in the memory 302 and capable of running on the processor 301, wherein:

[0129] The processor 301 runs the computer program of the crowd gathering recognition method stored in the memory 302 and performs the following steps:

[0130] Acquire the image to be identified in the target area;

[0131] Performing recognition processing based on the image to be recognized, obtaining the number of people and a person recognition frame in the image to be recognized;

[0132] When the number of personnel is greater than a preset first personnel number threshold, performing a first expansion process based on the personnel identification frame to obtain a personnel expansion frame corresponding to each of the personnel identification frames;

[0133] Perform fusion processing based on the personnel expansion frame to obtain at least one fusion area;

[0134] Based on the fusion area, determine whether there is a crowd gathering in the target area.

[0135] Optionally, the processor 301 performs the first expansion processing based on the person identification frame to obtain a person expansion frame corresponding to each person identification frame, including:

[0136] When the person recognition frame is a head recognition frame, performing a second expansion process based on the head recognition frame to obtain a human body expansion frame corresponding to the head recognition frame;

[0137] The first expansion process is performed based on the human body expansion frame to obtain a person expansion frame corresponding to the person identification frame.

[0138] Optionally, the processor 301 performs the second expansion processing based on the head recognition frame to obtain a human body expansion frame corresponding to the head recognition frame, including:

[0139] Determining an expansion ratio corresponding to the human body expansion frame based on the scene type of the image to be recognized;

[0140] Based on the expansion ratio, a second expansion process is performed on the head recognition frame in a preset direction to obtain a human body expansion frame corresponding to the head recognition frame.

[0141] Optionally, the processor 301 performs fusion processing based on the personnel expansion frame to obtain at least one fusion area, including:

[0142] Performing human body detection processing based on the person identification frame corresponding to the person expansion frame, and obtaining the human body confidence corresponding to each person expansion frame;

[0143] Filtering the personnel expansion frame based on the human body confidence to obtain a target personnel expansion frame;

[0144] Fusion processing is performed based on the target person expansion frame to obtain at least one fusion area.

[0145] Optionally, the processor 301 performs fusion processing based on the target person expansion frame to obtain at least one fusion area, including:

[0146] Calculate the intersection-and-union ratio between adjacent target person expansion frames;

[0147] Based on the intersection-and-union ratio, the person identification frames corresponding to the adjacent target person expansion frames are fused to obtain at least one fused area.

[0148] Optionally, the determining, based on the fusion area, whether there is a crowd gathering in the target area performed by the processor 301 includes:

[0149] determining the number of persons in said fusion area;

[0150] Comparing the number of personnel with a preset second number threshold of personnel to obtain a number comparison result;

[0151] Based on the quantity comparison result, determine whether there is a crowd gathering in the target area.

[0152] Optionally, the fusion area corresponds to area coordinates, and the area coordinates are determined according to image coordinates of the person identification frame in the image to be identified. The processor 301 determines whether there is a crowd gathering in the target area based on the fusion area, including:

[0153] determining the number of persons in said fusion area;

[0154] Based on the region coordinates, the region area of ​​the fusion region is calculated;

[0155] Based on the number of people in the fusion area and the area of ​​the area, a personnel aggregation coefficient of the fusion area is calculated;

[0156] Based on the personnel gathering coefficient and a preset coefficient threshold, determine whether there is a crowd gathering in the target area.

[0157] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the computer program implements the various processes of the crowd gathering identification method or the application-side crowd gathering identification method provided in the embodiment of the present invention, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0158] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing related hardware through a computer program, and the above-mentioned computer program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, the above-mentioned computer-readable storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.

[0159] The above disclosure is only the preferred embodiment of the present invention, which certainly cannot be used to limit the scope of the present invention. Therefore, equivalent changes made according to the claims of the present invention are still within the scope of the present invention.

Claims

1. A method for identifying crowd gatherings, characterized in that: The method comprises the following steps: Obtain an image to be identified in the target area; Performing recognition processing based on the image to be recognized, obtaining the number of people and a person recognition frame in the image to be recognized; When the number of personnel is greater than a preset first personnel number threshold, performing a first expansion process based on the personnel identification frame to obtain a personnel expansion frame corresponding to each of the personnel identification frames; Perform fusion processing based on the personnel expansion frame to obtain at least one fusion area; Based on the fusion area, determine whether there is a crowd gathering in the target area.

2. The crowd gathering recognition method according to claim 1, characterized in that: The first expansion process is performed based on the personnel identification frame to obtain a personnel expansion frame corresponding to each personnel identification frame, including: When the person recognition frame is a head recognition frame, performing a second expansion process based on the head recognition frame to obtain a human body expansion frame corresponding to the head recognition frame; The first expansion process is performed based on the human body expansion frame to obtain a person expansion frame corresponding to the person identification frame.

3. The crowd gathering recognition method according to claim 2, characterized in that: The performing a second expansion process based on the head recognition frame to obtain a human body expansion frame corresponding to the head recognition frame includes: Determining an expansion ratio corresponding to the human body expansion frame based on the scene type of the image to be recognized; Based on the expansion ratio, a second expansion process is performed on the head recognition frame in a preset direction to obtain a human body expansion frame corresponding to the head recognition frame.

4. The crowd gathering recognition method according to claim 1, characterized in that: The fusion processing is performed based on the personnel expansion frame to obtain at least one fusion area, including: Performing human body detection processing based on the person identification frame corresponding to the person expansion frame, and obtaining the human body confidence corresponding to each person expansion frame; Filtering the personnel expansion frame based on the human body confidence to obtain a target personnel expansion frame; Fusion processing is performed based on the target person expansion frame to obtain at least one fusion area.

5. The crowd gathering recognition method according to claim 4, characterized in that: The fusion processing is performed based on the target person expansion frame to obtain at least one fusion area, including: Calculate the intersection-and-union ratio between adjacent target person expansion frames; Based on the intersection-and-union ratio, the person identification frames corresponding to the adjacent target person expansion frames are fused to obtain at least one fused area.

6. The crowd gathering recognition method according to claim 1, characterized in that: The determining, based on the fusion area, whether there is a crowd gathering in the target area includes: determining the number of persons in said fusion area; Comparing the number of personnel with a preset second number threshold of personnel to obtain a number comparison result; Based on the quantity comparison result, determine whether there is a crowd gathering in the target area.

7. The crowd gathering recognition method according to claim 1, characterized in that: The fusion area corresponds to area coordinates, and the area coordinates are determined according to the image coordinates of the person identification frame in the image to be identified. The determining whether there is a crowd gathering in the target area based on the fusion area includes: determining the number of persons in said fusion area; Based on the region coordinates, the region area of ​​the fusion region is calculated; Based on the number of people in the fusion area and the area of ​​the area, a personnel aggregation coefficient of the fusion area is calculated; Based on the personnel gathering coefficient and a preset coefficient threshold, determine whether there is a crowd gathering in the target area.

8. A crowd gathering recognition device, characterized in that: The crowd gathering identification device comprises: An acquisition module, used for acquiring an image to be identified in a target area; A recognition module, used to perform recognition processing based on the image to be recognized, and obtain the number of people and a person recognition frame in the image to be recognized; An expansion module, configured to perform a first expansion process based on the personnel identification frames to obtain a personnel expansion frame corresponding to each of the personnel identification frames when the number of the personnel is greater than a preset first personnel number threshold; A fusion module, used for performing fusion processing based on the personnel expansion frame to obtain at least one fusion area; A determination module is used to determine whether there is a crowd gathering in the target area based on the fusion area.

9. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps in the crowd gathering identification method as described in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps in the crowd gathering identification method as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Method for calculating number of plate-shaped objects

    CN121095510A

  • Material identification method, material sorting method, material identification device and equipment

    CN122230996A