Camera image processing system using artificial intelligence and crowd density analysis method therefor
The camera image processing system uses AI to detect heads, calculate coordinates, and group people for accurate crowd density analysis, preventing overcrowding by generating timely warnings.
Patent Information
- Application Number
- PCT/KR2024/019091
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-05
- Filing Date
- 2024-11-28
- Publication Date
- 2025-07-10
AI Technical Summary
Existing methods for detecting and analyzing crowd density in camera images face challenges such as occlusion, poor joint position identification, poor pose estimation, and long computational times, leading to inaccurate analysis.
A camera image processing system using artificial intelligence that detects heads in images, calculates camera coordinates, and determines actual coordinates to group people based on these, outputting warnings when crowd density exceeds a threshold.
Enables accurate crowd density analysis, preventing overcrowding-related accidents by calculating people per unit area and generating timely warnings.
Smart Images

Figure KR2024019091_10072025_PF_FP_ABST
Abstract
Description
Camera image processing system using artificial intelligence and crowd density analysis method therefor
[0001] The present invention relates to a camera image processing system using artificial intelligence and a crowd density analysis method therefor.
[0002] Recently, stampede accidents and crimes targeting unspecified numbers of people have been occurring in crowded areas such as downtown areas, subway stations, and public facilities.
[0003] To prevent such accidents and crimes, preventive measures such as cameras, CCTV, and patrols are being implemented. Among these preventive measures, utilizing cameras and CCTV is an effective method for conducting them with a limited workforce. However, because these methods are operated by a small number of surveillance operators, it is practically impossible to continuously monitor and take action in densely populated areas.
[0004] Meanwhile, according to conventional technology, three methods are mainly used to detect and analyze people in camera images.
[0005] The first is a method for detecting an area where a person is located in a square shape, the second is a method for identifying the joint positions of a person in an image and estimating the person's posture in the form of a skeleton based on the identified joint positions, and the third is a method for generating a label map by displaying pixels representing a person in an image.
[0006] However, these methods have problems such as not being able to detect the exact location of a person due to occlusion when there are multiple people in the image, difficulty in identifying the location of the person's joints, resulting in poor pose estimation performance, long computational times for generating label maps, and inaccurate analysis based on pixel-by-pixel analysis.
[0007] Therefore, a crowd density analysis method that can solve the problems of the prior art is required.
[0008] The present invention is intended to solve the problems of the above-mentioned prior art, and provides a camera image processing system using artificial intelligence and a crowd density analysis method therefor.
[0009] Additionally, the present invention provides a method for analyzing crowd density based on a head so as not to identify an object as a person.
[0010] A computing device according to one embodiment of the present invention includes a communication unit that receives an image from an image acquisition device; and a processor that detects a head of a person in the received image, calculates camera coordinates of the detected head based on pixel data constituting the image, and calculates actual coordinates of the person corresponding to the head using the calculated camera coordinates.
[0011] A computing device according to another embodiment of the present invention includes a communication unit that receives an image from an image acquisition device; and a processor that detects a head of a person in the received image, groups people concentrated in a specific area based on the detected head, and outputs a warning message to the outside when the number of people in a specific group exceeds a reference value.
[0012] In an image processing system according to one embodiment of the present invention, a crowd density analysis method includes the steps of: detecting a head of a person in an image received from an image acquisition device; obtaining camera coordinates of the detected head; obtaining actual coordinates of a person corresponding to the head using the camera coordinates; and detecting crowd density by grouping people crowded in a specific area based on the obtained actual coordinates.
[0013]
[0014] The camera image processing system using artificial intelligence of the present invention and the crowd density analysis method therefor can capture an observation target area with a camera and analyze the captured image using artificial intelligence.
[0015] In addition, the camera image processing system can calculate the number of people per unit area in the image to analyze the crowd density in the area, and generate a warning message for areas where the crowd is excessively dense, thereby preventing accidents resulting from overcrowding in advance.
[0016] FIG. 1 is a diagram illustrating an image processing system according to one embodiment of the present invention.
[0017] FIG. 2 is a block diagram illustrating the configuration of a server according to one embodiment of the present invention.
[0018] FIG. 3 is a drawing illustrating a box display method according to the distance between a subject and a camera according to one embodiment of the present invention.
[0019] FIG. 4 is a diagram illustrating the concept of internal parameters and external parameters according to one embodiment of the present invention.
[0020] FIG. 5 is a drawing illustrating the installation of an identifier according to one embodiment of the present invention.
[0021] FIG. 6 is a drawing illustrating a camera installed on an exterior wall of a building according to one embodiment of the present invention.
[0022] FIG. 7 is a drawing showing a structure in which a camera according to one embodiment of the present invention is installed in a moving means.
[0023] Figure 8 is a flowchart illustrating a crowd density analysis process according to one embodiment of the present invention.
[0024] As used herein, singular expressions include plural expressions unless the context clearly dictates otherwise. In this specification, terms such as "consist of" or "include" should not be construed to necessarily include all components or steps described in the specification, and should be construed to mean that some of the components or steps may not be included, or that additional components or steps may be included. In addition, terms such as "part" and "module" described in the specification mean a unit that processes at least one function or operation, which may be implemented by hardware or software, or by a combination of hardware and software.
[0025]
[0026] The present invention relates to a camera image processing system using artificial intelligence and a crowd density analysis method therefor, wherein an observation target area is photographed with a camera (image acquisition device), the photographed image is analyzed using artificial intelligence to calculate the number of people per unit area within the image, and crowd density within the area can be detected based on the calculation result.
[0027]
[0028] Hereinafter, various embodiments of the present invention will be described in detail with reference to the attached drawings.
[0029] FIG. 1 is a diagram illustrating an image processing system according to an embodiment of the present invention, FIG. 2 is a block diagram illustrating a configuration of a server according to an embodiment of the present invention, FIG. 3 is a diagram illustrating a method for displaying boxes according to the distance between a subject and a camera according to an embodiment of the present invention, and FIG. 4 is a diagram illustrating the concept of internal parameters and external parameters according to an embodiment of the present invention. FIG. 5 is a diagram illustrating the installation of an identifier according to an embodiment of the present invention, FIG. 6 is a diagram illustrating a camera installed on an outer wall of a building according to an embodiment of the present invention, FIG. 7 is a diagram illustrating a structure in which a camera is installed on a moving vehicle according to an embodiment of the present invention, and FIG. 8 is a flowchart illustrating a crowd density analysis process according to an embodiment of the present invention.
[0030] Referring to FIG. 1, the image processing system of the present embodiment may include a server (100, computing device), a camera (200), and a warning device (300).
[0031] The server (100) can analyze crowd density by processing images acquired by a camera, and as shown in FIG. 2, can include a communication module, a memory storing a program for image processing, a processor executing the program, and a DB (DataBase) storing data for executing the program.
[0032] The above communication module is a communication connection channel with external devices, such as a camera (200) and an alarm device (300).
[0033] The processor can perform various functions based on the execution of a program stored in the memory. For example, the processor can detect crowd density by processing images acquired by the camera (100) described below, and output a warning through the warning device (300) based on the detection results.
[0034] A camera (200) is a type of image acquisition device that can be installed on high terrain, such as the exterior wall of a building or a telephone pole, to photograph an observation target area (surveillance area).
[0035] These cameras (200) may be CCTV cameras, include PTZ (Pan, Tilt, Zoom) functions, and can communicate with the server (100) via wired or wireless means.
[0036] According to one embodiment, the camera (200) may additionally include a beam projector module, and may display a predetermined image at a point within the field of view captured by the camera (200) using the beam projector module.
[0037] The warning device (300) can output a warning message to the surroundings under the control of the server (100). The warning message may be a visually identifiable image, text, or warning sound. To this end, the warning device (300) may include a display unit for displaying the image or text, or a speaker for outputting the warning message as sound.
[0038] According to one embodiment, the server (100) can control the warning device (300) to output a warning message when the crowd density exceeds a standard value through image analysis.
[0039]
[0040] Below, we will examine the image processing process and crowd density analysis process in detail.
[0041] The camera (200) captures an image of the observation target area and the server (100) can receive an image of the observation target area from the camera (200). Here, the observation target area is an area monitored by the camera (200), and may be a narrow area with a dense crowd or an area with unstable terrain (e.g., an area with a ventilation duct or sewer).
[0042] According to one embodiment, an identifier (400) as illustrated in FIG. 5 may be installed in the observation target area, and the camera (200) may capture an image of the observation target area in which the identifier (400) is installed. At this time, the identifier (400) may be installed within the field of view of the camera (200) that captures the observation target area, or may be projected and displayed from the camera (200). That is, the identifier (400) may be a physical object or a non-physical object (e.g., an image).
[0043] For example, the identifier (400) may be prepared in advance at the location before the camera (200) performs a shooting operation, and may include a predetermined pattern having a repeating pattern so as not to be confused with sign information or guidance information existing at the location.
[0044] These identifiers (400) are used to estimate coordinates and can be modified in various ways as long as the coordinates can be estimated. For example, the identifier (400) may have a plate shape composed of grids with two distinct colors, as illustrated in FIG. 5.
[0045] Next, the processor of the server (100) analyzes the image received from the camera (200) using a preset program, for example, a learned head detection model (artificial intelligence), to detect the head of a person in the image, and can calculate the camera coordinates of the person corresponding to the detected head based on the pixel data constituting the image.
[0046] According to one embodiment, the head detection model is a model trained using as input a data set that matches an image of the observation target area and a head detection result. Accordingly, the head detection model can output a head detection result when an image of the observation target area is input. Here, the machine learning for training the head detection model may be supervised learning or unsupervised learning. However, the method for training the head detection model is not limited to the above learning method as long as the head detection model can detect a head.
[0047] According to one embodiment, the head detection model can output a human head among the subjects as a box (head identifier) as shown in FIG. 3.
[0048] At this time, as illustrated in FIG. 3, if the distance between the subject in the image captured by the camera (200) and the camera (200) is relatively far, the box-shaped head identifier may be displayed small, and if the distance between the subject and the camera (200) is relatively close, the box-shaped head identifier may be displayed large. That is, the size of the head identifier may vary depending on the distance between the subject and the camera (200).
[0049] Meanwhile, in Fig. 3, the head identifier is indicated by a box, but the shape is not limited to a box and various shapes can be used as long as the head can be indicated by distinguishing it from other body parts.
[0050] Continuing, the processor of the server (100) can calculate the actual coordinates (absolute coordinates) of the person corresponding to the detected head.
[0051] To calculate these coordinates, the server (100) can utilize internal parameters and external parameters.
[0052] In one embodiment, the external parameter may be a three-dimensional coordinate system representing the distance between the camera (200) and the ground, and the internal parameter may be a two-dimensional coordinate system representing the focal distance between the camera (200) and an image captured by the camera (200).
[0053] These parameters can be expressed as a rotation matrix (R) and a translation matrix (t) composed of coordinate values, and the server (100) can obtain a transformation variable between the camera coordinates and the actual coordinates by calculating the external parameters and the internal parameters.
[0054] To this end, the processor of the server (100) can utilize identifiers (400) present at various locations within the image by the camera (200).
[0055] Since the identifier (400) is composed of grids with two different colors and the length and width of the grids are predetermined, the processor of the server (100) can detect the camera coordinates for the identifier (400) in the image received from the camera (200).
[0056] Referring to FIG. 4, when the external parameters are defined as a rotation matrix R and a translation matrix t, and the internal parameters are defined as K, the 3D real coordinates (X, Y, Z) and the 2D image coordinates (u, v) can have the following relationship. Here, W is a matrix (scale factor) that plays a role in adjusting the size, and the relationship between the real coordinate system, the camera coordinate system, and the image coordinate system can be configured as shown in FIG. 4.
[0057] According to one embodiment, the processor of the server (100) can compare the camera coordinates of the identifier (400) in the image captured by the camera (200) with the pre-stored identifier (400) to calculate the actual coordinates of the identifier (400) in the image, and can calculate the actual coordinates of the detected head based on the actual coordinates of the identifier (400) in the image and the camera coordinates of the head detected in the image.
[0058] Referring to FIG. 5, the server (100) detects camera coordinates for an identifier (400) in an image received from a camera (200), and then establishes an equation based on the actual coordinates and camera coordinates for the identifier (400) to obtain a transformation matrix (relational expression) between the actual coordinates and the camera coordinates, and inputs the camera coordinates of the head into the transformation matrix to calculate the actual coordinates of the head.
[0059] At this time, one identifier (400) may be placed in the observation target area, or multiple identifiers (400) may be placed in non-overlapping positions to more easily obtain conversion variables between camera coordinates and actual coordinates. At this time, the identifiers (400) may be placed in a number of identifiers (400) capable of calculating the corresponding variables according to the number of variables in the equation. For example, in the case of a cubic equation, since there are three variables, at least three identifiers (400) may be placed in the observation target area.
[0060] Next, the processor of the server (100) groups people crowded in a specific area based on the actual coordinates of the heads produced, determines whether the number of people placed in each group exceeds a preset number of people per unit area (reference value), and determines whether the area is in a normal or dangerous state, and outputs a warning message to the warning device (300) of the area determined to be in a dangerous state. In other words, the processor can determine the crowd density and dangerous state based on the actual coordinates of the heads.
[0061] At this time, the server (100) can compare the number of people per unit area set corresponding to the region corresponding to the calculated coordinates with the number of heads corresponding to the actual coordinates, and if the number of heads exceeds the number of people per unit area set above, determine the photographed area as being in a dangerous state.
[0062] Additionally, the above grouping process can be performed based on head detection information based on actual coordinates rather than head detection information based on an image captured by a camera (200).
[0063] Through this, the server (100) can obtain group information of people who are adjacent to each other by a preset number or more based on actual coordinates, and can calculate the number of people per unit area within the group to determine whether the density is at a dangerous level or not.
[0064] That is, the image processing method of the present invention can identify heads grouped together, such as parents holding a child or a couple in close contact, and thus, even if a large number of heads are detected, they may not be judged as being crowded. In other words, the processor of the server (100) can determine whether a group in a given area is crowded or not by considering various variables.
[0065] At this time, the result value after the grouping process may be expressed only in X and Y coordinates, assuming that the head detection information has the same Z value based on the actual coordinates.
[0066] Through this, the server (100) can determine whether people are gathered in a small area or are spaced apart from each other, even if the same number of people are included in the group.
[0067] Additionally, it can be used to determine how many high-risk groups are clustered in a small area, and it can also include a function to send an alert message to the administrator if there is a high-risk group.
[0068] Here, the grouping process can utilize a method such as K-means clustering, which requires prior input of the number of clusters, or a method such as affinity propagation, which automatically selects the number of clusters. However, different grouping methods may be utilized depending on various embodiments.
[0069] Meanwhile, the camera (200) may be installed at a predetermined height on the exterior wall of a building as shown in FIG. 6 to facilitate capturing the head of a human body, or may be implemented by being connected to a moving vehicle through a pole having a predetermined length as shown in FIG. 7.
[0070] Referring to FIG. 7, the result value of correction based on a structure in the form of parallel lines identified in an image captured by a camera (200) in consideration of the positional relationship between the vehicle body and the ground according to the number of people on board the vehicle, the boarding position, and the terrain on which it is driven can be utilized as an external parameter, and the result value of correction for the focal length of an image captured by a camera (200) in consideration of the length and flexibility of the pole can be utilized as an internal parameter.
[0071] This is to address the problem of errors in conversion to actual coordinates, as external parameters change due to changes in the positional relationship between the vehicle body and the ground, such as the number of passengers on board and their boarding positions. The present invention can utilize parallel line-shaped structures (e.g., lanes, parking lines) visible in the image to compensate for this.
[0072] For example, assuming that the external parameter information of the camera (200) is known in advance before the camera (200) is installed in the vehicle, the optical axis of the camera (200) and the ground are not perpendicular due to the terrain on which the vehicle is driving, and the camera image being captured is converted to a camera image in which the optical axis and the ground are perpendicular.
[0073] At this time, if the relationship between the camera (200) and the ground is normal, a parallel straight line in the real world also forms a parallel straight line in the transformed image coordinate system, whereas if a heavy object is loaded on the rear of the vehicle and the rear suspension of the vehicle becomes shorter than the front suspension, the camera (200) tilts, or the pole moves due to elasticity, so that the angular displacement value increases by the length of the camera (200) on the ground, causing more severe distortion.
[0074] Additionally, when the vehicle is going up a hill or down a slope, the distortion caused by the vehicle's tilt and the ground slope can be combined to create even more distortion.
[0075] To solve this problem, the present invention can optimize external parameters by setting the objective function of the non-linear least squares method to the concurrency of two or more straight lines in the transformed image coordinate system.
[0076] According to another embodiment, when a camera (200) is installed in a vehicle, the system of the present invention can be linked with the GPS (global positioning system) information built into the vehicle to record changes in crowd density according to location and time, and when the crowd becomes too dense and the danger level increases, a report message can be automatically sent to a nearby police station, fire station, etc. to help the area quickly take necessary measures.
[0077] In addition, in cases where the head of the human body is positioned higher than the vehicle due to the nature of the vehicle, or the angle of view of the image captured by the camera (200) is limited beyond a preset range due to an obstacle, as shown in FIG. 7, the server (100) automatically changes the length of the pole from the minimum length to the maximum length and performs the capture, thereby preventing problems in the collection of the image through the camera (200).
[0078]
[0079] Below, we will once again examine the method of analyzing crowd density in camera images using artificial intelligence with reference to Figure 8.
[0080] First, the server (100) receives an image of the observation target area from a camera (200) installed to capture the observation target area (S101).
[0081] Next, the server (100) detects the head of a human body in the received image using the head detection model learned according to the received image and a preset algorithm, and calculates the camera coordinates of the person corresponding to the detected head based on the pixel data constituting the image (S102).
[0082] Continuing, the server (100) calculates the detected actual coordinates of the person corresponding to the head (S103)
[0083] Next, the server (100) groups people concentrated in a specific area based on actual coordinates, determines whether the number of people placed in each group exceeds the preset number of people per unit area, determines a normal state or a dangerous state, and generates a warning message to a warning device (300) in an area determined to be a dangerous state (S104).
[0084]
[0085] Meanwhile, the components of the aforementioned embodiments can be easily understood from a process perspective. That is, each component can be understood as a separate process. Furthermore, the processes of the aforementioned embodiments can be easily understood from the perspective of the device components.
[0086] In addition, the technical contents described above may be implemented in the form of program commands that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program commands, data files, data structures, etc., alone or in combination. The program commands recorded on the medium may be those specially designed and configured for the embodiments or may be known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program commands, such as ROMs, RAMs, and flash memories. Examples of program commands include not only machine language codes generated by a compiler, but also high-level language codes that can be executed by a computer using an interpreter, etc. The hardware devices may be configured to operate as one or more software modules to perform the operations of the embodiments, and vice versa.
[0087] The above-described embodiments of the present invention are disclosed for the purpose of illustration, and those skilled in the art with common knowledge of the present invention will be able to make various modifications, changes, and additions within the spirit and scope of the present invention, and such modifications, changes, and additions should be considered to fall within the scope of the following patent claims.
Claims
1. A communication unit that receives an image from an image acquisition device; and A computing device characterized by including a processor that detects a head of a person in the received image, calculates camera coordinates of the detected head based on pixel data constituting the image, and calculates actual coordinates of the person corresponding to the head using the calculated camera coordinates.
2. In the first paragraph, the image acquisition device captures the observation target area to acquire the image, and detects the head of the person using a head detection model learned by machine learning, wherein at least one identifier exists within the observation target area. A computing device characterized in that the processor detects camera coordinates of the identifier, compares the detected camera coordinates with a pre-stored identifier to calculate actual coordinates of the identifier, obtains a relationship between the camera coordinates and the actual coordinates, and inputs the camera coordinates of the detected head into the relationship to calculate actual coordinates of the detected head.
3. In the second paragraph, the processor calculates the relationship by calculating an external parameter indicating the distance between the image acquisition device and the ground and an internal parameter indicating the focal distance between the image acquisition device and the image. A computing device, characterized in that the external parameters or the internal parameters are composed of a rotation matrix and a translation matrix.
4. In the third paragraph, when the image acquisition device is installed on a pole of the vehicle, the processor uses an external parameter reflecting the positional relationship between the vehicle and the ground and uses the focal length corrected by considering the length or flexibility of the pole as the internal parameter. A computing device.
5. In the first paragraph, the image acquisition device captures the observation target area and acquires the image, and at least one identifier exists within the observation target area. A computing device characterized in that the above identifier is composed of grids having two different colors and having a repeating pattern.
6. In the first paragraph, the processor displays a head identifier on the detected head, A computing device, characterized in that the size of the head identifier varies depending on the distance between the subject and the image acquisition device.
7. A computing device according to claim 1, characterized in that the processor groups people concentrated in a specific area based on the actual coordinates of the calculated heads, determines the area as being in a dangerous state if the number of people arranged in each group exceeds a reference value, and outputs a warning message to a warning device in the area if the area is determined to be in a dangerous state.
8. A computing device according to claim 1, characterized in that the processor groups people concentrated in a specific area based on the actual coordinates of the calculated heads, determines the number of groups with a risk level exceeding a reference value, and outputs a notification message to an administrator if there is a group with a risk level exceeding the reference value.
9. A communication unit for receiving an image from an image acquisition device; and A computing device characterized by including a processor that detects a human head in the received image, groups people concentrated in a specific area based on the detected head, and outputs a warning message to the outside when the number of people in a specific group exceeds a reference value.
10. A computing device according to claim 9, characterized in that the processor groups people concentrated in a specific area based on the actual coordinates of the head produced.
11. A computing device according to claim 10, wherein the processor obtains a relationship between the camera coordinates of the identifier in the image and the actual coordinates, and inputs the camera coordinates of the detected head into the relationship to calculate the actual coordinates of the detected head.
12. A computing device according to claim 9, wherein, when the image acquisition device is installed in a vehicle, the processor detects changes in crowd density according to the location and time of the vehicle, and transmits a reporting message to a preset device when the crowd density exceeds a reference value.
13. A step of detecting a human head in an image received from an image acquisition device; A step of obtaining the camera coordinates of the detected head; A step of obtaining the actual coordinates of the person corresponding to the head using the above camera coordinates; and A method for analyzing crowd density in an image processing system, characterized by including a step of detecting crowd density by grouping people densely gathered in a specific area based on the above-determined actual coordinates.
14. In paragraph 13, A method for analyzing crowd density in an image processing system, characterized in that it further includes a step of outputting a warning message externally when the number of people in a specific group exceeds a reference value.
15. In paragraph 13, the step of obtaining the actual coordinates of the person is: A step of obtaining a relationship between the camera coordinates and the actual coordinates of the identifier in the above image; and A method for analyzing crowd density in an image processing system, characterized by including a step of inputting the camera coordinates of the detected head into the above relational expression to calculate the actual coordinates of the detected head.
Citation Information
Patent Citations
Crowd state recognition device, method, and program
JP2021073606A
Crowd density calculation device, crowd density calculation method, and crowd density calculation program
JP6678835B2
Aptamer specifically binding to a nickel ion
KR102378495B1
Crowd density-based hazardous area automated alert system
KR102579542B1
KR20220155700A