Indoor personnel congestion assessment method and device

By obtaining the depth image of the monitoring area and performing grid division, combined with the head detection model, the problem of difficulty in accurately assessing indoor crowding in the existing technology is solved, and accurate assessment and safety monitoring of the degree of crowding in people are achieved.

CN120298971APending Publication Date: 2025-07-11SHARETRONIC DATA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510411248.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the prior art, it is difficult to accurately reflect whether the indoor personnel are crowded due to relying solely on two-dimensional images.

Method used

By obtaining the depth image of the monitoring area, meshing is performed based on the depth value, and using the head detection model to identify the head center points in the video frame, and counting the number of head center points in each grid for crowding evaluation.

Benefits of technology

It can accurately reflect the real three-dimensional spatial information of the monitoring area, improve the accuracy of personnel crowding assessment and ensure safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298971A_ABST
    Figure CN120298971A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of machine learning, and provides an indoor personnel congestion assessment method and device, and the method comprises the steps: obtaining a depth image of a monitoring region; performing grid division on the depth image based on the depth value, so that the size of a single grid corresponds to a preset area in the real world; monitoring the monitoring area and acquiring a video frame, identifying a head in the video frame based on a head detection model, and acquiring the position of the center point of the head; and mapping the positions of the head center points to the depth image, counting the number of the head center points of each grid in the depth image, and performing indoor personnel congestion assessment based on the number of the head center points of each grid. Thus, the real three-dimensional space information of the monitoring area is reflected through the depth image, so that the single grid in the depth image corresponds to the preset area in the real world, then the number of people in the preset area is counted, and the crowdedness degree of people can be accurately evaluated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of machine learning, and particularly relates to an indoor personnel crowding assessment method and device. Background Art

[0002] With the acceleration of the urbanization process and the increase in indoor public places, such as shopping malls, stations, airports, exhibition halls, and schools, the management and safety issues of indoor crowded scenarios have become particularly important. Crowded people not only affect the utilization efficiency of the venue but also may lead to potential safety hazards, such as stampede accidents and difficulties in emergency evacuation. Therefore, effective monitoring and analysis of indoor crowding situations are of great significance for improving venue management levels and ensuring personnel safety.

[0003] Traditional methods for detecting crowded people mostly rely on manual observation or simple sensor data, such as infrared sensors or people counters. These methods often have problems such as high costs, difficult deployment, and low data accuracy. In recent years, with the development of computer vision and deep learning technologies, through the combination of video surveillance systems and object detection algorithms, accurate analysis of indoor crowd density can be achieved, providing real-time and reliable data support for managers.

[0004] The existing technology mainly focuses on detecting human heads, and uses the ratio of the total pixel area of the detection frames of human heads to the pixel area of the overlapping region of the detection frames as an index to evaluate the degree of crowding. However, relying solely on two-dimensional images for judgment is difficult to accurately reflect whether there is crowding of people. Summary of the Invention

[0005] The purpose of the embodiments of the present application is to provide an indoor personnel crowding assessment method, aiming to solve the problem in the existing technology that it is difficult to accurately reflect whether there is crowding of people by relying solely on two-dimensional images for judgment.

[0006] The embodiments of the present application are implemented as follows. An indoor personnel crowding assessment method, the method includes:

[0007] Obtain a depth image of the monitoring area, where the pixel value of a pixel point of the depth image is the depth value of the corresponding point in the monitoring area;

[0008] Based on the depth value, perform grid division on the depth image so that the size of a single grid corresponds to a preset area in the real world;

[0009] Monitor the monitoring area and obtain video frames, identify human heads in the video frames based on a human head detection model, and obtain the positions of the center points of the human heads;

[0010] Map the position of the center point of the human head to the depth image, count the number of center points of the human head in each grid of the depth image, and perform indoor personnel crowding assessment based on the number of center points of the human head in each grid.

[0011] In some preferred embodiments of the present application, the method for obtaining the depth image of the monitoring area includes:

[0012] Synchronously collect the left-eye image and the right-eye image of the monitoring area through the left-eye camera and the right-eye camera of the binocular camera;

[0013] Based on the feature matching algorithm, calculate the position difference of each point in the monitoring area on the left-eye image and the right-eye image to obtain the disparity value of the point;

[0014] Based on the formula Calculate the depth value of each point in the monitoring area, where Z is the depth value, f is the focal length of the binocular camera, B is the baseline length of the binocular camera, and d is the disparity value;

[0015] Generate the depth image based on the depth values of each point in the monitoring area, and the pixel value of each pixel point in the depth image is the depth value of the corresponding point in the monitoring area.

[0016] In some preferred embodiments of the present application, the method for dividing the depth image into grids based on the depth value includes:

[0017] Based on the focal length of the binocular camera and the depth value of each pixel point in the depth image, calculate the actual physical area corresponding to each pixel point in the depth image;

[0018] Perform area accumulation along the horizontal and vertical directions of the depth image until the accumulated actual physical area reaches the preset area, and mark the accumulated area as a grid;

[0019] Traverse the entire depth image and divide the entire depth image into grids, where each grid corresponds to a preset area in the real world.

[0020] In some preferred embodiments of the present application, the actual physical area of a single pixel point is calculated by the following formula:

[0021]

[0022] Among them, S is the actual physical area corresponding to the pixel point, Z is the depth value of the current pixel point, and f is the focal length of the binocular camera.

[0023] In some preferred embodiments of the present application, before calculating the position difference of each point in the monitoring area on the left-eye image and the right-eye image based on the feature matching algorithm, it further includes:

[0024] Obtain the intrinsic matrix, distortion coefficients of the left-eye camera and the right-eye camera, as well as the rotation matrix and translation vector between the left-eye camera and the right-eye camera;

[0025] Based on the intrinsic matrix and distortion coefficients, perform undistortion correction on the left-eye image and the right-eye image to eliminate the non-linear distortion introduced by the camera lens;

[0026] Based on the rotation matrix and translation vector, perform epipolar rectification on the undistorted left-eye image and right-eye image, so that the corresponding pixel points of the left-eye image and the right-eye image are located on the same horizontal line.

[0027] In some preferred embodiments of the present application, before calculating the actual physical area corresponding to each pixel point, it further includes:

[0028] Perform median filtering on the depth image to smooth its depth value.

[0029] In some preferred embodiments of the present application, the training method of the human head detection model includes:

[0030] Obtain training images containing human heads, perform human head annotation on the training images to obtain a human head detection data set;

[0031] Based on the human head detection data set, generate a YOLO format data set configuration file containing data paths and class information, and a model training configuration file containing network structures and hyperparameters respectively;

[0032] Call the YOLO framework to load the data set configuration file and the model configuration file, calculate the loss function through forward propagation and update the network parameters through backpropagation, and iteratively train to obtain the final human head target detection model.

[0033] In some preferred embodiments of the present application, after performing human head annotation on the training images, it further includes:

[0034] Horizontally and vertically flip the training images to obtain enhanced images, and the training images and the enhanced images form the human head detection data set.

[0035] In some preferred embodiments of the present application, the method for evaluating indoor personnel congestion based on the number of human head center points in each grid includes:

[0036] Obtain the number of human head center points in each cell of the depth map;

[0037] If the number of human head center points in any grid is greater than a preset threshold, an alarm is issued.

[0038] Another object of the embodiments of the present application is an indoor personnel crowding assessment device based on object detection and binocular depth estimation. The device includes:

[0039] An image acquisition device, configured to acquire a depth image of a monitoring area, where the pixel value of the depth image is the depth value of the corresponding point in the monitoring area;

[0040] A grid division device, configured to perform grid division on the depth image based on the depth value, so that the size of a single grid corresponds to a preset area in the real world;

[0041] A head recognition device, configured to monitor the monitoring area and acquire video frames, recognize heads in the video frames based on a pre-trained head detection model, and acquire the positions of the head center points;

[0042] A crowding assessment device, configured to map the positions of the head center points to the depth image, count the number of head center points in each grid in the depth image, and perform indoor personnel crowding assessment based on the number of head center points.

[0043] An indoor personnel crowding assessment method provided by the embodiments of the present application includes: acquiring a depth image of a monitoring area, where the depth image contains depth value information of each point in the monitoring area; performing grid division on the depth image based on the depth value, so that the size of a single grid corresponds to a preset area in the real world; acquiring the positions of the head center points through a head detection model; mapping the positions of the head center points in the depth image; counting the number of head center points in the grids in the depth image; and performing indoor personnel crowding assessment through the number of head center points. In this way, the present application reflects the true three-dimensional space information of the monitoring area through the depth map, makes a single grid in the depth image correspond to a preset area in the real world, converts the area on the map into the actual area in the real world, and then counts the number of people in the preset area, so as to accurately evaluate the degree of personnel crowding. Description of the Drawings

[0044] Figure 1 It is a flowchart of an indoor personnel crowding assessment method provided by the embodiments of the present application;

[0045] Figure 2 It is a flowchart of a method for acquiring a depth image of a monitoring area provided by the embodiments of the present application;

[0046] Figure 3 It is a flowchart of a method for performing grid division on the depth image based on the depth value provided by the embodiments of the present application;

[0047] Figure 4 It is a flowchart of a method for correcting a left-eye image and a right-eye image provided by the embodiments of the present application;

[0048] Figure 5 Flow chart of the training method for the head detection model provided by the embodiment of the present application;

[0049] Figure 6 Structural block diagram of an indoor personnel crowding assessment device provided by the embodiment of the present application. Detailed implementation manners

[0050] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0051] It can be understood that the terms "first", "second", etc. used in the present application can be used in this article to describe various elements, but unless otherwise specified, these elements are not limited by these terms. These terms are only used to distinguish the first element from another element. For example, without departing from the scope of the present application, the first xx script can be called the second xx script, and similarly, the second xx script can be called the first xx script.

[0052] In the field of personnel crowding assessment, the existing technology mainly focuses on detecting human heads. Specifically, image data is collected using a surveillance camera, and a target detection algorithm is used to detect the heads of personnel. The crowding degree of the image acquisition area is evaluated based on the head detection results. The crowding degree calculation formula is the ratio of the total pixel area of all detection frames to the pixel area of the overlapping area of the detection frames. An alarm is given when the crowding degree of the crowd exceeds the set threshold. However, using the ratio of the total pixel area of the detection frames of human heads to the pixel area of the overlapping area of the detection frames as an index to evaluate the crowding degree has no depth information and cannot reflect the real physical space. Relying only on two-dimensional images for judgment is difficult to accurately reflect whether there is personnel crowding. To address the above problems, the present application proposes an indoor personnel crowding assessment method, which can improve the accuracy of the assessment. The solution of the present application will be described below.

[0053] As Figure 1 shown, in one embodiment, an indoor personnel crowding assessment method is proposed, which may specifically include the following steps:

[0054] Step S100: Obtain a depth image of the monitoring area, where the pixel value of the pixel point of the depth image is the depth value of the corresponding point in the monitoring area.

[0055] In this embodiment, the depth image can be obtained by a binocular camera. The pixel value of each pixel point of the depth image is the depth value of the corresponding point in the monitoring area. The depth value is specifically the distance between the corresponding point on the monitoring area and the camera.

[0056] Step S200: Based on the depth values, perform grid division on the depth image so that the size of a single grid corresponds to a preset area in the real world.

[0057] In this embodiment, the correspondence between the area on the graph and the actual area can be obtained through the depth values and the parameters of the binocular camera. In this way, grid division can be performed on the depth map so that the size of a single grid corresponds to a preset area in the real world. In some embodiments, the preset area is 1 m 2 .

[0058] Step S300: Monitor the monitoring area and obtain video frames. Based on the human head detection model, identify the human heads in the video frames and obtain the positions of the centers of the human heads.

[0059] In this embodiment, the monitoring area is monitored by a monitoring device, video frames are intercepted at regular intervals, the human heads in the video frames are identified through the human head detection model, and the positions of the centers of the human heads are obtained.

[0060] Step S400: Map the positions of the centers of the human heads to the depth image, count the number of centers of the human heads in each grid of the depth image, and perform an indoor crowding assessment based on the number of centers of the human heads in each grid. In some embodiments, the positions of each point in the monitoring area on the depth map and the video frames are the same. In this way, the positions of the centers of the human heads can be correspondingly mapped into the depth image.

[0061] In this embodiment, the depth value information of each point in the monitoring area is included in the obtained depth image of the monitoring area. Based on the depth values, grid division is performed on the depth image so that the size of a single grid corresponds to a preset area in the real world. The positions of the centers of the human heads are obtained through the human head detection model and mapped onto the depth image. The number of centers of the human heads in the grids of the depth image is counted, and an indoor crowding assessment is performed through the number of centers of the human heads. In this way, the real three-dimensional space information of the monitoring area can be reflected through the depth image. When performing grid division on the depth map, a single grid can correspond to a preset area in the real world. In this way, the area on the graph is converted into the actual area in the real world, and then the number of people within the preset area is counted, enabling an accurate assessment of the degree of crowding.

[0062] Specifically, in some embodiments, it is possible to start from the lower left corner of the depth map and accumulate the area simultaneously upward and to the right. After accumulating to a preset area, a first square grid is obtained. Then, based on the right side length of the first square grid, continue to accumulate the area and divide the grid to the right. If the last grid is less than the preset area, it is directly cropped. In this way, the division of the first layer of grids is completed. For the second layer, start from the upper left corner of the first square grid and accumulate the area simultaneously upward and to the right. After accumulating to the preset area, a second square grid is obtained. Then, based on the right side length of the second square grid, continue to accumulate the area and divide the grid to the right. Repeat the above process until the entire depth image is divided. In some other embodiments, a dynamic grid division technique can be used to divide the depth image, which can reduce grid fragmentation.

[0063] As Figure 2 shown, in some embodiments of the present application, the method for obtaining the depth image of the monitoring area includes:

[0064] Step S120: Synchronously collect the left-eye image and the right-eye image of the monitoring area through the left-eye camera and the right-eye camera of the binocular camera.

[0065] In this embodiment, the binocular camera has two horizontally arranged cameras, namely the left-eye camera and the right-eye camera. There is a certain distance between the two cameras, and the distance between the two cameras is specifically usually referred to as the baseline length. The imaging positions of the same object in the left-eye camera and the right-eye camera are different. For example, the position difference of nearby objects in the left and right images is large, and the difference of distant objects is small.

[0066] Step S140: Calculate the position difference of each point in the monitoring area on the left-eye image and the right-eye image based on the feature matching algorithm to obtain the disparity value of the point.

[0067] In this embodiment, corresponding points are searched in the left-eye image and the right-eye image through the feature matching algorithm, and their horizontal position differences are calculated to obtain the disparity value. In some embodiments, the feature matching algorithm can use methods such as semi-global matching, block matching, or deep learning models. The specific algorithm is an existing mature technology and will not be elaborated here.

[0068] Step S160: Based on the formula calculate the depth value of each point in the monitoring area, where Z is the depth value, f is the focal length of the binocular camera, B is the baseline length of the binocular camera, and d is the disparity value. In some embodiments, the focal lengths of the left-eye camera and the right-eye camera in the binocular camera are the same, and the baseline length of the binocular camera is the distance between the left-eye camera and the right-eye camera.

[0069] Step S180: Generate the depth image based on the depth values of each point in the monitoring area. The pixel value of each pixel point in the depth image is the depth value of the corresponding point in the monitoring area.

[0070] In this embodiment, the left camera and the right camera of the binocular camera are used to capture the monitoring area to obtain a left-eye image and a right-eye image. Subsequently, a feature matching algorithm is used to search for corresponding pixel points in the left-eye image and the right-eye image, calculate the disparity value, and substitute the disparity value into the formula , calculate the obtained depth value, and generate a depth image with the depth value as the pixel value of the corresponding pixel point. In this way, a depth image can be obtained through the binocular camera, which is simple and reliable.

[0071] As Figure 3 shown, in some embodiments of the present application, the method for grid-dividing the depth image based on the depth value includes:

[0072] Step S220: Calculate the actual physical area corresponding to each pixel point in the depth image based on the focal length of the binocular camera and the depth value of each pixel point in the depth image.

[0073] In some embodiments, the actual physical area of a single pixel point is calculated by the following formula:

[0074]

[0075] where S is the actual physical area corresponding to the pixel point, Z is the depth value of the current pixel point, and f is the focal length of the binocular camera.

[0076] Step S240: Perform area accumulation along the horizontal and vertical directions of the depth image until the accumulated actual physical area reaches a preset area, and mark the accumulated area as a grid.

[0077] Step S260: Traverse the entire depth image and perform grid division on the entire depth image, where each grid corresponds to a preset area in the real world.

[0078] In this embodiment, through the depth value and the focal length of the binocular camera, the actual physical area corresponding to a single pixel point is calculated, area accumulation is performed along the horizontal and vertical directions of the depth image, and when the accumulated area reaches the preset area, it is marked as a grid and the border is generated. By traversing the entire depth image, the entire image can be divided.

[0079] In some embodiments of the present application, before calculating the position difference of each point in the monitoring area on the left-eye image and the right-eye image based on the feature matching algorithm, it further includes:

[0080] Step S112: Obtain the internal parameter matrix, distortion coefficient of the left camera and the right camera, and the rotation matrix and translation vector between the left camera and the right camera.

[0081] As Figure 4 shown, in this embodiment, the left-eye camera and the right-eye camera are respectively calibrated for single object, and the internal parameter matrix and the distortion coefficient can be obtained. By performing binocular calibration on the binocular camera, the rotation matrix and the translation vector between the left-eye camera and the right-eye camera can be obtained. For specific single object calibration and binocular calibration, reference can be made to the existing calibration method based on checkerboard, which will not be elaborated here.

[0082] Step S114: Based on the internal parameter matrix and the distortion coefficient, perform distortion correction on the left-eye image and the right-eye image to eliminate the non-linear distortion introduced by the camera lens.

[0083] In this embodiment, distortion correction is a process of using the internal parameter matrix and the distortion coefficient obtained by single object calibration to restore the deformed area caused by lens distortion in the original image to a straight line and a plane, making the image closer to the real world, facilitating subsequent disparity value calculation, and being able to reduce the error of disparity value calculation.

[0084] Step S116: Based on the rotation matrix and the translation vector, perform epipolar correction on the distortion-corrected left-eye image and right-eye image, so that the corresponding pixel points of the left-eye image and the right-eye image are located on the same horizontal line. In this way, when performing feature matching on the left-eye image and the right-eye image, only the pixel points on the same horizontal line need to be matched, reducing the workload and improving the matching speed.

[0085] In this embodiment, after the left-eye image and the right-eye image are subjected to distortion correction and epipolar correction, the non-linear distortion introduced by the camera lens is eliminated, and the corresponding pixel points of the left-eye image and the right-eye image are located on the same horizontal line, improving the acquisition speed and accuracy of the disparity value.

[0086] In some embodiments of the present application, before calculating the actual physical area corresponding to each pixel point, it further includes: performing median filtering on the depth image to smooth its depth value.

[0087] In this embodiment, the depth map often generates outlier noise due to environmental factors or defects of the camera itself. Through the effect of median filtering, isolated outliers are effectively filtered out. The smoothed depth value is more stable in the local area, reducing the drastic fluctuation of the depth value within the grid, and ensuring that a single grid more accurately corresponds to 1m in the real world 2 .

[0088] In some embodiments of the present application, the training method of the human head detection model includes:

[0089] Step S320: Obtain training images containing human heads, perform head annotation on the training images, and obtain a head detection dataset. In this embodiment, a labeling tool LabelImg can be used to perform bounding box annotation on human heads to generate a labeling file in a standard format.

[0090] Step S340: Based on the head detection dataset, generate a YOLO format dataset configuration file containing data paths and class information, and a model training configuration file containing network structures and hyperparameters respectively.

[0091] In this embodiment, the data configuration file specifies the image and label paths of the training set, validation set, and test set, ensuring that the model can correctly load the data and clarify the class names and quantities of detection targets. By uniformly managing all data paths and class labels through the dataset configuration file, training failures caused by path errors or label confusion can be avoided. The model training configuration file is used to define the model structure, select a pre-trained model, customize network layers, number of channels, anchor parameters, and configure hyperparameters. By reasonably setting the training configuration file, the accuracy and efficiency of the head detection model can be significantly improved.

[0092] Step S360: Invoke the YOLO framework to load the dataset configuration file and the model configuration file, calculate the loss function through forward propagation and update the network parameters through backpropagation, and iteratively train to obtain the final head target detection model.

[0093] In this embodiment, using the YOLO framework to construct a head detection model can improve the accuracy of detection while taking into account flexible deployment and development efficiency.

[0094] In some embodiments of the present application, the training images are horizontally and vertically flipped to obtain enhanced images, and the training images and the enhanced images form the head detection dataset.

[0095] In this embodiment, by horizontally and vertically flipping the training images to obtain enhanced images, the training samples can be expanded, and using the flipped normal images for training the model can improve the robustness of the model.

[0096] In some embodiments of the present application, a method for evaluating indoor personnel crowding based on the number of human head center points in each grid includes: obtaining the number of human head center points in each cell of the depth map; if the number of human head center points in any grid is greater than a preset threshold, an alarm is issued.

[0097] In this embodiment, the monitoring device captures a video frame at regular intervals, identifies the position of the center point of the human head through a human head detection model and maps it onto the depth map, and counts the number of center points of human heads in each grid of the depth map. When the number exceeds a preset threshold, an alarm is triggered to remind the staff to limit the flow. In this way, the congestion situation of the personnel can be monitored at all times to ensure the safety of the personnel. In some embodiments, the preset threshold is 2 and the preset area is 1m 2 。

[0098] The embodiment of the present application also provides an indoor personnel congestion assessment device, and the device includes:

[0099] An image acquisition device 100 for acquiring a depth image of a monitoring area, where the pixel value of the depth image is the depth value of the corresponding point in the monitoring area.

[0100] A grid division device 200 for dividing the depth image based on the depth value, so that the size of a single grid corresponds to a preset area in the real world.

[0101] A human head recognition device 300 monitors the monitoring area and acquires video frames, identifies the human heads in the video frames based on a pre-trained human head detection model, and acquires the positions of the center points of the human heads.

[0102] A congestion assessment device 400 for mapping the positions of the center points of the human heads to the depth image, counting the number of center points of the human heads in each grid of the depth image, and performing an indoor personnel congestion assessment based on the number of center points of the human heads.

[0103] In this embodiment, a depth image of the monitoring area is acquired through the image acquisition device 100. The depth image contains the depth value information of each point in the monitoring area. The grid division device 200 divides the depth image based on the depth value, so that the size of a single grid corresponds to a preset area in the real world. The human head recognition device 300 uses the human head detection model to acquire the positions of the center points of the human heads. The congestion assessment device 400 maps the positions of the center points of the human heads in the depth image and counts the number of center points of the human heads in the grids of the depth image, and performs an indoor personnel congestion assessment through the number of center points of the human heads. In this way, the real three-dimensional space information of the monitoring area can be reflected through the depth image. When dividing the grid of the depth map, it can be ensured that a single grid corresponds to a preset area in the real world, convert the area on the map into the actual area in the real world, and then count the number of people in the preset area, so as to accurately evaluate the degree of personnel congestion.

[0104] It should be understood that although the steps in the flowcharts of the embodiments of the present application are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order limit for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in each embodiment may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0105] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0106] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0107] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

[0108] The foregoing is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An indoor personnel congestion assessment method, characterized in that, The method includes: Obtain a depth image of the monitoring area, where the pixel value of each pixel point in the depth image is the depth value of the corresponding point in the monitoring area; Based on the depth value, perform grid division on the depth image so that the size of a single grid corresponds to a preset area in the real world; Monitor the monitoring area and obtain a video frame. Based on a head detection model, identify the heads in the video frame and obtain the positions of the head center points; Map the positions of the head center points to the depth image, count the number of head center points in each grid in the depth image, and perform indoor personnel crowding assessment based on the number of head center points in each grid.

2. The indoor personnel crowding evaluation method according to claim 1, wherein, The method for obtaining a depth image of the monitoring area includes: Synchronously collect a left-eye image and a right-eye image of the monitoring area through the left-eye camera and the right-eye camera of a binocular camera; Based on a feature matching algorithm, calculate the position difference of each point in the monitoring area on the left-eye image and the right-eye image to obtain the disparity value of the point; Based on the formula Calculate the depth values of each point in the monitoring area, where Z is the depth value, f is the focal length of the binocular camera, B is the baseline length of the binocular camera, and d is the disparity value; Generate the depth image based on the depth values of each point in the monitoring area, and the pixel value of each pixel point in the depth image is the depth value of the corresponding point in the monitoring area.

3. The indoor personnel crowding assessment method according to claim 2, wherein The method for performing grid division on the depth image based on the depth value includes: Based on the focal length of the binocular camera and the depth value of each pixel point in the depth image, calculate the actual physical area corresponding to each pixel point in the depth image; Perform area accumulation along the horizontal and vertical directions of the depth image until the accumulated actual physical area reaches the preset area, and mark the accumulated area as a grid; Traverse the entire depth image and perform grid division on the entire depth image, where each grid corresponds to a preset area in the real world.

4. The indoor personnel congestion assessment method according to claim 3, wherein The actual physical area of a single pixel point is calculated by the following formula: where S is the actual physical area corresponding to the pixel point, Z is the depth value of the current pixel point, and f is the focal length of the binocular camera.

5. The indoor personnel crowding assessment method according to claim 2, characterized in that Before calculating the position difference of each point in the monitoring area on the left-eye image and the right-eye image based on the feature matching algorithm, it further includes: Obtain the internal parameter matrix, distortion coefficient of the left-eye camera and the right-eye camera, and the rotation matrix and translation vector between the left-eye camera and the right-eye camera; Based on the internal parameter matrix and distortion coefficient, perform distortion correction on the left-eye image and the right-eye image to eliminate the non-linear distortion introduced by the camera lens; Based on the rotation matrix and translation vector, perform epipolar correction on the undistorted left-eye image and right-eye image so that the corresponding pixel points of the left-eye image and the right-eye image are located on the same horizontal line.

6. The indoor personnel crowding evaluation method according to claim 3, characterized in that Before calculating the actual physical area corresponding to each pixel point, it further includes: Perform median filtering on the depth image to smooth its depth value.

7. The indoor personnel crowding assessment method according to claim 1, characterized in that, The training method of the head detection model includes: Obtain training images containing heads, perform head annotation on the training images to obtain a head detection data set; Based on the head detection data set, generate a YOLO format data set configuration file containing data paths and class information, and a model training configuration file containing network structures and hyperparameters respectively; Call the YOLO framework to load the dataset configuration file and the model configuration file, calculate the loss function through forward propagation, and update the network parameters through backpropagation. Iteratively train to obtain the final head target detection model.

8. The indoor personnel crowding assessment method according to claim 7, wherein After head annotation of the training images, it further includes: Horizontally and vertically flip the training images to obtain enhanced images. The training images and the enhanced images form the head detection dataset.

9. The indoor personnel congestion assessment method according to claim 1, characterized in that The method for indoor personnel crowding assessment based on the number of head center points in each grid includes: Obtain the number of head center points in each cell of the depth map; If the number of head center points in any grid is greater than the preset threshold, an alarm is issued.

10. An indoor personnel crowding assessment device, characterized in that, The device includes: An image acquisition device for acquiring a depth image of a monitoring area, where the pixel value of the depth image is the depth value of the corresponding point in the monitoring area; A grid division device for dividing the depth image based on the depth value so that the size of a single grid corresponds to a preset area in the real world; A head recognition device for monitoring the monitoring area and acquiring video frames, recognizing the heads in the video frames based on a pre-trained head detection model, and obtaining the positions of the head center points; A crowding assessment device for mapping the positions of the head center points to the depth image, counting the number of head center points in each grid of the depth image, and performing indoor personnel crowding assessment based on the number of head center points.