Regional information estimation method and system and non-transient computer readable storage medium
Through the coordinated work of multiple monitoring devices and processing devices, two-dimensional and three-dimensional density maps are generated, which solves the problems of calculation errors and high burdens in large-scale population counting, and realizes accurate population counting and computing burden sharing in crowded and masked environments.
Patent Information
- Application Number
- CN202410728726.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-01-17
- Filing Date
- 2024-06-06
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art is susceptible to crowding and shading when counting crowds in large areas, resulting in errors in calculation results, and multi-field-based methods require frequent calibration and high computational burden.
Multiple monitoring devices are used to collect images from different fields of view, and a two-dimensional density map is generated through a two-dimensional neural network model. A three-dimensional density map is generated by combining a three-dimensional neural network model to calculate the number of target objects. The monitoring device is movable without additional calibration. The processing device fuses the output and shares the operation burden.
Accurately counting large-scale populations in crowded and shading reduces the calibration needs of monitoring devices and reduces the computing burden.
Smart Images

Figure CN120339370A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a method and a system, and particularly to a method and a system for estimating regional information. Background Art
[0002] In the field of crowd counting, some related technologies perform crowd counting based on a single view. However, these related technologies based on a single view are prone to incorrect calculation results due to crowding and / or occlusion, and thus are not suitable for estimating the number of people in a large area. Some related technologies perform crowd counting based on multiple views, and are generally implemented by a system including an operation unit and multiple cameras. However, these related technologies based on multiple views must calibrate the cameras every time the camera position is changed, which is not convenient for users. In addition, the operation unit needs to fuse the outputs of multiple cameras to further obtain the final calculation result, which brings a heavy operation burden to the operation unit. Therefore, it is necessary to propose a new method to perform crowd counting. Summary of the Invention
[0003] One aspect of the present disclosure is a method for estimating regional information. The method for estimating regional information is applicable to a regional information estimation system including a processing device and multiple monitoring devices, and includes: capturing multiple images of a region from different views through the multiple monitoring devices; generating multiple two-dimensional density maps of at least one target object in the region through the multiple monitoring devices according to the multiple images; generating a three-dimensional density map by the processing device according to the multiple two-dimensional density maps; and calculating a quantity of the at least one target object by the processing device according to the three-dimensional density map.
[0004] In some embodiments, generating the multiple two-dimensional density maps of the at least one target object according to the multiple images includes: transforming the multiple images into the multiple two-dimensional density maps by the multiple monitoring devices using a two-dimensional neural network model.
[0005] In some embodiments, the two-dimensional neural network model is a convolutional neural network.
[0006] In some embodiments, the method for estimating regional information further includes: providing, by the multiple monitoring devices, multiple image capture data corresponding to the multiple images to the processing device.
[0007] In some embodiments, when at least one of the multiple monitoring devices is moved, the method for estimating regional information further includes: calculating, by the at least one of the multiple monitoring devices, at least one device pose information using a vision-based positioning technique, so as to generate at least one of the multiple image capture data.
[0008] In some embodiments, the regional information estimation method further includes: accessing, by the plurality of monitoring devices, a plurality of camera parameter information of a plurality of cameras of the plurality of monitoring devices as the plurality of image capture data.
[0009] In some embodiments, generating the three-dimensional density map based on the plurality of two-dimensional density maps includes: projecting the plurality of two-dimensional density maps based on a plurality of image capture data to generate an aggregation volume model; and generating the three-dimensional density map based on the aggregation volume model.
[0010] In some embodiments, projecting the plurality of two-dimensional density maps based on the plurality of image capture data to generate the aggregation volume model includes: calculating, based on the plurality of image capture data, a position of at least one characteristic pixel point of the plurality of two-dimensional density maps within the aggregation volume model, thereby forming at least one voxel point of the aggregation volume model.
[0011] In some embodiments, generating the three-dimensional density map based on the aggregation volume model includes: using a three-dimensional neural network model to transform the aggregation volume model into the three-dimensional density map.
[0012] In some embodiments, the three-dimensional neural network model is a convolutional neural network.
[0013] Another aspect of the present disclosure is a regional information estimation system. The regional information estimation system includes a plurality of monitoring devices and a processing device. The plurality of monitoring devices are configured to be disposed in a region, capture a plurality of images of the region from different perspectives, and generate a plurality of two-dimensional density maps of at least one target object within the region based on the plurality of images. The processing device is coupled to the plurality of monitoring devices, configured to generate a three-dimensional density map based on the plurality of two-dimensional density maps, and calculate a quantity of the at least one target object based on the three-dimensional density map.
[0014] In some embodiments, each of the plurality of monitoring devices includes a camera and a processor. The camera is configured to capture a corresponding one of the plurality of images. The processor is coupled to the camera and configured to transform the corresponding one of the plurality of images into a corresponding one of the plurality of two-dimensional density maps using a two-dimensional neural network model.
[0015] In some embodiments, the plurality of monitoring devices are configured to provide a plurality of image capture data corresponding to the plurality of images to the processing device. Each of the plurality of monitoring devices includes a camera, a memory, and a processor. The camera is configured to capture a corresponding one of the plurality of images. The memory is configured to store camera parameter information of the camera, where the camera parameter information includes an intrinsic camera parameter, an extrinsic camera parameter, and a distortion coefficient. The processor is coupled to the camera and the memory and is configured to access the camera parameter information as a corresponding one of the plurality of image capture data.
[0016] Another aspect of the present disclosure is a non-transitory computer-readable storage medium having a computer program for executing a region information estimation method, where the region information estimation method is applicable to a region information estimation system including a processing device and a plurality of monitoring devices, and includes: capturing, by the plurality of monitoring devices, a plurality of images of a region from different fields of view; generating, by the plurality of monitoring devices, a plurality of two-dimensional density maps of at least one target object within the region based on the plurality of images; generating, by the processing device, a three-dimensional density map based on the plurality of two-dimensional density maps; and calculating, by the processing device, a quantity of the at least one target object based on the three-dimensional density map.
[0017] In summary, through the monitoring devices capable of sensing their own postures in a region, the region information estimation system and the region information estimation method of the present disclosure can not only overcome the problems generated in crowded and / or blocked situations to perform crowd counting in a large-scale region, but also allow the monitoring devices to move in the region without additional calibration. In addition, by the monitoring devices predicting two-dimensional density maps based on the images of the region and the processing device fusing the outputs of the monitoring devices to generate a three-dimensional density map and calculating the quantity of the target object, the region information estimation system and the region information estimation method of the present disclosure have advantages such as sharing the heavy computing burden. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 A schematic diagram of a region information estimation system disposed in a region according to some embodiments of the present disclosure.
[0019] Figure 2 A block diagram of a region information estimation system according to some embodiments of the present disclosure.
[0020] Figure 3 A flowchart of a region information estimation method according to some embodiments of the present disclosure.
[0021] Figure 4Schematic diagrams of an image, a two-dimensional density map, an aggregated volume model, and a three-dimensional density map according to some embodiments of the present disclosure.
[0022] Symbol Explanation:
[0023] 10[1], 10[2]~10[N]: Monitoring device
[0024] 20: Processing device
[0025] 100: Region information estimation system
[0026] 101: Processor
[0027] 103: Camera
[0028] 105: Sensor
[0029] 107: Memory
[0030] 110, 110[1]~110[N]: Two-dimensional neural network model
[0031] 210: Three-dimensional neural network model
[0032] 300: Region information estimation method
[0033] 2DM1~2DMN: Two-dimensional density map
[0034] 3DM: Three-dimensional density map
[0035] A1: Region
[0036] B1, B2~BM: Target object
[0037] DC1~DCN: Image capture data
[0038] IB11, IB12, IBN1, IBN2: Target part
[0039] IMG1~IMGN: Image
[0040] P1: Camera parameter information
[0041] PL11, PL12, PLN1, PLN2: Feature pixel points
[0042] POS1: Device attitude information
[0043] S301~S304: Operations
[0044] T1, T2~TN: Trajectory
[0045] VL1, VL1’, VL2, VL3: Voxel points
[0046] VM: Aggregate Volume Model Detailed Implementation Manner
[0047] The following examples are provided in conjunction with the accompanying drawings for detailed description. However, the specific examples described are only used to explain this case and do not limit this case. The description of the structure and operation is not used to limit the execution order. Any structure formed by recombination of components and the resulting device with equivalent functions are all within the scope covered by this disclosure content.
[0048] Regarding the "coupled" or "connected" used in this article, it can refer to two or more components making direct physical or electrical contact with each other, or making indirect physical or electrical contact with each other, and can also refer to two or more components operating or acting on each other.
[0049] Please refer to Figure 1 , Figure 1 It is a schematic diagram of a regional information estimation system 100 shown in some embodiments according to this disclosure content. In some embodiments, the regional information estimation system 100 includes a plurality of monitoring devices 10[1] to 10[N] and a processing device 20, and is used to obtain the regional information of a region A1. It should be understood that N can be an integer greater than 1. In some practical applications, the regional information of region A1 can be the traffic flow of a road, the population density of a space, the number of people waiting for entertainment equipment, etc.
[0050] In Figure 1 's embodiment, the regional information estimation system 100 is used to obtain the quantities of a plurality of target objects B1 and B2 to BM within the region A1, where the plurality of target objects B1 to BM can each be pedestrians, people, etc. It should be understood that M can be an integer greater than 1. In order to obtain the quantities of the plurality of target objects B1 to BM, as Figure 1 shown, the monitoring devices 10[1] to 10[N] are arranged to be evenly distributed within the region A1. It should be noted that when arranged within the region A1, each of the monitoring devices 10[1] to 10[N] is movable. For example, in Figure 1 , a trajectory T1 represents the moving path of the monitoring device 10[1], a trajectory T2 represents the moving path of the monitoring device 10[2], and a trajectory TN represents the moving path of the monitoring device 10[N]. The processing device 20 can be arranged within the region A1 or in another region away from the region A1 (not shown in the figure). The processing device 20 is electrically and / or communicatively coupled to the monitoring devices 10[1] to 10[N]. In addition, the monitoring devices 10[1] to 10[N] are communicatively coupled to each other.
[0051] With the above settings, the monitoring devices 10[1] to 10[N] can take pictures from different perspectives within the area A1, and can transmit (processed or unprocessed) signals, data, and / or information to the processing device 20, enabling the processing device 20 to calculate the number of target objects B1 to BM. The operations of the monitoring devices 10[1] to 10[N] and the processing device 20 will be described in detail later. Next, first in conjunction with Figure 2 Describe in detail the structures of the monitoring devices 10[1] to 10[N] and the processing device 20.
[0052] Please refer to Figure 2 , Figure 2 FIG. 100 is a block diagram of an area information estimation system 100 according to some embodiments of the present disclosure. It should be understood that the monitoring devices 10[1] to 10[N] may have the same structure. Therefore, the structure of the monitoring devices 10[1] to 10[N] will be described by taking the monitoring device 10[1] as an example. As Figure 2 shown, the monitoring device 10[1] includes a processor 101, a camera 103, a sensor 105, and a memory 107. The processor 101 is electrically coupled to the camera 103, the sensor 105, and the memory 107.
[0053] The camera 103 is used to record and convert the optical signal from the area A1 into an electrical signal, and can be implemented through at least one lens unit, a photosensitive element (e.g., an image sensor such as a complementary metal-oxide semiconductor, a charge-coupled device, etc.), and an image processor.
[0054] The sensor 105 is used to generate and provide sensed data. Specifically, the sensor 105 may include a tracking camera and / or at least one inertial measurement unit, and the at least one inertial measurement unit can be implemented through an accelerometer, a magnetometer, a gyroscope, etc. In some embodiments, the sensed data can be used as auxiliary information for the calculation of the processor 101, thereby improving the computing efficiency of the processor 101. It should be understood that the sensor 105 is an optional component.
[0055] The memory 107 is used to store the signals, data, and / or information required for the operation of the monitoring device 10[1]. For example, the memory 107 can store the camera parameter information P1 of the camera 103, the sensed data sensed by the sensor 105, etc. The memory 107 can be implemented through at least one volatile memory unit, at least one non-volatile memory unit, or both.
[0056] The processor 101 is configured to process signals, data, and / or information required for the operation of the monitoring device 10[1]. In some embodiments, the processor 101 may use at least one vision-based positioning technology (e.g., Simultaneous Localization and Mapping technology) to calculate the pose of the monitoring device 10[1] based on the image data generated by the camera 103 and / or the tracking camera in the sensor 105. Specifically, the pose calculated by the processor 101 may indicate the six degrees of freedom of the monitoring device 10[1]. Further, the processor 101 may also use the motion data generated by at least one inertial measurement unit in the sensor 105 to assist in calculating the pose of the monitoring device 10[1], thereby improving the accuracy of the pose of the monitoring device 10[1]. In some further embodiments, as Figure 2 shown, the processor 101 utilizes a two-dimensional neural network model 110 for processing. Specifically, the two-dimensional neural network model 110 may be a convolutional neural network (e.g., network for Congested Scene Recognition (CSRNet), multi-column convolutional neural network (MCNN), deep convolutional neural networks for cross-scene crowd counting, etc.) that has been well-trained to perform at least one specific task such as two-dimensional image transformation (to be described later). The processor 101 may be implemented by a central processing unit, a graphics processing unit, an application-specific integrated circuit, a microprocessor, a system-on-chip, or other suitable processing units.
[0057] The processing device 20 is configured to process signals, data, and / or information transmitted from the monitoring device 10[1]. In some further embodiments, as Figure 2As shown, the processing device 20 uses a three-dimensional neural network model 210 for processing. Specifically, the three-dimensional neural network model 210 can be a type of convolutional neural network (e.g., fully convolutional neural networks for volumetric image segmentation (V-Net), Learning on Compressed Output (LoCO), U-shaped convolutional neural network transformers (UNETR), etc.) that has been well-trained to perform at least one specific task such as three-dimensional model transformation (to be described later). The processing device 20 can be implemented by a desktop computer, a laptop computer, a server, a tablet computer, a mobile phone, or other suitable computing devices.
[0058] Figure 2 The operations of each element in Figures 3 - 4 will be described in detail below. Please refer to Figure 3 , Figure 3 FIG. 300 is a flowchart of a region information estimation method 300 according to some embodiments of the present disclosure. The region information estimation method 300 is applicable to Figures 1 - 2 the region information estimation system 100. In some embodiments, as Figure 3 shown, the region information estimation method 300 includes a plurality of operations S301 to S304.
[0059] In operation S301, the monitoring devices 10[1] to 10[N] capture a plurality of images IMG1 to IMGN of region A1 from different fields of view. In some embodiments, as Figure 1 shown, the monitoring device 10[1] can use Figure 2 the camera 103 in Figure 2 to take a picture of region A1 in a preset viewing direction of the camera 103, so that
[0060] the camera 103 in Figure 2 can correspondingly generate the image IMG1. It should be understood that the remaining images IMG1 to IMGN can be captured in a manner similar to the image IMG1, so the description thereof is omitted here. Figure 2 DM1 to 2DMN. In operation S303, the processing device 20 generates a three-dimensional density Figure 3 DM based on the two-dimensional densitiesFigure 4 Describe operations S302 - S303 in detail.
[0061] Please refer to Figure 4 , Figure 4 images IMG1 - IMGN, two - dimensional densities Figure 2 DM1 - 2DMN, an aggregation volume model VM, and three - dimensional density Figure 3 DM as shown in some embodiments according to the present disclosure. In Figure 4 , the two - dimensional neural network model 110[1] is the same as the two - dimensional neural network model 110 in Figure 2 , and the two - dimensional neural network model 110[N] belongs to the monitoring device 10[N] in Figures 1 - 2 . For the sake of convenience of description, the two - dimensional neural network models of the remaining ones among the monitoring devices 10[1] - 10[N] are not shown in the drawings.
[0062] In some embodiments of operation S302, as Figure 4 shown, the monitoring devices 10[1] - 10[N] use the two - dimensional neural network models 110[1] - 110[N] to generate two - dimensional densities Figure 2 DM1 - 2DMN based on the images IMG1 - IMGN. Specifically, the two - dimensional neural network models 110[1] - 110[N] can perform convolution operations on the images IMG1 - IMGN to generate two - dimensional densities Figure 2 DM1 - 2DMN. As Figure 4 shown, the two - dimensional neural network model 110[1] that performs convolution operations on the image IMG1 can identify the target parts IB11 - IB12 (which can correspond to Figure 1 a part of the target objects B1 - BM in Figure 2 ) from the image IMG1, and can predict the target parts IB11 - IB12 with a preset pixel value (for example, a pixel value close to 1), so as to form characteristic pixel points PL11 - PL12 on the two - dimensional density Figure 1 DM1. It should be understood that the non - target parts on the image IMG1 (which may not correspond to Figure 2 any of the target objects B1 - BM in Figure 2 ) can be set to the minimum pixel value (for example: 0) significantly less than the preset pixel value, so as to form a non - characteristic pixel part on the two - dimensional density Figure 2 DM1. Similarly, the two - dimensional neural network model 110[N] that performs convolution operations on the image IMGN can also transform the target parts IBN1 - IBN2 and non - target parts on the image IMGN into characteristic pixel points PLN1 - PLN2 and non - characteristic pixel parts respectively, so as to form the two - dimensional density Figure 4 DMN. The remaining ones among the two - dimensional densitiesFigure 2 It is generated in the manner of DM1 and 2DMN, so its description is omitted here.
[0063] From Figure 2 and the description of operation S302, it can be seen that in some embodiments, the processor 101 is used to transform the image IMG1 into a two-dimensional density Figure 2 DM1 using the two-dimensional neural network model 110.
[0064] In some embodiments of operation S303, as Figure 4 shown, the processing device 20 is used to generate the aggregated volume model VM by projecting the two-dimensional density Figure 2 DM1~2DMN, and is used to generate a three-dimensional density Figure 3 DM according to the aggregated volume model VM. In some further embodiments, as Figure 2 shown, the processing device 20 projects the two-dimensional density Figure 2 DM1~2DMN based on a plurality of image capture data DC1~DCN transmitted from the monitoring devices 10[1]~10[N].
[0065] Since the generation of the image capture data DC1~DCN can be analogized, then the generation of the image capture data DC1~DCN will be described by taking the image capture data DC1 as an example. In some embodiments, as Figure 2 shown, the processor 101 accesses the camera parameter information P1 stored in the memory 107 as the image capture data DC1, and provides the image capture data DC1 to the processing device 20. Specifically, the camera parameter information P1 may include camera intrinsic parameters, camera extrinsic parameters, and distortion coefficients. The camera intrinsic parameters may indicate a projection transformation from a two-dimensional image coordinate system to a camera coordinate system, and can be determined after the camera 103 is manufactured. The camera extrinsic parameters may indicate a rigid transformation from the camera coordinate system to a three-dimensional world coordinate system, and can be determined according to a specific posture of the camera 103 of the monitoring device 10[1] after the monitoring device 10[1] is set in the area A1. Also, the distortion coefficients indicate calibration for various lens distortions (such as: radial distortion, tangential distortion, etc.), and can be determined after the camera 103 is manufactured.
[0066] As can be seen from the above description, the image capture data DC1 can be used to indicate the relationship between the image IMG1 and a specific three-dimensional space (e.g., area A1) captured by the camera 103 of the monitoring device 10[1]. In short, the image capture data DC1~DCN corresponds to the images IMG1~IMGN.
[0067] Continuing with the embodiment where the monitoring devices 10[1]~10[N] are movable when arranged in area A1, the pose of the camera 103 may change when the monitoring device 10[1] moves. Accordingly, the extrinsic parameters of the camera in the camera parameter information P1 should be updated when the monitoring device 10[1] moves. In some embodiments, as Figure 2 shown, the monitoring device 10[1] can use the processor 101 to generate the device pose information POS1 to update the extrinsic parameters of the camera in the camera parameter information P1. Continuing with the above description, the processor 101 can calculate the device pose information POS1 through at least one vision-based positioning technology. In other words, the device pose information POS1 can be the pose of the monitoring device 10[1] (which can generally be regarded as the pose of the camera 103). Therefore, the processor 101 can update the extrinsic parameters of the camera in the camera parameter information P1 through the device pose information POS1, and can transmit the camera parameter information P1 (with the updated extrinsic parameters of the camera) to the processing device 20 as the image capture data DC1. It can be seen from this that the processor 101 is used to generate the image capture data DC1 based on the device pose information POS1.
[0068] In some embodiments, based on the image capture data DC1~DCN, the processing device 20 can obtain the position of each of the monitoring devices 10[1]~10[N] in area A1 in real time.
[0069] Continuing with the embodiment of generating the aggregation volume model VM by projecting the two-dimensional densities Figure 2 DM1~2DMN, as Figure 4 shown, a three-dimensional cube related to area A1 can be predefined as a framework of the aggregation volume model VM. In some embodiments, the processing device 20 is used to calculate the two-dimensional densities Figure 2 DM1~2DMN according to the image capture data DC1~DCN, and calculate the position of each of the characteristic pixel points PL11~PL12 and PLN1~PLN2 in the aggregation volume model VM. For example, through the intrinsic parameters and extrinsic parameters of the camera in the image capture data DC1, the processing device 20 can transform the characteristic pixel point PL11 of the two-dimensional density Figure 2 DM1 from the two-dimensional image coordinate system to a three-dimensional coordinate system (e.g., three-dimensional world coordinate system) applied by the aggregation volume model VM, and can use the pixel value of the characteristic pixel point PL11 as the voxel value of a voxel point VL1 at the transformed coordinate.
[0070] Similarly, a two-dimensional density can be calculated based on the image capture data DC1 Figure 2 The position of the characteristic pixel point PL12 of DM1 in the aggregation volume model VM is obtained, thereby forming another voxel point VL2 of the aggregation volume model VM. Also, a two-dimensional density can be calculated based on the image capture data DCN Figure 2 The positions of the characteristic pixel points PLN1 to PLN2 of DMN in the aggregation volume model VM are obtained, thereby forming two voxel points VL1' and VL3 of the aggregation volume model VM.
[0071] As Figure 4 shown, in the aggregation volume model VM, the voxel point VL1 and the voxel point VL1' may overlap each other or be located at the same three-dimensional coordinates. In this case, the voxel value of the voxel point VL1 may be combined with the voxel value of the voxel point VL1' to generate a larger voxel value. However, this larger voxel value may cause the processing device 20 to calculate unacceptable results. For example, the voxel point VL1 and the voxel point VL1' located at the same three-dimensional coordinates may mean that the target parts IB11 and IBN1 correspond to the same one of the target objects B1 to BM, but the processing device 20 may calculate a number greater than 1 based on the larger voxel value corresponding to the voxel point VL1 (and / or the voxel point VL1').
[0072] In view of the above problems, the processing device 20 then uses the three-dimensional neural network model 210 to transform the aggregation volume model VM into a three-dimensional density Figure 3 DM. In some embodiments, the three-dimensional neural network model 210 performs a convolution operation on the aggregation volume model VM, thereby generating a three-dimensional density Figure 3 DM. Specifically, the three-dimensional neural network model 210 that performs the convolution operation on the aggregation volume model VM can eliminate the overlapping voxel points (e.g., Figure 4 the voxel point VL1' in) from the aggregation volume model VM, thereby generating a three-dimensional density Figure 3 DM.
[0073] In operation S304, the processing device 20 calculates the number of the target objects B1 to BM based on the three-dimensional density Figure 3 DM. In some embodiments, the processing device 20 performs at least one known counting method on the three-dimensional density Figure 3 DM to calculate the number of the target objects B1 to BM. For example, the processing device 20 can calculate the number of the target objects B1 to BM by summing or integrating the voxel points in the three-dimensional density Figure 3 DM.
[0074] As can be seen from the description of the above embodiments, the monitoring devices 10[1] to 10[N] should not be limited to as Figure 2The structures shown. For example, in some embodiments where the monitoring devices 10[1] to 10[N] are fixed in area A1, the sensor 105 can be from Figure 2 omitted. In some embodiments, the memory 107 can be integrated into the processor 101 such that the processor 101 can store the camera parameter information P1 and the memory 107 can be from Figure 2 omitted. In short, the structure of each of the monitoring devices 10[1] to 10[N] can be adjusted according to practical requirements. Further, any of the various operations described in the above embodiments (e.g., the generation of two-dimensional densities Figure 2 DM1 to 2DMN, the generation of the aggregated volume model VM, the generation of three-dimensional density Figure 3 DM, the counting of the number of target objects B1 to BM, etc.) can be performed in a cloud computing environment or by other computing resources, so that the monitoring devices 10[1] to 10[N], the processing device 20, and the cloud computing environment (or other computing resources) can share the heavy computing burden together.
[0075] As can be seen from the above embodiments of the present disclosure, through the monitoring devices 10[1] to 10[N] that can sense their own postures in area A1, the area information estimation system 100 and the area information estimation method 300 of the present disclosure can not only overcome the problems generated in crowded and / or blocked situations to perform crowd counting in a large-scale area, but also allow the monitoring devices 10[1] to 10[N] to move in area A1 without additional calibration. In addition, by the monitoring devices 10[1] to 10[N] predicting two-dimensional densities Figure 2 DM1 to 2DMN based on the images IMG1 to IMGN of area A1 and the processing device 20 fusing the outputs of the monitoring devices 10[1] to 10[N] to generate a three-dimensional density Figure 3 DM and calculating the number of target objects B1 to BM, the area information estimation system 100 and the area information estimation method 300 of the present disclosure have advantages such as sharing the heavy computing burden.
[0076] The method of the present disclosure can exist in the form of program code. The program code can be included in a physical medium, such as a floppy disk, an optical disc, a hard disk, or any other transient or non-transient computer-readable storage medium. Wherein, when the program code is loaded and executed by a computer, this computer becomes a device for implementing the method. The program code can also be transmitted through some transmission media, such as wires or cables, through optical fibers, or through any other transmission form. Wherein, when the program code is received, loaded, and executed by a computer, this computer becomes a device for implementing the method. When implemented on a general-purpose processor, the program code combines with the processor to provide a unique device that operates similar to application-specific logic circuits.
[0077] Although the present disclosure has been disclosed as above in embodiments, it is not intended to limit the present disclosure. Those of ordinary skill in the art can make various changes and modifications without departing from the spirit and scope of the present disclosure. Therefore, the protection scope of the present disclosure shall be subject to that defined by the appended claims.
Claims
1. A method for estimating regional information, characterized in that, Applicable to a regional information estimation system including a processing device and multiple monitoring devices, and comprising: Through the multiple monitoring devices, multiple images of an area are captured from different perspectives; Through the multiple monitoring devices, based on the multiple images, multiple two-dimensional density maps of at least one target object within the area are generated; Through the processing device, based on the multiple two-dimensional density maps, a three-dimensional density map is generated; And Through the processing device, based on the three-dimensional density map, the quantity of the at least one target object is calculated.
2. The regional information estimation method according to claim 1, wherein Generating the multiple two-dimensional density maps of the at least one target object based on the multiple images includes: Through the multiple monitoring devices, using a two-dimensional neural network model to transform the multiple images into the multiple two-dimensional density maps.
3. The regional information estimation method according to claim 2, wherein The two-dimensional neural network model is a convolutional neural network.
4. The regional information estimation method according to claim 1, wherein It further comprises: Through the multiple monitoring devices, multiple image capture data corresponding to the multiple images are provided to the processing device.
5. The regional information estimation method according to claim 4, wherein When at least one of the multiple monitoring devices is moved, the regional information estimation method further comprises: Through the at least one of the multiple monitoring devices, using a vision-based positioning technique to calculate at least one device pose information, thereby generating at least one of the multiple image capture data.
6. The regional information estimation method according to claim 4, wherein It further comprises: Through the multiple monitoring devices, multiple camera parameter information of multiple cameras of the multiple monitoring devices is accessed as the multiple image capture data.
7. The regional information estimation method according to claim 1, characterized in that Generating the three-dimensional density map based on the multiple two-dimensional density maps includes: Based on multiple image capture data, projecting the multiple two-dimensional density maps to generate an aggregated volume model; and Based on the aggregated volume model, generating the three-dimensional density map.
8. The regional information estimation method according to claim 7, wherein Projecting the multiple two-dimensional density maps based on the multiple image capture data to generate the aggregated volume model includes: Based on the multiple image capture data, calculating the position of at least one feature pixel point of the multiple two-dimensional density maps within the aggregated volume model, thereby forming at least one voxel point of the aggregated volume model.
9. The regional information estimation method according to claim 7, wherein Generating the three-dimensional density map based on the aggregated volume model includes: Using a three-dimensional neural network model to transform the aggregated volume model into the three-dimensional density map.
10. The regional information estimation method according to claim 9, wherein The three-dimensional neural network model is a convolutional neural network.
11. A regional information estimation system, characterized in that, It comprises: Multiple monitoring devices, used to be arranged within an area, used to capture multiple images of the area from different perspectives, and used to generate multiple two-dimensional density maps of at least one target object within the area based on the multiple images; And A processing device, coupled to the multiple monitoring devices, used to generate a three-dimensional density map based on the multiple two-dimensional density maps, and used to calculate the quantity of the at least one target object based on the three-dimensional density map.
12. The regional information estimation system according to claim 11, characterized in that, Each of the multiple monitoring devices comprises: A camera, used to capture a corresponding one of the multiple images; and A processor, coupled to the camera, and used to use a two-dimensional neural network model to transform the corresponding one of the multiple images into a corresponding one of the multiple two-dimensional density maps.
13. The regional information estimation system according to claim 11, wherein The plurality of monitoring devices are configured to provide a plurality of image capture data corresponding to the plurality of images to the processing device, wherein each of the plurality of monitoring devices includes: a camera configured to capture a corresponding one of the plurality of images; a memory configured to store camera parameter information of the camera, wherein the camera parameter information includes an intrinsic camera parameter, an extrinsic camera parameter, and a distortion coefficient; and a processor coupled to the camera and the memory and configured to access the camera parameter information as a corresponding one of the plurality of image capture data.
14. A non-transitory computer-readable storage medium, characterized in that, There is a computer program for executing a region information estimation method, wherein the region information estimation method is applicable to a region information estimation system including a processing device and a plurality of monitoring devices, and includes: capturing, by the plurality of monitoring devices, a plurality of images of a region from different fields of view; generating, by the plurality of monitoring devices, a plurality of two-dimensional density maps of at least one target object within the region based on the plurality of images; generating, by the processing device, a three-dimensional density map based on the plurality of two-dimensional density maps; and calculating, by the processing device, a quantity of the at least one target object based on the three-dimensional density map.