Counting method, counting device, and counting system

The method addresses the limitations of existing object recognition and building information modeling by using a trained discriminator to convert two-dimensional images into three-dimensional point cloud data for accurate and efficient object counting.

JP2026088876APending Publication Date: 2026-05-29UNIV OF TSUKUBA +1

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
UNIV OF TSUKUBA
Filing Date
2024-11-19
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing object recognition technologies using 3D laser scanners and machine learning models face challenges with labeling costs and stability, and building information modeling for facility management is time-consuming and costly, necessitating a low-cost and efficient method for object counting.

Method used

A method and system that utilizes a trained discriminator to discriminate specific objects from two-dimensional images, converts them into three-dimensional point cloud data, and generates clusters for accurate counting, using unsupervised learning and 3D shape reconstruction, with optional color marking and transparency for improved accuracy.

Benefits of technology

Enables efficient and accurate counting of objects without duplication, reducing processing time and costs by leveraging two-dimensional images and three-dimensional point cloud data generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026088876000001_ABST
    Figure 2026088876000001_ABST
Patent Text Reader

Abstract

To provide a counting method, counting device, and counting system that can count objects accurately at low cost. [Solution] A counting method for a counting device 3 that counts specific objects, comprising the steps of: acquiring a group of images covering a predetermined area; discriminating specific objects using a trained discriminator 29 and the group of images; marking the discriminated specific objects; creating point cloud data of a predetermined area using the group of images by a 3D shape reconstruction process; generating clusters of the marked specific objects from the point cloud data; and counting the generated clusters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a counting method, a counting device, and a counting system.

Background Art

[0002] When carrying out renovation work on a building, it is necessary to conduct an on-site survey. In old non-residential buildings, the existing equipment conditions are often not recorded. Therefore, there is a need for a technology to quickly grasp (count) equipment-related information and convert it into data. In addition, from the perspective of achieving SDGs (sustainable society), it is also important to maintain the building performance through equipment maintenance.

[0003] In relation to this, there are technologies such as object recognition using a 3D laser scanner and a machine learning model (see, for example, Patent Documents 1 to 3). Also, as a social trend, facility management using Building Information Modeling (BIM) is progressing.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Patent Document 2

Patent Document 3

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, in object recognition using 3D laser scanners and machine learning models, qualitative and quantitative limitations and stability issues have been pointed out regarding labeling and supervised learning. Semi-automated methods for automatically detecting object boundaries and regions have been proposed, but fundamental automation is not possible because labeling is required on a region-by-region basis. There are also methods that label 3D point cloud data of the object, but the cost of labeling 3D point clouds is far higher than that of 2D images, making it impractical.

[0006] Furthermore, it is estimated that it will take 5 to 6 years to implement building information modeling for facility management in existing buildings. Therefore, low-cost facility management is required.

[0007] From this perspective, the present invention provides a counting method, a counting device, and a counting system that can count objects accurately at low cost. [Means for solving the problem]

[0008] The counting method according to the present invention is a method for counting specific objects. This counting method includes the steps of: acquiring a group of images covering a predetermined region; discriminating specific objects using a trained discriminator and the group of images; marking the discriminated specific objects; creating point cloud data of a predetermined region using the group of images by a 3D shape reconstruction process; generating clusters of the marked specific objects from the point cloud data; and counting the generated clusters.

[0009] The counting method according to the present invention enables computationally efficient processing by performing discrimination from a two-dimensional planar image. Furthermore, by converting the two-dimensional planar image into three-dimensional point cloud data, accurate counting without duplication can be performed.

[0010] In the step of generating the aforementioned clusters, an automatic aggregater may be used to generate the clusters. An example of an automatic aggregater may be cluster analysis used in multivariate analysis. Cluster analysis is an unsupervised learning method.

[0011] This method allows for the automatic generation of clusters of specific objects, thereby improving the efficiency of counting.

[0012] In the marking process step, a specific object on the image in a predetermined area may be overwritten with color information, or color information and transparency.

[0013] In this way, it is possible to form point cloud data that characterizes specific objects, thereby improving counting accuracy.

[0014] The procedure may include the steps of setting a reference known point in a predetermined area and shooting a video while moving the imaging device so that the number of frames per unit distance remains constant.

[0015] This method allows you to determine the location of a specific object based on known points, making it easier to align it with drawings and other documents.

[0016] In the step of acquiring the aforementioned image set, the following steps may be performed: extracting video from the 360-degree video in multiple directions, and extracting images from the extracted video under predetermined conditions.

[0017] This method ensures that all specific objects within a given area are extracted without omission. Furthermore, by reducing the size of the extracted data, the processing time for both the image acquisition step and the image extraction step can be shortened.

[0018] The counting device according to the present invention is a counting device that counts specific objects. This counting device includes an image group acquisition unit, an object discrimination unit, a marking processing unit, a three-dimensional shape creation unit, a clustering processing unit, and a counting processing unit. The image group acquisition unit acquires an image group that covers a predetermined area. The object discrimination unit discriminates specific objects using a learned discriminator and the image group. The marking processing unit performs a marking process on the discriminated specific objects. The three-dimensional shape creation unit creates point cloud data of a predetermined area using the image group by three-dimensional shape restoration processing. The clustering processing unit generates the marked specific objects from the point cloud data as clusters. The counting processing unit counts the generated clusters.

[0019] In the counting device according to the present invention, highly efficient processing can be performed by discriminating from a two-dimensional planar image. Also, accurate counting without duplication can be performed by converting the two-dimensional planar image into three-dimensional point cloud data.

[0020] The clustering processing unit may use an automatic aggregator to generate the clusters. An example of the automatic aggregator may be cluster analysis used in multivariate analysis. Cluster analysis is unsupervised learning.

[0021] By doing so, clusters of specific objects can be automatically generated. Therefore, the efficiency of counting can be improved.

[0022] The marking processing unit may overwrite a specific object on the image of the predetermined area by designating color information or color information and transparency.

[0023] By doing so, it is possible to form point cloud data that characterizes specific objects, and the counting accuracy can be improved.

[0024] In a predetermined area, known points serving as references are installed, and the image group may be generated based on a moving image captured while moving such that the number of frames per unit distance is constant.

[0025] By doing so, since the position of a specific object can be grasped based on the known points, it is easy to match with drawings and the like.

[0026] The image group acquisition unit may cut out videos in a plurality of directions from an omnidirectional video and extract images from the cut-out videos under predetermined conditions.

[0027] By doing so, a specific object in a predetermined area can be extracted without omission. Further, by reducing the capacity of the cut-out data, the processing time for acquiring the image group and the processing time for extracting the images can be shortened.

[0028] The counting system according to the present invention is a counting system for counting a specific object. This counting system includes the above-described counting device and an imaging device that captures the image group.

[0029] In the counting system according to the present invention, highly efficient processing can be performed by discriminating from a two-dimensional planar image. Further, by converting the two-dimensional planar image into three-dimensional point cloud data, accurate counting without duplication can be performed.

Effects of the Invention

[0030] According to the present invention, an object can be accurately counted at low cost.

Brief Description of the Drawings

[0031] [Figure 1] It is a configuration diagram of a counting system according to a first embodiment of the present invention. [Figure 2] It is a functional configuration diagram of a counting device according to a first embodiment of the present invention. [Figure 3]This is an illustrative diagram summarizing the processing related to the counting method implemented by the counting system according to the first embodiment of the present invention. [Figure 4] This is an image of point cloud data after 3D shape reconstruction processing. [Figure 5] This is a diagram showing the configuration of a counting system according to a second embodiment of the present invention. [Modes for carrying out the invention]

[0032] Hereinafter, embodiments for carrying out the present invention will be described in detail with reference to the drawings as appropriate. Each figure is only a schematic representation to the extent that the present invention can be fully understood. Therefore, the present invention is not limited to the illustrated examples. In each figure, common or similar components are denoted by the same reference numerals, and their redundant descriptions may be omitted.

[0033] [First Embodiment] <Configuration of the remote control system according to the first embodiment> Referring to Figure 1, the counting system 1 according to the first embodiment will be described. Figure 1 is a configuration diagram of the counting system 1 according to the first embodiment. The counting system 1 detects specific objects (for example, measuring instruments) using images taken of the area to be counted, and counts the number of specific objects based on the detection results. The images may also be images (also called frames) that make up a video.

[0034] In this embodiment, the area to be counted (i.e., the space where the specified object is located) is assumed to be a renovation construction site, and the specified object is assumed to be measuring instruments. In other words, the case of counting measuring instruments installed inside a building undergoing renovation is described. Note that the specified object may not be a measuring instrument, and the space in which the specified object is located is not limited to inside a building undergoing renovation. The space in which the specified object is located may be outdoors, and it is also possible to count specified objects located outdoors using the counting system 1.

[0035] When counting is performed in a target area (i.e., the space where specific objects are placed), the device or person (e.g., a drone with an imaging device, or a worker performing renovation work) to count the specific objects can be arbitrarily set. The area to be counted may be, for example, the entire room, or a part of the room (e.g., a single wall that makes up the room). The area to be counted is referred to as the "predetermined area."

[0036] As shown in Figure 1, the counting system 1 mainly comprises an imaging device 2 and a counting device 3. The number of imaging devices 2 is not particularly limited; for example, a predetermined area can be captured by multiple imaging devices 2. The counting device 3 may be located within the predetermined area or outside of it. The imaging device 2 and the counting device 3 can communicate with each other at all times or temporarily via some means of communication (e.g., wireless LAN (Local Area Network) or communication cable). The video (or image) captured by the imaging device 2 may be recorded on a computer-readable recording medium (e.g., a memory card), and the video may be acquired by the counting device 3 by inserting the recording medium into the counting device 3 as needed. The imaging device 2 and the counting device 3 may be configured as a single device. That is, the imaging device 2 may have the hardware configuration and functions of the counting device 3 (or the counting device 3 may have the hardware configuration and functions of the imaging device 2).

[0037] The imaging device 2 shown in Figure 1 captures still or moving images and outputs the captured still or moving images. The imaging device 2 is, for example, a digital video camera capable of capturing dozens of frames (images) per second. It is desirable that the imaging device 2 is capable of capturing a wide area at once. The imaging device 2 is, for example, a 360-degree camera, and in this embodiment, a 360-degree camera is assumed to be the imaging device 2 and will be described accordingly. A 360-degree camera is a camera that can capture all directions of "360 degrees". A 360-degree camera is equipped with, for example, two wide-angle lenses that can capture "180 degrees" or more, and processes images by simultaneously capturing with each lens and stitching them together. As a result, a 360-degree camera can capture 360 ​​degrees in all directions in one shot. The imaging device 2 may also capture the target space using an illumination device.

[0038] As shown in Figure 1, the predetermined area contains a specific object to be counted (here, instrument 8) and an object not to be counted (e.g., pipe 9). The imaging device 2 takes images including the specific object (here, instrument 8) and the object not to be counted (e.g., pipe 9). The imaging device 2 is attached to a mobile body that moves within the predetermined area (for example, a human, robot, or drone), and takes images of the predetermined area while moving with the mobile body. The means of movement of the mobile body are not particularly limited, and methods of movement include walking, running, and flying. For example, the imaging device 2 is fixed to the end of a rod-shaped handle 7 (e.g., a monopod), and the worker moves within the predetermined area while holding the handle 7. The imaging device 2 may also be fixed to a part of the worker's body (e.g., the head). As another example, the imaging device 2 is installed on a drone. The drone flies within the predetermined area, for example, by remote control or according to a pre-set route or conditions.

[0039] The counting device 3 shown in Figure 1 counts the number of specific objects (here, instruments 8 are assumed) from video (or images) captured by the imaging device 2. The counting device 3 may also count the total number of specific objects (here, thermometers 8a and pressure gauges 8b) from video (or images) captured by the imaging device 2. The counting device 3 may be, for example, a management PC installed in the management room of a renovation construction site, or a server located away from the renovation construction site. The counting device 3 may also be part of a cloud system.

[0040] Referring to Figure 2, the configuration and functions of the counting device 3 will be explained. Figure 2 is a functional configuration diagram of the counting device 3. As shown in Figure 2, the counting device 3 mainly comprises a storage unit 10 and a control unit 20. Although not shown in the illustration, the counting device 3 may also include components such as an input unit, a display unit, and a communication unit.

[0041] The counting device 3 shown in Figure 2 counts the number of specific objects using artificial intelligence (AI) technology. In this embodiment, we will explain assuming that the processing related to the learning of a machine learning model (for example, a neural network, but a Transformer or similar may also be used) (learning stage processing), the discrimination of specific objects using the trained machine learning model, and the processing related to counting by an automatic aggregater (counting stage processing) are all performed by a single device. An example of an automatic aggregater may be cluster analysis used in multivariate analysis. Cluster analysis is an unsupervised learning method. In other words, the counting device 3 performs the learning of the machine learning model, discrimination using the trained machine learning model, and counting without prior training. It is also possible to perform the learning stage processing and the counting stage processing by different devices. The device that performs the learning stage processing will be specifically referred to as the "learning device".

[0042] The memory unit 10 is a component that stores information necessary for counting specific objects. The memory unit 10 is, for example, a storage medium such as RAM (Random Access Memory), ROM (Read Only Memory), HDD (Hard Disk Drive), SSD (Solid State Drive), or flash memory. In addition, the memory unit 10 according to this embodiment stores information necessary for training a machine learning model. The memory unit 10 has, for example, a 360-degree video and training data. The 360-degree video is a video of a predetermined area captured by the imaging device 2.

[0043] The control unit 20 is a component that counts specific objects based on video (or images) captured from a predetermined area. The control unit 20 is composed of, for example, a CPU (Central Processing Unit) and its peripheral devices, reads various processing programs, loads them into RAM, and performs various processing in cooperation with those programs. Through program execution, the control unit 20 realizes functions such as, for example, an image group acquisition unit 21, an object discrimination unit 22, a marking processing unit 23, a 3D shape creation unit 24, a clustering processing unit 25, and a counting processing unit 26.

[0044] The object discrimination unit 22 has a trained discriminator 29. The discriminator 29 is an example of a machine learning model. The trained discriminator 29 has been trained using training data to discriminate specific objects in an image.

[0045] The image group acquisition unit 21, target discrimination unit 22, marking processing unit 23, 3D shape creation unit 24, clustering processing unit 25, and counting processing unit 26 are functions related to the counting stage.

[0046] The functions of the control unit 20 will now be described. Only an overview of each function will be provided here; the details of each function will be explained later in the section describing the counting method. The image acquisition unit 21 acquires a set of images covering a predetermined area from the 360-degree video. The object discrimination unit 22 uses a trained discriminator 29 and a set of images covering a predetermined region to discriminate against a specific object. The marking processing unit 23 marks the identified specific object on the image group.

[0047] The 3D shape creation unit 24 creates point cloud data of a predetermined region by 3D shape reconstruction processing using a set of images. The point cloud data obtained through 3D shape reconstruction processing reflects the results of the process that marked a specific object. The clustering processing unit 25 generates clusters of specific objects that have been marked from the point cloud data. The counting processing unit 26 counts the generated clusters. The number of clusters is the number of specific objects.

[0048] <Regarding the counting method in the counting system according to the first embodiment> Referring to Figure 3 (and Figures 1 and 2 as appropriate), the processing of the counting method performed by the counting system 1 will be described. Figure 3 is an illustrative diagram summarizing an example of the processing of the counting method performed by the counting system 1 according to the first embodiment of the present invention.

[0049] Regarding the processing at the counting stage: (Image capture of the designated area "Step S10") In the predetermined area imaging step S10, for example, the photographer (here, assumed to be a worker) uses the imaging device 2 to image the predetermined area. The photographer walks through all walkable areas within the predetermined area. For example, while maintaining a certain distance (for example, about 1 to 3 meters) from the subject (for example, a wall or an object placed in a room), the photographer walks around the predetermined area in a single continuous line.

[0050] The photographer should hold the handle 7, keep their elbows close to their body, and maintain an upright shooting posture as much as possible. The photographer should move in such a way that the number of frames per unit distance remains constant (if the frame rate is constant, it is desirable to maintain a constant walking speed). The photographer should move slowly, for example, with a stride the size of their shoes. The image sensor 2 should be oriented so that one of its lenses is pointed in the direction of movement. Care should be taken not to bring the image sensor 2 too close to the photographer.

[0051] Photographers may use lighting equipment to photograph the subject space. Lighting equipment is useful, for example, when the shape of the subject is not clearly visible in a dimly lit space. It is especially desirable to use lighting equipment in narrow and dark spaces. When taking close-up shots with lighting equipment, care must be taken to avoid overexposure.

[0052] (Image acquisition "Step S20") The image acquisition step S20 is a step of acquiring an image set that covers a predetermined area from a 360-degree video. The image acquisition step S20 mainly consists of a video cutting step S21 and an image extraction step S22.

[0053] The image acquisition unit 21 extracts 2D videos in multiple directions from the 360-degree video captured in step S10 (step S21). The 2D video extraction process (cropping) is performed for directions such as "8 to 12". The 2D video extraction process should satisfy conditions such as there being no gaps between the 2D videos and the degree of overlap between adjacent 2D videos being appropriate. This helps to avoid overlooking objects with the instrument 8 and leads to improved accuracy in the 3D shape reconstruction process described later.

[0054] Next, the image group acquisition unit 21 extracts 2D images (frames) from the 2D video in multiple directions extracted in step S21 under predetermined conditions (step S22). In other words, it extracts a portion of the image divided into frame units. The image group acquisition unit 21 extracts images continuously from the 2D video at predetermined intervals, for example. This generates an image group that covers a predetermined area. In this embodiment, the case of extracting 2D video from a 360-degree video and then extracting it under predetermined conditions has been described, but it is also possible to acquire 2D video that has been extracted in advance. In other words, acquiring an image group here includes acquiring an image group that has already been generated.

[0055] (Identification of specific objects "Step S30") Step S30, which involves identifying a specific object, uses the trained classifier 29 and the image set generated in step S20 to identify a specific object (here, an instrument 8) that appears in the image set. "Identification" here includes (1) detecting one type of specific object, and (2) detecting multiple types of specific objects. If (1) is the case, for example, the instrument 8 that appears in the image can be detected. If (2) is the case, for example, the thermometer 8a and pressure gauge 8b that appear in the image can be detected. The trained classifier 29 has been machine-trained to identify specific objects that appear in an image when an image is input.

[0056] In this embodiment, the trained discriminator 29 is used to distinguish between multiple types of specific objects. The object discrimination unit 22 uses the trained discriminator 29 and a group of images covering a predetermined region to detect the thermometer 8a and pressure gauge 8b that appear in the images.

[0057] (Marking process "Step S40") The marking process step S40 is a step in which the identified specific object is marked on the image group. The marking process is a process of overwriting the pixel values ​​(RGB values) that make up the image. In the marking process step S40, the specific object on the image in a predetermined area is overwritten with color information, or color information and transparency.

[0058] For example, the area of ​​thermometer 8a is overwritten with red, and the area of ​​pressure gauge 8b is overwritten with green. It is preferable to use transparent markings, and it is desirable to set an appropriate transparency (= number of pixels not marked / total number of pixels) and perform the marking process. By making the markings transparent, the problem of the marking information not being properly reproduced in the 3D data due to cancellation when they overlap during the 3D reconstruction process described later can be suppressed.

[0059] (Step S50: Creating point cloud data) In step S50, which creates point cloud data, point cloud data for a predetermined region is created by a 3D shape reconstruction process using an image to which predetermined marking information has been attached to a specific object. The 3D shape creation unit 24 reconstructs the 3D shape of the predetermined region using, for example, SfM (Structure from Motion) technology. SfM is a technology that reconstructs a 3D shape from image data from different viewpoints. The point cloud data after 3D shape reconstruction processing reflects the results of the process in which the specific object was marked (marking information). Figure 4 is an image diagram of the point cloud data after 3D shape reconstruction processing.

[0060] (Clustering "Step S60") In the clustering step S60, specific objects marked are generated as clusters from the point cloud data. The clustering processing unit 25 filters the point cloud data according to RGB conditions to extract points with marking information (for example, red or green), and constructs clusters based on the x, y, and z coordinates of the extracted points. For example, the number of clusters generated from extracted red points is the number of thermometers 8a. Also, for example, the number of clusters generated from extracted green points is the number of pressure gauges 8b.

[0061] (Counting "Step S70") In the counting step S70, the total number of thermometers 8a and pressure gauges 8b is counted as the number of instruments 8, based on the clustering results performed in step S60.

[0062] As described above, the counting device 3 and its counting method according to the first embodiment can perform computationally efficient processing by discriminating from a two-dimensional planar image. Furthermore, by converting the two-dimensional planar image into three-dimensional point cloud data, accurate counting without duplication can be performed.

[0063] [Second Embodiment] The counting system 201 according to the second embodiment (see Figure 5) is provided with known point marks 6 in a predetermined area compared to the counting system 1 according to the first embodiment. For example, a two-dimensional barcode with coordinate information can be used as the known point marks 6. For example, by combining two-dimensional barcodes with different coordinate information, the scale can be reflected in the 3D point cloud data. For example, four two-dimensional barcodes with origin coordinates (0, 0, 0) and other absolute coordinates (0.15, 0, 0), (0, 0, 0.2), and (0.15, 0, 0.2) can be used. The known point marks 6 may also be assigned the role of assigning an origin to the 3D point cloud data.

[0064] If a known point mark 6 is placed in a predetermined area, the location where the known point mark 6 is attached may be used as the starting and ending point of the shooting. In other words, shooting starts from the location where the known point mark 6 is attached, circles the predetermined area, and returns to the location where the known point mark 6 is attached to end the shooting. In this case, for example, shooting starts from within a range of "20 to 30 cm" from the location where the known point mark 6 is attached, and the camera moves away to a distance of "1.5 m" over a period of "10 seconds" before circling the area.

[0065] Alternatively, a mark for feature point extraction may be provided instead of, or in conjunction with, the known point mark 6. Photogrammetry techniques struggle with spaces that have a patterned shape (such as symmetrical buildings). By adding a mark for feature point extraction, accurate 3D point cloud data can be created even in spaces without distinctive features. Furthermore, when using a mobile device such as a drone equipped with an imaging device 2, it is desirable to move in such a way that the number of frames per unit distance remains constant.

[0066] As described above, the counting system 201 according to the second embodiment can determine the position of a specific object based on known points, making it easy to align it with drawings and the like.

[0067] Although embodiments of the present invention have been described above, the present invention is not limited thereto and can be implemented without changing the spirit of the claims.

[0068] Each embodiment assumed the use of a 360-degree camera for imaging. However, when the imaging device 2 is attached to a moving object moving on the floor, or when the imaging device 2 is attached to a moving object moving near the ceiling or in the air (for example, a drone), images in the direction of the floor and ceiling may not be necessary. In such cases, a hemispherical camera can also be used. [Explanation of Symbols]

[0069] 1,201 Counting Systems 2. Imaging device 3. Counting device 6. Mark known points 7. Handle 8. Instruments (Specific Objects) 8a Thermometer (Specific object) 8b Pressure gauge (specific object) 9 Piping 10 Storage section 20 Control Unit 21 Image acquisition unit 22 Target discrimination unit 23 Marking Processing Unit 24 3D Shape Creation Unit 25 Clustering Processing Unit 26 Counting Processing Unit 29 Discriminator

Claims

1. A method for counting specific objects, The steps include acquiring a set of images that cover a predetermined area, A step of identifying a specific object using a trained classifier and the image set, The steps include marking the identified specific object, A step of creating point cloud data of a predetermined region using the image group by 3D shape reconstruction processing, A step of generating the specified object, which has been marked from the point cloud data, as a cluster, The steps include counting the generated clusters, A method for counting specific objects.

2. The counting method according to claim 1, The method for generating the aforementioned cluster is characterized by using an automatic aggregater. A method for counting specific objects.

3. The counting method according to claim 1, In the marking process step, a specific object on the image in a predetermined area is overwritten with color information, or color information and transparency. A method for counting specific objects.

4. The counting method according to claim 1, The steps include: setting a reference known point in a predetermined area, The process involves the steps of: capturing video while moving using an imaging device so that the number of frames per unit distance remains constant; A method for counting specific objects.

5. The counting method according to claim 1, In the step of acquiring the aforementioned image set, Steps to extract video from a 360-degree video in multiple directions, The process involves extracting images from the extracted video according to predetermined conditions, and then performing the following steps: A method for counting specific objects.

6. A counting device for counting specific objects, An image group acquisition unit that acquires a group of images covering a predetermined area, A target discrimination unit that uses a trained discriminator and the aforementioned image group to discriminate a specific object, A marking processing unit that marks the identified specific object, A 3D shape creation unit creates point cloud data of a predetermined region using the image group by 3D shape reconstruction processing, A clustering processing unit that generates the specified objects marked from point cloud data as clusters, The system comprises a counting processing unit that counts the generated clusters, Counting device.

7. A counting device according to claim 6, The clustering processing unit is characterized by using an automatic aggregater to generate the clusters. Counting device.

8. A counting device according to claim 6, The marking processing unit overwrites a specific object on an image in a predetermined area with color information, or color information and transparency. Counting device.

9. A counting device according to claim 6, A reference known point is set within the designated area. The aforementioned image set was generated based on a video captured while moving so that the number of frames per unit distance remained constant. Counting device.

10. A counting device according to claim 6, The aforementioned image acquisition unit, Extract video clips from a 360-degree video in multiple directions. Extract images from the extracted video according to specified conditions. Counting device.

11. A counting system for counting specific objects, The counting device according to claim 6, The system comprises an imaging device for capturing the aforementioned group of images, Counting system.