Detection device, detection system, detection method, and program

WO2026203462A1PCT designated stage Publication Date: 2026-10-01MITSUBISHI HEAVY IND LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/034919
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-26
Filing Date
2025-10-01
Publication Date
2026-10-01

Smart Images

  • Figure JP2025034919_01102026_PF_FP_ABST
    Figure JP2025034919_01102026_PF_FP_ABST
Patent Text Reader

Abstract

This detection device comprises: an acquisition unit that acquires an evaluation image in which the inside of a refuse pit is captured; an extraction unit that extracts a feature vector from the evaluation image by using a pre-trained image recognition model; an evaluation unit that calculates, by using an evaluation model that has learned a reference vector obtained by reducing the dimensions of a feature vector extracted from a normal image captured when the inside of the refuse pit is in a normal state, the distance between the reference vector and the feature vector extracted from the evaluation image, and evaluates an abnormality score for each region of the evaluation image on the basis of the calculated distance; and a determination unit that determines that an inappropriate object is present in a region in which the abnormality score is greater than or equal to an abnormality threshold, when the area of the region is greater than or equal to an upper limit area.
Need to check novelty before this filing date? Find Prior Art

Description

Detection Apparatus, Detection System, Detection Method, and Program

[0001] The present disclosure relates to a detection apparatus, a detection system, a detection method, and a program. The present application claims priority based on Japanese Patent Application No. 2025-051066 filed in Japan on March 26, 2025, the content of which is incorporated herein by reference.

[0002] Patent Document 1 describes a technique for detecting the presence or absence of an abnormality included in a target image using an AI model (for example, PatchCore) that has learned normal images. PatchCore stores, in a memory bank, feature amounts for each region (patch) obtained by inputting a normal image to a trained image recognition model, and detects an abnormality included in the target image based on the distance between the feature amounts stored in the memory bank and the feature amounts obtained by inputting the target image to the image recognition model. Further, Patent Document 1 describes that, in consideration of the possibility that the normal state of an abnormality detection target may change due to environmental changes or temporal changes, a user appropriately performs relearning according to the changes.

[0003] Further, Patent Document 2 describes a technique for detecting whether or not unsuitable objects are present in a garbage pit of a garbage incineration facility by combining detection results of a plurality of trained models.

[0004] Japanese Unexamined Patent Application Publication No. 2024-79411 Japanese Unexamined Patent Application Publication No. 2021-64139

[0005] For example, since the normal state of a garbage pit in a garbage incineration facility is extremely diverse, in order to improve the detection accuracy of unsuitable objects, it is necessary to cause a model to learn a large number of normal images showing various normal states and store them in the memory bank. As the memory bank increases, the calculation time and the amount of memory used required to calculate the distance between the feature amounts of the memory bank and the feature amounts of the target image also increase accordingly. In this case, if a machine with small memory is used, a memory error may occur, which may make it difficult to detect an abnormality.

[0006] Therefore, conventional techniques reduce memory consumption by clustering memory banks according to similar vectors using a greedy algorithm and using the representative points of each cluster to calculate distances during evaluation. However, since the greedy algorithm requires calculating the distance between all vectors in the memory bank, the memory consumption during the greedy algorithm calculation increases as the number of memory banks increases. In other words, there is a possibility of memory errors occurring even during the calculation of the greedy algorithm.

[0007] Therefore, there is a need for a technology that can accurately detect anomalies while keeping computation time and memory consumption within a range that is feasible for machines actually used in waste incineration facilities.

[0008] The purpose of this disclosure is to provide a detection device, detection system, detection method, and program that can accurately detect anomalies while reducing the computation time and memory consumption of the process for detecting unsuitable materials in a waste pit.

[0009] According to one aspect of the present disclosure, the detection device includes: an acquisition unit that acquires evaluation images taken inside a garbage pit; an extraction unit that extracts feature vectors from the evaluation images using a pre-trained image recognition model; an evaluation unit that calculates the distance between the reference vectors and the feature vectors extracted from the evaluation images using an evaluation model that has learned a reference vector obtained by reducing the dimensionality of feature vectors extracted from a normal image taken when the inside of the garbage pit is in a normal state, and evaluates the abnormality score for each region of the evaluation image based on the calculated distance; and a determination unit that determines that there are unsuitable objects in a region when the area of ​​the region in which the abnormality score is equal to or greater than an abnormal threshold is equal to or greater than an upper limit area.

[0010] According to one aspect of this disclosure, the detection system comprises the above-described detection device and a learning device which includes a learning unit that learns the reference vector obtained by quantizing the feature vector extracted from the normal image to create the evaluation model.

[0011] According to one aspect of the present disclosure, the detection method includes the steps of: acquiring an evaluation image taken inside a garbage pit; extracting feature vectors from the evaluation image using a pre-trained image recognition model; calculating the distance between the reference vector and the feature vector extracted from the evaluation image using an evaluation model trained on a reference vector obtained by reducing the dimensionality of feature vectors extracted from a normal image taken when the inside of the garbage pit is in a normal state, and evaluating the abnormality score for each region of the evaluation image based on the calculated distance; and determining that an unsuitable object exists in a region if the area of ​​the region in which the abnormality score is equal to or greater than an abnormality threshold is equal to or greater than an upper limit area.

[0012] According to one aspect of the present disclosure, the program causes a detection device to perform the following steps: acquire an evaluation image taken inside a garbage pit; extract feature vectors from the evaluation image using a pre-trained image recognition model; calculate the distance between the reference vector and the feature vector extracted from the evaluation image using an evaluation model that has learned a reference vector obtained by reducing the dimensionality of feature vectors extracted from a normal image taken when the inside of the garbage pit is in a normal state, and evaluate the abnormality score for each region of the evaluation image based on the calculated distance; and determine that an unsuitable object exists in a region if the area of ​​the region in which the abnormality score is equal to or greater than an abnormal threshold is equal to or greater than an upper limit area.

[0013] According to the above embodiment, it is possible to detect abnormalities with high accuracy while reducing the computation time and memory consumption of the process for detecting unsuitable materials in the waste pit.

[0014] This is a diagram showing the functional configuration of the detection system according to the first embodiment. This is a flowchart showing an example of the learning phase processing of the evaluation model according to the first embodiment. This is a diagram showing an example of the process of reducing the dimensionality of feature vectors using a vector quantization method. This is a flowchart showing an example of the evaluation phase processing using the evaluation model according to the first embodiment. This is a diagram showing an example of the measurement results of the detection accuracy and memory consumption of the detection device according to the second embodiment. This is a diagram for explaining the evaluation model according to the third embodiment. This is a flowchart showing an example of the learning phase processing of the evaluation model according to the third embodiment. This is a flowchart showing an example of the evaluation phase processing using the evaluation model according to the third embodiment. This is a diagram showing an example of the measurement results of the detection accuracy of the detection device according to the third embodiment. This is a diagram for explaining the evaluation model according to the fourth embodiment. This is a flowchart showing an example of the learning phase processing of the evaluation model according to the fourth embodiment. This is a flowchart showing an example of the evaluation phase processing using the evaluation model according to the fourth embodiment. This is a diagram showing the functional configuration of the detection system according to the fifth embodiment. This is a flowchart showing an example of the learning image selection process according to the sixth embodiment. This is a flowchart showing an example of the learning image selection process according to the seventh embodiment. This is a flowchart showing an example of the dimensionality reduction algorithm determination process according to the seventh embodiment. This is a flowchart showing an example of the learning image selection process according to the eighth embodiment. This is a schematic block diagram showing the configuration of a computer.

[0015] <First Embodiment> The first embodiment will be described in detail below with reference to Figures 1 to 4.

[0016] (Overall configuration of the detection system) Figure 1 is a diagram showing the functional configuration of the detection system 1 according to the first embodiment. As shown in Figure 1, the detection system 1 comprises a camera 10 and a detection device 20.

[0017] Camera 10 photographs the inside of the waste pit 11 of the plant (waste incineration facility). The waste pit 11 stores the waste that will be incinerated in the incinerator. The waste pit 11 also sends the stored waste to the incinerator using a crane, hopper, chute, etc. (not shown). The images captured by camera 10 include the internal structure of the waste pit 11 and the waste stored in the waste pit 11. The images captured by camera 10 are transmitted to the detection device 20.

[0018] The detection device 20 detects unsuitable materials inside the waste pit 11 based on images (evaluation images) captured by the camera 10. Unsuitable materials are those that cannot be incinerated in the incinerator.

[0019] (Functional configuration of the detection device) As shown in Figure 1, the detection device 20 includes an acquisition unit 201, an extraction unit 202, an evaluation unit 203, a determination unit 204, a control unit 205, an output unit 206, a learning unit 207, and a storage unit 208.

[0020] During the evaluation phase of the waste pit 11, the acquisition unit 201 acquires evaluation images taken by the camera 10 of the inside of the waste pit 11. In addition, during the learning phase of the evaluation model M2, which will be described later, the acquisition unit 201 acquires normal images taken when the inside of the waste pit 11 is in a normal state.

[0021] The extraction unit 202 extracts feature vectors from an image using a pre-trained image recognition model M1. A feature vector is a vector consisting of N-dimensional feature quantities. In the evaluation phase, the extraction unit 202 extracts feature vectors from evaluation images. In the training phase, the extraction unit 202 also extracts feature vectors from normal images.

[0022] The evaluation unit 203 uses an evaluation model M2, which has learned a reference vector obtained by reducing the dimensionality of feature vectors extracted from a normal image taken when the inside of the garbage pit is in a normal state, to calculate the distance between the reference vector and the feature vector extracted from the evaluation image, and evaluates the anomaly score for each region of the evaluation image based on the calculated distance. For example, the evaluation unit 203 evaluates the anomaly score for each pixel of the evaluation image.

[0023] The determination unit 204 determines that an unsuitable object exists in a region if the area of ​​the region where the abnormal score is equal to or greater than the abnormal threshold is equal to or greater than the upper limit area.

[0024] If the control unit 205 determines that unsuitable material is present, it controls the crane according to the operator's instructions to remove the unsuitable material so that it is not mistakenly fed into the incinerator.

[0025] The output unit 206 includes a display device that notifies the operator of the detection of an unsuitable material in text or images when it is determined that an unsuitable material is present. The output unit 206 may also include a speaker that notifies the operator of the detection of an unsuitable material in voice or other means.

[0026] The learning unit 207 learns a reference vector obtained by quantizing a feature vector extracted from a normal image captured by the camera 10 when the waste pit 11 is in a normal state, and creates an evaluation model M2. Quantization includes various quantization methods such as direct product quantization and residual quantization. In other embodiments, the learning unit 207 may also use dimensionality reduction algorithms such as Principal Component Analysis (PCA), advanced UMAP (Uniform Manifold Approximation and Projection), TriMAP, and PaCMAP (Pairwise Controlled Manifold Approximation and Projection) to reduce the dimensionality of the feature vector instead of these quantization methods.

[0027] The memory unit 208 stores the image recognition model M1 and the evaluation model M2, etc. The image recognition model M1 is an existing pre-trained image recognition model such as ResNet or Wide-ResNet.

[0028] (Detection device processing example 1: Learning phase) Figure 2 is a flowchart showing an example of the learning phase processing of the evaluation model according to the first embodiment. Referring to Figure 2, the process flow in which the detection device 20 learns the evaluation model M2 will be explained.

[0029] The acquisition unit 201 acquires multiple normal images taken by the camera 10 when the waste pit 11 is in a normal state (step S101).

[0030] The extraction unit 202 inputs multiple normal images into the image recognition model M1 read from the storage unit 208 and extracts a feature vector (a vector consisting of N-dimensional feature quantities) from each normal image (step S102).

[0031] The learning unit 207 trains an evaluation model M2 based on feature vectors extracted from each normal image. The evaluation model M2 has a memory bank that stores each feature vector.

[0032] As described above, the normal state of the waste pit 11 is highly diverse, so it is necessary to learn a large number of normal images. However, if all the feature vectors extracted from each of the large number of normal images are saved as they are, the calculation time and memory consumption required for the evaluation unit 203 to calculate the distance will become very large, and if this exceeds the memory capacity of the detection device 20, a memory error will occur and it will become impossible to detect unsuitable objects.

[0033] Therefore, in this embodiment, the learning unit 207 first obtains a reference vector by reducing the dimensionality of the feature vector (step S103). Then, the learning unit 207 creates a lightweight evaluation model M2 by learning (storing in a memory bank) the reduced dimensionality reference vector (step S104).

[0034] Specifically, the learning unit 207 reduces the dimensionality of the feature vector by using vector quantization methods such as direct product quantization and residual quantization.

[0035] Figure 3 shows an example of a process that reduces the dimensionality of a feature vector using a vector quantization method. Figure 3 shows an example of the process (step S103) in which the learning unit 207 reduces the dimensionality of a feature vector using a direct product quantization method.

[0036] As shown in Figure 3, the learning unit 207 first divides the feature vector x, which consists of N-dimensional feature quantities extracted from a normal image, into M sub-vectors x1, x2, ..., xM. The learning unit 207 also performs vector quantization on each of the M sub-vectors to obtain a representative point (codeword) for each sub-vector. As a result, the N-dimensional feature vector is reduced in dimension (made less dimensional) to an M-dimensional vector. The learning unit 207 stores this M-dimensional vector as a reference vector in the memory bank. This allows the memory bank to be compressed. The evaluation model M2, including the memory bank, is stored in the storage unit 208.

[0037] Furthermore, as mentioned above, conventional techniques use a greedy method to compress the memory bank. However, with a greedy method, it is necessary to calculate the distance between all vectors, so when attempting to train a large number of normal images, if the amount of memory available during the training phase is small, a memory error occurs and the evaluation model cannot be trained. In contrast, in this embodiment, the feature vector x is first divided, and the dimensionality is reduced by quantizing each divided sub-vector, so the computational cost can be significantly reduced compared to a greedy method. Therefore, even if the amount of memory available to the training unit 207 is small during the training phase, the occurrence of memory errors is suppressed, and it becomes possible to train the evaluation model M2 using a large number of normal images. As a result, the accuracy of detecting unsuitable objects using the evaluation model M2 can be improved.

[0038] (Detection device processing example 2: Evaluation phase) Figure 4 is a flowchart showing an example of the evaluation phase processing using the evaluation model according to the first embodiment. Referring to Figure 4, the processing flow of the detection device 20 evaluating the presence or absence of unsuitable materials using the learned evaluation model M2 will be explained.

[0039] The acquisition unit 201 acquires the evaluation image captured by the camera 10 (step S201).

[0040] The extraction unit 202 inputs the evaluation image to the image recognition model M1 read from the storage unit 208 and extracts a feature vector (a vector consisting of N-dimensional feature quantities) from the evaluation image (step S202).

[0041] The evaluation unit 203 calculates the minimum distance between a reference vector stored in the evaluation model M2 (memory bank) read from the storage unit 208 and a feature vector extracted from the evaluation image by the extraction unit 202 (step S203).

[0042] The evaluation unit 203 also generates an anomaly score map based on the minimum distance between each feature vector and the reference vector (step S204). The anomaly score map is a heatmap in which each region (each pixel) is color-coded according to the anomaly score. The anomaly score map may be output (displayed) by the output unit 206, for example, superimposed on the evaluation image, so that an operator can visually check it.

[0043] Next, the determination unit 204 binarizes the anomaly score map (step S205). For example, the determination unit 204 binarizes the anomaly score map by coloring regions where the anomaly score is equal to or higher than an anomaly threshold white, and coloring regions where the anomaly score is less than the anomaly threshold black. The anomaly threshold may be arbitrarily changed by the operator according to the image quality of the image captured by the camera 10 and the configuration of the waste pit 11. For example, the detection device 20 has a plurality of anomaly thresholds in advance, and the operator selects one of the anomaly thresholds according to the camera 10 and the waste pit 11.

[0044] The determination unit 204 also calculates the area of a spatially continuous region where the anomaly score is equal to or higher than the anomaly threshold, based on the binarized anomaly score map (step S206). If there are a plurality of regions where the anomaly score is equal to or higher than the anomaly threshold, the determination unit 204 calculates the area of each region.

[0045] The determination unit 204 determines whether at least one area calculated in step S206 is equal to or larger than an upper limit area (step S207). The upper limit area is set to a value that allows detection of an object exceeding the size processable in the waste pit 11, based on, for example, the size of a crane, hopper, or the like of the waste pit 11. For example, if the hopper of the waste pit 11 can accept objects up to 1 meter square, the upper limit area is set to a value corresponding to 1 square meter in the evaluation image. The upper limit area may also be set as a percentage (%) relative to the total size of the evaluation image, for example.

[0046] When at least one area is equal to or larger than the upper limit area (step S207: YES), the determination unit 204 determines that an unsuitable object exists in this region (step S208). In this case, the output unit 206 notifies an operator that an unsuitable object has been detected. Further, the control unit 205 controls the crane in accordance with the operation of the operator who has received the notification, and removes the detected unsuitable object (step S209). At this time, the output unit 206 may display an image obtained by superimposing an abnormality score map on the evaluation image, and may highlight (e.g., frame) a region where the unsuitable object is detected. The operator performs the work of confirming and removing the unsuitable object based on the notification from the output unit 206.

[0047] On the other hand, when all areas are less than the upper limit area (step S207: NO), the determination unit 204 determines that no unsuitable object exists in the garbage pit 11 (step S210), and ends the processing for the evaluation image acquired in step S201.

[0048] Note that the operator may refer to the abnormality score map or the like output by the output unit 206, and perform an instruction operation to remove an object in the garbage pit 11 as necessary. In this case, the control unit 205 controls the crane to remove the object in accordance with the instruction from the operator.

[0049] The detection device 20 repeatedly executes the series of processes in FIG. 4 at predetermined intervals while the garbage pit 11 is in operation. Accordingly, when an unsuitable object is stored in the garbage pit 11, it can be promptly detected and removed. If an unsuitable object is accidentally charged into an incinerator, the combustion state becomes unstable, and it may be necessary to shut down the operation of the plant. However, in the present embodiment, since an unsuitable object can be detected and removed in advance as described above, even when an unsuitable object is stored in the garbage pit 11, accidental charging of the unsuitable object into the incinerator is suppressed, and the plant can be operated without being shut down.

[0050] (Effects) As described above, the detection device 20 according to this embodiment includes: an acquisition unit 201 that acquires evaluation images taken inside the garbage pit 11; an extraction unit 202 that extracts feature vectors from the evaluation images using a pre-learned image recognition model M1; an evaluation unit 203 that calculates the distance between the reference vector and the feature vector extracted from the evaluation image using an evaluation model M2 that has learned a reference vector obtained by reducing the dimensionality of the feature vector extracted from a normal image taken when the inside of the garbage pit 11 is in a normal state, and evaluates the abnormality score for each region of the evaluation image based on the calculated distance; and a determination unit 204 that determines that there are unsuitable objects in a region when the area of ​​the region in which the abnormality score is equal to or greater than the abnormal threshold is equal to or greater than the upper limit area.

[0051] The detection device 20 can reduce computation time and memory consumption by using an evaluation model M2 that has learned the reduced-dimensional reference vectors, thereby reducing the computation time and memory consumption of the distance between the reference vector and the feature vector of the evaluation image. Furthermore, because the amount of data is reduced by reducing the dimensionality of the feature vectors, it is possible to use an evaluation model M2 that has learned more normal images, thus achieving both a reduction in computation time and memory consumption for the process of detecting unsuitable objects in the waste pit 11 and an improvement in the accuracy of detecting unsuitable objects. In addition, by detecting unsuitable objects based on area, it is possible to appropriately detect objects that are difficult to process due to the configuration of the waste pit 11 (e.g., cranes, hoppers, etc.).

[0052] Furthermore, the detection device 20 is further equipped with a control unit 205 that removes unsuitable materials if it is determined that such materials are present.

[0053] In this way, the detection device 20 can quickly detect and remove unsuitable materials when they are stored in the waste pit 11. Therefore, even if unsuitable materials are stored in the waste pit 11, it is possible to prevent them from being mistakenly fed into the incinerator and to operate the plant without shutting it down.

[0054] Furthermore, the detection device 20 includes a learning unit 207 that learns a reference vector obtained by quantizing a feature vector extracted from a normal image to create an evaluation model M2.

[0055] As described above, dimensionality reduction of feature vectors using a greedy method requires calculating the distance between all vectors. Therefore, when attempting to train a large number of normal images, if the amount of memory available during the training phase is small, a memory error occurs, making it impossible to train the evaluation model. In contrast, the detection device 20 according to this embodiment, for example, using a Cartesian product quantization method, first divides the feature vector x and then quantizes each divided sub-vector to reduce dimensionality. This significantly reduces computational cost compared to the greedy method. Consequently, even if the amount of memory available to the training unit 207 during the training phase is small, it is possible to suppress the occurrence of memory errors and train the evaluation model M2 using a large number of normal images. This improves the accuracy of detecting unsuitable objects using the evaluation model M2.

[0056] <Second Embodiment> Next, the second embodiment will be described in detail with reference to Figure 5. Components common to the above-described embodiment are denoted by the same reference numerals and their detailed descriptions are omitted.

[0057] In the first embodiment, an example was described in which, in step S203 of the evaluation phase (Figure 4), the evaluation unit 203 calculates the distance between the feature vector extracted from the evaluation image and all the reference vectors stored in the memory bank of the evaluation model M2. In contrast, the evaluation unit 203 in this embodiment calculates the distance between the feature vector extracted from the evaluation image and some of the reference vectors stored in the memory bank of the evaluation model M2.

[0058] For example, the evaluation unit 203 optimizes the combination of vectors used for distance calculation using a vector search method such as Inverted File Flat (IVF). Specifically, the evaluation unit 203 divides the memory bank of the evaluation model M2 into multiple clusters and assigns each reference vector to one of the clusters. In step S203 of the evaluation phase (Figure 4), the evaluation unit 203 identifies clusters of reference vectors that approximate the feature vectors extracted from the evaluation image, and calculates the minimum distance between the reference vectors included in the identified clusters and the feature vectors extracted from the evaluation image. In this way, the detection device 20 can reduce the number of distance calculations and reduce the computation time and memory consumption in the evaluation phase.

[0059] Figure 5 shows an example of the measurement results for the detection accuracy and memory consumption of the detection device according to the second embodiment. Figure 5 also shows a comparison of the measurement results for the detection rate of unsuitable objects [%], the false detection rate [%], and the memory consumption per evaluation image [GB] for the conventional technology and the detection device 20, respectively. As described above, the detection device 20 of this embodiment reduces the dimensionality of the evaluation model M2 (memory bank) by vector quantization and reduces the number of distance calculations using IVF in step S203 of the evaluation phase (Figure 4). On the other hand, the conventional technology does not perform dimensionality reduction by vector quantization and calculates the distance to all feature vectors stored in the memory bank without using IVF in the evaluation phase.

[0060] As shown in Figure 5, the detection device 20 according to the second embodiment was able to significantly reduce memory consumption while maintaining the detection accuracy of unsuitable materials at the same level as the conventional technology. This makes it possible to configure the detection device 20 with an inexpensive machine with small memory. In other words, it is possible to reduce the cost of the detection device 20 and satisfy the requirements for detection performance of unsuitable materials (detection rate of 90% or more, false detection rate of 5% or less).

[0061] <Third Embodiment> Next, the third embodiment will be described in detail with reference to Figures 6 to 9. Components common to the above embodiments are denoted by the same reference numerals and their detailed descriptions are omitted.

[0062] Figure 6 is a diagram illustrating an evaluation model according to the third embodiment. As shown in Figure 6, in the detection device 20 according to this embodiment, the learning unit 207 creates a plurality of evaluation models M2.

[0063] For example, suppose that due to the memory capacity limitations of the detection device 20, the maximum number of normal images that can be loaded into memory and trained during the training phase is 500. As mentioned above, in the waste pit 11, all images that do not contain unsuitable objects are classified as normal images, so a particularly large number of images must be trained. However, with such memory capacity limitations, it is not possible to create an evaluation model M2 that has been trained on more than 500 normal images.

[0064] Therefore, the detection device 20 according to this embodiment divides normal images into multiple image groups and creates multiple evaluation models M2 that have been trained on each image group. In the example in Figure 6, 1000 normal images are divided into two image groups 1 and 2. Each image group 1 and 2 contains 500 normal images. The detection device 20 also creates two evaluation models M2: evaluation model M2_1 that has been trained on image group 1, and evaluation model M2_2 that has been trained on image group 2.

[0065] (Detection device processing example 1: Learning phase) Figure 7 is a flowchart showing an example of the learning phase of an evaluation model according to the third embodiment. Referring to Figure 7, the process flow in which the detection device 20 learns multiple evaluation models M2 will be explained.

[0066] The acquisition unit 201 acquires multiple normal images taken by the camera 10 when the waste pit 11 is in a normal state (step S401).

[0067] Furthermore, the learning unit 207 divides the acquired normal images into multiple image groups (step S402). At this time, the learning unit 207 divides the normal images in such a way that the number of image groups is minimized and the number of normal images contained in each of the multiple image groups is approximately equal.

[0068] Increasing the number of image groups also increases the number of evaluation models M2, which increases the computation time and memory consumption in the evaluation phase described later (steps S503A and S503B in Figure 9). Therefore, the learning unit 207 determines the number of divisions for the image groups so as close as possible to the maximum number that can be loaded into memory. As in the example above, suppose that due to memory capacity limitations, the maximum number of normal images that can be loaded into memory and trained during the learning phase is 500. In this case, since the maximum number of images in an image group is 500, if the total number of normal images is 1000, the learning unit 207 allocates 500 normal images to each of the two image groups 1 and 2.

[0069] Furthermore, if the total number of normal images is 1001, and 500 normal images are assigned to image groups 1 and 2, and 1 image is assigned to image group 3, the accuracy of the evaluation model M2, which is trained using image group 3, will decrease. For this reason, it is desirable to assign normal images to each of image groups 1 to 3 in roughly equal proportions. Therefore, if the total number of normal images is 1001, for example, 334 normal images should be assigned to image groups 1 and 2, and 333 normal images should be assigned to image group 3.

[0070] Furthermore, if, for example, normal images are distributed roughly equally among the image groups, and the number of normal images in each image group falls below a specified number (e.g., 400 images), the accuracy of the evaluation model M2 may decrease. For this reason, it is possible to remove some normal images before dividing them into image groups so that each image group contains at least the specified number of normal images. For example, if the total number of normal images is 1001, one normal image can be removed, and the remaining 1000 normal images can be distributed equally among two image groups, 1 and 2. The operator may specify which normal images to remove, or the learning unit 207 may automatically decide based on a predetermined rule (e.g., removing older normal images first, removing images randomly, etc.).

[0071] Next, the extraction unit 202 selects one of the multiple image groups (for example, image group 1). Then, the extraction unit 202 inputs the normal images included in the selected image group 1 into the image recognition model M1 read from the storage unit 208, and extracts feature vectors from each normal image (step S403).

[0072] The learning unit 207 reduces the dimensionality of the feature vectors extracted from each normal image in image group 1 using a vector quantization method (step S404). The learning unit 207 also creates an evaluation model M2_1 by learning (storing in a memory bank) the reduced dimensionality reference vectors (step S405). The processing in steps S403 to S405 is the same as the processing in steps S102 to S104 of the first embodiment (Figure 2).

[0073] Next, the learning unit 207 determines whether the learning of all image groups has been completed (step S406). For example, if the learning of image group 1 has been completed but the learning of image group 2 has not been completed (step S406; NO), steps S403 to S405 are executed again to create an evaluation model M2_2 that has been learned from the normal images of the next image group 2. On the other hand, if the learning of all image groups has been completed (step S406; YES), the learning unit 207 terminates the learning phase processing. Each evaluation model M2_1 and M2_2 created by the learning unit 207 is stored in the storage unit 208.

[0074] (Detection device processing example 2: Evaluation phase) Figure 8 is a flowchart showing an example of the evaluation phase processing using the evaluation model according to the third embodiment. Referring to Figure 8, the processing flow in which the detection device 20 evaluates the presence or absence of unsuitable materials using multiple evaluation models M2 that have been learned will be explained. Here, an example using two evaluation models M2 (first evaluation model M2_1, second evaluation model M2_2) will be explained.

[0075] The acquisition unit 201 acquires the evaluation image captured by the camera 10 (step S501). The extraction unit 202 inputs the evaluation image into the image recognition model M1 read from the storage unit 208 and extracts feature vectors from the evaluation image (step S502). The processing in steps S501 to S502 is the same as the processing in steps S201 to S202 of the first embodiment (Figure 4).

[0076] Next, the evaluation unit 203 calculates a first minimum distance between the reference vector stored in the first evaluation model M2_1 (memory bank) read from the memory unit 208 and the feature vector extracted from the evaluation image by the extraction unit 202 (step S503A). In parallel with this, the evaluation unit 203 also calculates a second minimum distance between the reference vector stored in the second evaluation model M2_2 (memory bank) read from the memory unit 208 and the feature vector extracted from the evaluation image by the extraction unit 202 (step S503B).

[0077] Furthermore, the evaluation unit 203 selects the smaller of the first minimum distance calculated using the first evaluation model M2_1 and the second minimum distance calculated using the second evaluation model M2_2 (step S504).

[0078] Furthermore, the evaluation unit 203 creates an anomaly score map based on the minimum distance selected in step S504 (step S505).

[0079] Next, the determination unit 204 binarizes the abnormal score map (step S506).

[0080] Furthermore, the determination unit 204 calculates the area of ​​spatially continuous regions where the abnormal score is equal to or greater than the abnormal threshold, based on the binarized abnormal score map (step S507). If there are multiple regions where the abnormal score is equal to or greater than the abnormal threshold, the determination unit 204 calculates the area of ​​each region.

[0081] The determination unit 204 determines whether at least one area calculated in step S206 is equal to or greater than the upper limit area (step S508).

[0082] The determination unit 204 determines that if at least one area is greater than or equal to the upper limit area (step S508; YES), an unsuitable object exists in that area (step S509). In this case, the output unit 206 notifies the operator that an unsuitable object has been detected. The control unit 205 then controls the crane according to the operator's instructions to remove the detected unsuitable object (step S510).

[0083] On the other hand, if all areas are less than the upper limit area (step S508; NO), the determination unit 204 determines that there are no unsuitable materials in the waste pit 11 (step S511) and terminates processing of the evaluation image acquired in step S501.

[0084] Note that the processing in steps S505 to S511 is the same as the processing in steps S204 to S210 of the first embodiment (Figure 4). The detection device 20 repeatedly executes the series of processes shown in Figure 8 at predetermined intervals while the waste pit 11 is in operation.

[0085] Figure 9 shows an example of the measurement results of the detection accuracy of the detection device according to the third embodiment. Figure 9 also shows a comparison of the measurement results of the detection rate [%] of unsuitable objects and the false detection rate [%] in two cases: one using a single evaluation model M2 that has been trained on 500 normal images, and another using two evaluation models M2_1 and M2_2 that have each been trained on 500 normal images.

[0086] As shown in Figure 9, by using two evaluation models M2_1 and M2_2, we were able to slightly improve the detection rate and reduce the false positive rate. In this measurement, 1280 evaluation images were used, so we were able to reduce the number of false positives by 9 compared to using only one evaluation model M2.

[0087] (Effects) As described above, in the detection device 20 according to this embodiment, the learning unit 207 divides a plurality of normal images into a plurality of image groups such that the number of images is less than the upper limit of memory available when training the evaluation model M2, and creates a plurality of evaluation models M2 corresponding to each image group. The evaluation unit 203 evaluates the abnormality score for each region of the evaluation image based on the minimum distance between the reference vector of each of the plurality of evaluation models M2 and the feature vector extracted from the evaluation image.

[0088] In this way, the detection device 20 can learn more normal images without exceeding its memory capacity. This further improves the accuracy of detecting unsuitable materials in the waste pit 11. In particular, as shown in Figure 9, the false detection rate can be reduced. When unsuitable materials are detected in the waste pit 11, the operator needs to perform a visual inspection. Therefore, by reducing the false detection of unsuitable materials, the time required for the operator to check for unsuitable materials can be reduced, thus reducing the human cost of the operator.

[0089] Furthermore, the learning unit 207 divides multiple normal images so that the number of image groups is minimized and the number of normal images in each of the multiple image groups approaches equality.

[0090] As described above, increasing the number of image groups increases the number of evaluation models M2, which increases the computation time and memory consumption in the evaluation phase (steps S503A and S503B in Figure 9). However, the learning unit 207 can minimize the increase in computation time and memory consumption in the evaluation phase by dividing the normal images in a way that minimizes the number of image groups. In addition, the learning unit 207 can prevent an imbalance in the number of normal images in each image group, which would reduce the accuracy of some evaluation models M2, by adjusting the number of normal images in each image group to be roughly equal.

[0091] <Fourth Embodiment> Next, the fourth embodiment will be described in detail with reference to Figures 10 to 12. Components common to the above embodiments are denoted by the same reference numerals and their detailed descriptions are omitted.

[0092] Figure 10 is a diagram illustrating an evaluation model according to the fourth embodiment. As shown in Figure 10, after creating a plurality of evaluation models M2_1 and M2_2 using the method of the third embodiment, the learning unit 207 may perform further learning using additional normal images and add new evaluation models M2_3 and M2_4.

[0093] (Detection device processing example 1: Learning phase) Figure 11 is a flowchart showing an example of the learning phase processing of the evaluation model according to the fourth embodiment. Referring to Figure 11, the process flow in which the detection device 20 additionally learns the new evaluation model M2 will be explained.

[0094] In this embodiment, the learning unit 207 can increase the number of evaluation models M2 up to the limit of the memory available to the detection device 20 during the evaluation phase. Therefore, the learning unit 207 first determines whether it is possible to add evaluation models M2 based on the upper limit of memory available during the evaluation phase (step S601). For example, the learning unit 207 determines that it is possible to add evaluation models M2 if the difference between the total memory consumption of multiple evaluation models M2 and the upper limit of memory available to the evaluation unit 203 when performing evaluation using the evaluation models M2 is greater than the memory consumption of one evaluation model M2 (step S601; YES). In this case, the learning unit 207 executes the processes in steps S602 to S607 to create and add a new evaluation model M2 based on the newly added normal image. The processes in steps S602 to S607 are the same as the processes in steps S401 to S406 of the third embodiment (Figure 7).

[0095] On the other hand, the learning unit 207 determines that it is not possible to add an evaluation model M2 if the difference between the total memory consumption of multiple evaluation models M2 and the upper limit of available memory of the evaluation unit 203 is smaller than the memory consumption of a single evaluation model M2 (step S601; NO). In this case, the learning unit 207 does not create a new evaluation model M2 and terminates the process.

[0096] (Detection device processing example 2: Evaluation phase) Figure 12 is a flowchart showing an example of the evaluation phase processing using the evaluation model according to the fourth embodiment. The processing of steps S701 to S702 and S705 to S711 of the detection device 20 in this embodiment is the same as the processing of steps S501 to S502 and S505 to S511 of the third embodiment (Figure 8). As shown in Figure 12, in steps S703A to S704D, the evaluation unit 203 calculates the minimum distance from the reference vectors of the previously created evaluation models M2_1 and M2_2, and also calculates the minimum distance from the reference vectors of the newly added evaluation models M2_3 and M2_4. At this time, since each evaluation model M2 is simultaneously expanded into memory, in the learning phase (step S601 in Figure 11), the total memory consumption of the evaluation models M2 is adjusted so that it is less than the upper limit of memory that the evaluation unit 203 can use in the evaluation phase.

[0097] Furthermore, the evaluation unit 203 selects the smallest distance among the minimum distances calculated using each evaluation model M2_1 to M2_4 (step S704). The subsequent processing is the same as in the third embodiment, so a detailed explanation is omitted.

[0098] (Effects) As described above, in the detection device 20 according to this embodiment, the learning unit 207 creates and adds a new evaluation model M2 based on a newly added set of normal images when the difference between the total memory consumption of the multiple evaluation models M2 and the upper limit of memory available to the evaluation unit 203 when performing evaluation using the evaluation models M2 is greater than the memory consumption of a single evaluation model M2.

[0099] In this way, the detection device 20 can learn from a large number of normal images newly acquired during the operation of the waste pit 11, thereby further improving the accuracy of detecting unsuitable materials. Furthermore, the detection device 20 adds the evaluation model M2 in a way that does not exceed the upper limit of available memory during the evaluation phase, making it possible to improve the accuracy of detecting unsuitable materials while suppressing the occurrence of memory errors during the evaluation phase.

[0100] <Fifth Embodiment> In the fourth embodiment, in order to avoid affecting the detection process of unsuitable materials in the waste pit 11, for example, only the evaluation phase process (Figure 12) may be executed while the waste pit 11 is in operation, and the learning phase process (Figure 11) may be performed when the plant is shut down, such as at night or on holidays (a period when the evaluation phase is not executed). Also, if the memory capacity of the detection device 20 is sufficiently large, the evaluation phase and the learning phase may be executed in parallel.

[0101] Figure 13 is a diagram showing the functional configuration of the detection system according to the fifth embodiment. As shown in Figure 13, the detection system 1 may also consist of two computers: a detection device 20 that performs only the detection phase and a learning device 30 that performs only the learning phase.

[0102] The learning device 30 comprises an acquisition unit 301, an extraction unit 302, a learning unit 307, and a storage unit 308. These functions are the same as those of the acquisition unit 201, extraction unit 202, learning unit 207, and storage unit 208 in the above-described embodiment.

[0103] Furthermore, the learning device 30 transmits the created (added) evaluation model M2 to the detection device 20. The detection device 20 stores the evaluation model M2 received from the learning device 30 in the storage unit 208.

[0104] In this way, the evaluation phase and the learning phase can be executed in parallel while the waste pit 11 is in operation. As a result, the detection device 20 can acquire the newly added evaluation model M2, for example, at the next startup, and use it to detect unsuitable materials. This makes it possible to further improve the accuracy of detecting unsuitable materials.

[0105] Furthermore, the configuration shown in Figure 13 may be applied not only to the fourth embodiment but also to the first to third embodiments.

[0106] <Sixth Embodiment> Next, the sixth embodiment will be described in detail with reference to Figure 14. Components common to the embodiments described above are denoted by the same reference numerals and their detailed descriptions are omitted.

[0107] The number of normal images that can be used for training the evaluation model M2 is limited by the amount of memory the machine has. Due to this limitation, training images that are similar to each other is inefficient. Furthermore, if there is a bias in the training images, it will lead to a decrease in the accuracy of the evaluation model M2. For this reason, the detection device 20 of this embodiment performs a process to select training images (normal images) that will serve as the training dataset for the evaluation model M2 before the training phase of the evaluation model M2 in any one of the embodiments described above (for example, Figure 2, Figure 7, or Figure 11).

[0108] In this embodiment, an example is described in which the detection device 20 performs the learning image selection process described below. In other embodiments, the learning device 30 may perform the learning image selection process instead of the detection device 20. In this case, the detection device 20, acquisition unit 201, extraction unit 202, and learning unit 207 in the following description shall be interpreted as being replaced by the learning device 30, acquisition unit 301, extraction unit 302, and learning unit 307, respectively.

[0109] (Selection Process for Learning Images) Figure 14 is a flowchart showing an example of the selection process for learning images according to the sixth embodiment. The learning unit 207 selects a learning dataset (normal images) to be used in the next learning phase by performing the series of processes shown in Figure 14 on a plurality of normal images acquired and stored by the acquisition unit 201 at a predetermined timing. The predetermined timing can be set arbitrarily, for example, at a time specified by the operator, after a certain period of time has elapsed since the previous selection process, or when a predetermined number of normal images have been stored.

[0110] First, the extraction unit 202 inputs multiple normal images to the image recognition model M1 and extracts a feature vector from each normal image (step S801). This process is the same as, for example, step S102 in Figure 2. One feature vector is extracted for each normal image.

[0111] Next, the learning unit 207 classifies the feature vectors of each normal image into multiple clusters (step S802). The learning unit 207 performs a known non-hierarchical clustering method, such as k-means or k-means++. At this time, the learning unit 207 performs clustering while changing the number of clusters and adopts the clustering result with the largest number of clusters that maximizes the silhouette coefficient after clustering. In this way, the number of clusters can be determined robustly even if the appearance or number of normal images changes.

[0112] As a result, feature vectors of similar normal images are stored within each cluster. To prevent similar normal images from being learned redundantly, the learning unit 207 selects one cluster at a time from the N clusters and performs the following steps S804 to S806 for the selected cluster i (step S803).

[0113] Cluster i contains M feature vectors. The learning unit 207 executes the processes in steps S805 to S806 to delete feature vectors until the number of vectors M in cluster i reaches the upper limit Mr (step S804). The upper limit Mr is the number of vectors to remain in each cluster. The maximum number of normal images that can be learned is determined by the memory capacity of the detection device 20. Therefore, the upper limit Mr for each cluster is determined so that the total number of normal images acquired in step S801 is equal to the maximum number of learnable images determined according to the memory capacity. The upper limit Mr for each cluster is determined in proportion to the number of vectors initially classified into each cluster (initial number of vectors). For example, if the number of normal images is reduced by half, the upper limit Mr for each cluster will be half the value of the initial number of vectors for each cluster.

[0114] Specifically, the learning unit 207 calculates the inter-vector distance d for each of the M feature vectors x1, x2, x3, ..., xM contained in cluster i (step S805).

[0115] Next, the learning unit 207 deletes one of the combinations of vectors that minimizes the distance d between them (step S806). A small distance between feature vectors indicates that the images are similar. Therefore, by deleting vectors in this way, one of the images that are similar to each other can be deleted. For example, suppose the distance d12 between feature vector x1 and feature vector x2 is minimized. In this case, the learning unit 207 deletes either feature vector x1 or feature vector x2. Which one to delete may follow a predetermined order (for example, the one with the older shooting date) or it may be random. As a result, the number of feature vectors included in cluster i becomes M = M - 1. The learning unit 207 repeats steps S805 to S806 until the number of feature vectors included in cluster i reaches the upper limit Mr.

[0116] Furthermore, when the number of vectors in a given cluster is reduced to the upper limit Mr, the learning unit 207 selects the next cluster i and repeats the above process.

[0117] Furthermore, in the learning phase of any one of the embodiments described above, the acquisition unit 201 acquires normal images included in the learning image dataset created by the learning unit 207 (step S101 in Figure 2, step S401 in Figure 7, and step S602 in Figure 11). Therefore, in the learning phase, the learning unit 207 learns a reference vector obtained by reducing the dimensionality of the feature vectors of the normal images included in the learning image dataset (steps S103 to S104 in Figure 2, steps S404 to S405 in Figure 7, and steps S605 to S606 in Figure 11).

[0118] (Effects) As described above, in the detection device 20 according to this embodiment, the learning unit 207 classifies the feature vectors extracted from each of the multiple normal images into multiple clusters, creates a learning image dataset by deleting feature vectors in order of increasing distance between feature vectors in each cluster, and learns a reference vector which is a low-dimensional representation of the feature vectors of the normal images included in the learning image dataset.

[0119] In this way, the learning unit 207 can create a training image dataset by reducing the number of similar normal images until the number of images can be processed by the memory capacity of the detection device 20. Furthermore, by removing similar normal images in this manner, the learning unit 207 can suppress the bias of normal images included in the training image dataset and learn a diverse range of normal images. This improves the accuracy of the evaluation model M2.

[0120] <Seventh Embodiment> Next, the seventh embodiment will be described in detail with reference to Figure 15. Components common to the above embodiments are denoted by the same reference numerals and their detailed descriptions are omitted.

[0121] In the sixth embodiment, depending on the image recognition model M1 used by the extraction unit 202, the dimensionality of the feature vector extracted from a single normal image may become enormous, for example, several hundred thousand or more. In this case, the memory of the detection device 20 becomes insufficient for the amount of data to be processed, and the number of images that the learning unit 207 can handle in a single selection process becomes small. For this reason, in this embodiment, a process is added to reduce the dimensionality of the feature vector extracted by the extraction unit 202 before the learning unit 207 performs clustering.

[0122] (Training image selection process) Figure 15 is a flowchart showing an example of the training image selection process according to the seventh embodiment.

[0123] First, the extraction unit 202 performs the same process as in step S801 of the sixth embodiment (Figure 14) to extract feature vectors from each normal image (step S901).

[0124] Next, the learning unit 207 reduces the dimensionality of the extracted feature vectors (step S902). This dimensionality reduction process is the same as the process in the learning phase in which the learning unit 207 reduces the dimensionality of the feature vectors (step S103 in Figure 2, step S404 in Figure 7, and step S605 in Figure 11). The learning unit 207 may use quantization methods such as the above-mentioned product quantization or residual quantization as the dimensionality reduction algorithm for reducing the dimensionality of the feature vectors, or it may use other dimensionality reduction algorithms such as PCA, UMAP, TriMAP, or PaCMAP.

[0125] The learning unit 207 classifies the reduced-dimensional feature vectors into multiple clusters (step S903). This process is the same as step S802 in the sixth embodiment (Figure 14). Furthermore, the processes in steps S904 to S907 are the same as steps S803 to S806 in the sixth embodiment (Figure 14), so their explanation is omitted.

[0126] (Process for determining the dimensionality reduction algorithm) Figure 16 is a flowchart showing an example of the process for determining the dimensionality reduction algorithm according to the seventh embodiment. Before actually performing the selection process (Figure 15), the learning unit 207 may evaluate each dimensionality reduction algorithm using the process shown in Figure 16 and present it as an indicator for determining the algorithm to be used in step S902 of Figure 15.

[0127] First, the extraction unit 202 selects a predetermined number (N images) of test images from the normal images and extracts the feature vectors of each test image (step S911). Here, in order to reduce processing load, instead of using all the normal images, only a portion of the normal images are used.

[0128] Next, the learning unit 207 performs the k-nearest neighbor method on the feature vectors of each test image to obtain the top k vector indices closest to the vector j (j = 1, ..., N) (step S912).

[0129] Furthermore, the learning unit 207 reduces the dimensionality of the feature vectors of each test image using each dimensionality reduction algorithm (step S913). For example, when evaluating three dimensionality reduction algorithms A, B, and C, the learning unit 207 creates a feature vector reduced in dimensionality by dimensionality reduction algorithm A, a feature vector reduced in dimensionality by dimensionality reduction algorithm B, and a feature vector reduced in dimensionality by dimensionality reduction algorithm C, respectively.

[0130] Next, the learning unit 207 performs the k-nearest neighbor method on the feature vectors reduced in dimensionality by each dimensionality reduction algorithm to obtain the top k vector indices closest to the vector j (j = 1, ..., N) (step S914). When evaluating the three dimensionality reduction algorithms A, B, and C as described above, the learning unit 207 obtains three sets of vector indices: the top k vector indices of the feature vector reduced in dimensionality by dimensionality reduction algorithm A, the top k vector indices of the feature vector reduced in dimensionality by dimensionality reduction algorithm B, and the top k vector indices of the feature vector reduced in dimensionality by dimensionality reduction algorithm C.

[0131] The learning unit 207 compares the top k vector indices in the feature vector before dimensionality reduction (vector indices obtained in step S912) with the top k vector indices in the feature vector after dimensionality reduction using each dimensionality reduction algorithm (vector indices obtained in step S914), and calculates an evaluation value for each dimensionality reduction algorithm (step S915).

[0132] Specifically, the learning unit 207 considers each vector j to be a match if the vector index obtained before and after dimensionality reduction contains the same number, and calculates the match rate for each vector using the formula "match rate = number of matches / k". Here, if the top k elements contain the same vector index, it is considered a match even if the order has changed.

[0133] The learning unit 207 then calculates the average of the matching rates for each vector j (j=1, ..., N) as the overall matching rate (evaluation value) of the dimensionality reduction algorithm.

[0134] The detection device 20 presents the operator with the evaluation value (agreement rate with the state before dimensionality reduction) of each dimensionality reduction algorithm, and the operator determines the dimensionality reduction algorithm to be used in step S902 of Figure 15, taking into account the agreement rate and memory consumption. The optimal dimensionality reduction algorithm may change depending on the content of the image. However, in this embodiment, by evaluating each algorithm using a portion of an actual normal image, the training image selection process can be performed using the optimal dimensionality reduction algorithm according to the environment in which the detection device 20 is operated. This makes it possible to create an optimal training image dataset, thereby improving the performance of the evaluation model M2.

[0135] (Effects) As described above, in the detection device 20 according to this embodiment, the learning unit 207 reduces the dimensionality of the feature vectors extracted from each of the multiple normal images using a dimensionality reduction algorithm, and classifies the reduced dimensionality of the feature vectors into multiple clusters.

[0136] In this way, the learning unit 207 can reduce the dimensionality of feature vectors extracted from normal images before clustering, even if the dimensionality of the feature vectors is enormous. This prevents the learning unit 207 from reducing the number of vectors (number of images) that can be processed in the clustering and vector selection process for each cluster (steps S903 to S907). This improves the efficiency of selecting normal images to be used in the training image dataset.

[0137] <Eighth Embodiment> Next, the eighth embodiment will be described in detail with reference to Figure 17. Components common to the embodiments described above are denoted by the same reference numerals and their detailed descriptions are omitted.

[0138] While the detection device 20 is in operation, new images that could become training image candidates (normal image candidates) may be acquired. The amount of normal images included in the training image dataset is enormous, and it is difficult for an operator, for example, to compare a new image with a normal image already included in the training image dataset to determine whether it is similar to the new image and decide whether it should be added as a training image. For this reason, in this embodiment, the detection device 20 automatically determines whether to include newly acquired training image candidates (normal image candidates) in the training image selection process.

[0139] (Training image selection process) Figure 17 is a flowchart showing an example of the training image selection process according to the eighth embodiment.

[0140] For example, the acquisition unit 201 acquires and saves a new image taken when the inside of the waste pit 11 is in a normal state as a candidate for a normal image (step S1001). This candidate for a normal image is, for example, an evaluation image in which it was determined that there were no unsuitable materials in the evaluation phase. Alternatively, this candidate for a normal image may be an image taken when the operator confirmed that the inside of the waste pit 11 was in a normal state (i.e., an image specified by the operator).

[0141] If the detection device 20 acquires (stores) a predetermined number or more of normal image candidates, it will determine whether to create a training image dataset including these normal image candidates by processing in steps S1002 to S1004.

[0142] Specifically, first, the extraction unit 202 extracts the feature vectors of both the normal images (existing normal images) included in the training image dataset and the candidate normal images (step S1002).

[0143] Next, the learning unit 207 classifies the feature vectors of each candidate normal image and the feature vectors of each existing normal image into multiple clusters (step S1003). The clustering method is the same as in step S802 of the sixth embodiment (Figure 14).

[0144] Furthermore, the learning unit 207 determines whether the number of clusters, including the normal image candidates, has increased by a predetermined threshold or more compared to the number of clusters before including the normal image candidates (i.e., when the learning image dataset was created last time) (step S1004). For this reason, in this embodiment, the learning unit 207 is assumed to have recorded the number of clusters when the learning image dataset was created last time.

[0145] If the number of clusters does not increase by more than a threshold when normal image candidates are included (step S1004; NO), the learning unit 207 determines that the normal image candidates do not contribute to the increase in clusters, that is, they do not contain diverse images, and terminates processing without using these normal image candidates as candidates for the training image dataset.

[0146] On the other hand, if the number of clusters increases by more than a threshold when normal image candidates are included (step S1004; YES), the learning unit 207 creates a training image dataset based on the dataset including normal image candidates. In other words, the learning unit 207 performs a series of processes from steps S1005 to S1008 for each cluster that has been classified by the feature vectors of the normal image candidates and existing normal images. Since these processes from steps S1005 to S1008 are the same as the processes from steps S803 to S806 in the sixth embodiment (Figure 14), their explanation is omitted. In other embodiments, the learning unit 207 may add a process to reduce the dimensionality of the feature vectors before performing clustering (step S1003), as in the seventh embodiment (Figure 15).

[0147] (Effects) As described above, in the detection device 20 according to this embodiment, if the number of clusters obtained by classifying the feature vectors extracted from a new image which is a candidate for a normal image and the feature vectors extracted from each of the multiple normal images included in the learning image dataset into multiple clusters (second clusters) is greater than the number of clusters obtained by classifying only the feature vectors extracted from each of the multiple normal images included in the learning image dataset into multiple clusters (first clusters), then the learning unit 207 creates a new learning image dataset by deleting feature vectors in order of increasing distance between feature vectors in each of the multiple second clusters.

[0148] In this way, the learning unit 207 can automatically determine whether the new normal image candidates include sufficiently diverse images (i.e., whether they may include images that are not similar to existing normal images) each time new normal image candidates are accumulated, and create a training image dataset. In other words, the learning unit 207 can create a diverse training image dataset suitable for training the evaluation model M2 by including newly acquired images as needed.

[0149] <Other Embodiments> In the embodiments described above, an example was described in which the detection device 20 and the detection system 1 are used to learn normal images of the inside of the waste pit 11 and to detect unsuitable items inside the waste pit 11, but the embodiments are not limited to this.

[0150] For example, in other embodiments, the detection device 20 and detection system 1 may be used to learn the state of the seabed and detect specific features from the seabed. In this case, various sensors such as a visible light camera, an infrared camera, or an acoustic sonar may be used as the camera 10. The learning unit 207 of the detection device 20 learns sensor images that do not contain the features to be detected as normal images. Because the appearance of the seabed varies greatly even in normal images due to the presence of suspended matter and turbidity in the water, as well as differences in the attenuation rates of light and sound, it is necessary to learn a large number of normal images of various patterns. Therefore, the detection device 20 and detection system 1 of each embodiment described above are effective in learning a vast number of normal images of the seabed while suppressing increases in memory consumption and computation time, and in generating an evaluation model M2 that can accurately detect the features to be detected.

[0151] <Computer Configuration> Figure 18 is a schematic block diagram showing the configuration of the computer. The computer 900 comprises a processor 901, main memory 902, auxiliary memory 903, and interface 904. The detection device 20 and learning device 30 described above are each implemented in the computer 900. The operation of each processing unit described above is stored in the auxiliary memory 903 in the form of a program. The processor 901 reads the program from the auxiliary memory 903, expands it into the main memory 902, and executes the above processing according to the program. The processor 901 also allocates memory area in the main memory 902 to be used for the above processing according to the program.

[0152] The program may be for implementing a part of the functions to be performed by the computer 900. For example, the program may perform functions in combination with other programs already stored in the auxiliary storage device 903, or in combination with other programs implemented in other devices. In other embodiments, the computer may be equipped with a custom LSI (Large Scale Integrated Circuit) such as a PLD (Programmable Logic Device) in addition to or instead of the above configuration. Examples of PLDs include PAL (Programmable Array Logic), GAL (Generic Array Logic), CPLD (Complex Programmable Logic Device), FPGA (Field Programmable Gate Array), etc. In this case, some or all of the functions implemented by the processor may be implemented by the integrated circuit.

[0153] Examples of auxiliary storage devices 903 include HDDs (Hard Disk Drives), SSDs (Solid State Drives), magnetic disks, magneto-optical disks, CD-ROMs (Compact Disc Read Only Memory), DVD-ROMs (Digital Versatile Disc Read Only Memory), and semiconductor memory. The auxiliary storage device 903 may be an internal medium directly connected to the bus of the computer 900, or it may be an external medium (external storage device 910) connected to the computer 900 via an interface 904 or a communication line. Furthermore, if this program is distributed to the computer 900 via a communication line, the computer 900 that receives the distribution may expand the program into the main memory 902 and execute the above processing. In at least one embodiment, the auxiliary storage device 903 is a tangible storage medium that is not temporary.

[0154] <Note> The above-described embodiment can be understood, for example, as follows.

[0155] (1) According to the first embodiment, the detection device 20 includes an acquisition unit 201 that acquires evaluation images taken inside the garbage pit 11, an extraction unit 202 that extracts feature vectors from the evaluation images using a pre-learned image recognition model M1, an evaluation unit 203 that calculates the distance between the reference vector and the feature vector extracted from the evaluation image using an evaluation model M2 that has learned a reference vector obtained by reducing the dimensionality of the feature vector extracted from a normal image taken when the inside of the garbage pit 11 is in a normal state, and evaluates the abnormality score for each region of the evaluation image based on the calculated distance, and a determination unit 204 that determines that there is an unsuitable object in a region when the area of ​​the region in which the abnormality score is equal to or greater than the abnormal threshold becomes equal to or greater than the upper limit area.

[0156] The detection device 20 can reduce computation time and memory consumption by using an evaluation model M2 that has learned the reduced-dimensional reference vectors, thereby reducing the computation time and memory consumption of the distance between the reference vector and the feature vector of the evaluation image. Furthermore, because the amount of data is reduced by reducing the dimensionality of the feature vectors, it is possible to use an evaluation model M2 that has learned more normal images, thus achieving both a reduction in computation time and memory consumption for the process of detecting unsuitable objects in the waste pit 11 and an improvement in the accuracy of detecting unsuitable objects. In addition, by detecting unsuitable objects based on area, it is possible to appropriately detect objects that are difficult to process due to the configuration of the waste pit 11 (e.g., cranes, hoppers, etc.).

[0157] (2) According to the second embodiment, the detection device 20 according to the first embodiment further comprises a control unit 205 that removes unsuitable substances when it is determined that unsuitable substances are present.

[0158] In this way, the detection device 20 can quickly detect and remove unsuitable materials when they are stored in the waste pit 11. Therefore, even if unsuitable materials are stored in the waste pit 11, it is possible to prevent them from being mistakenly fed into the incinerator and to operate the plant without shutting it down.

[0159] (3) According to the third embodiment, in the detection device 20 according to the first or second embodiment, the evaluation unit 203 identifies clusters of reference vectors that approximate the feature vectors extracted from the evaluation image, and calculates the distance between the reference vectors included in the identified clusters and the feature vectors extracted from the evaluation image.

[0160] In this way, the detection device 20 can reduce the number of distance calculations and lower the calculation time and memory consumption during the evaluation phase. Therefore, the detection device 20 can be configured with an inexpensive machine with small memory while maintaining the detection accuracy of unsuitable objects.

[0161] (4) According to the fourth aspect, the detection device 20 according to any one of the first to third aspects further comprises a learning unit 207 that learns a reference vector obtained by quantizing a feature vector extracted from a normal image to create an evaluation model M2.

[0162] As described above, dimensionality reduction of feature vectors using a greedy method requires calculating the distance between all vectors. Therefore, when attempting to train a large number of normal images, if the amount of memory available during the training phase is small, a memory error occurs, making it impossible to train the evaluation model. In contrast, the detection device 20 according to this embodiment, for example, using a Cartesian product quantization method, first divides the feature vector x and then quantizes each divided sub-vector to reduce dimensionality. This significantly reduces computational cost compared to the greedy method. Consequently, even if the amount of memory available to the training unit 207 during the training phase is small, it is possible to suppress the occurrence of memory errors and train the evaluation model M2 using a large number of normal images. This improves the accuracy of detecting unsuitable objects using the evaluation model M2.

[0163] (5) According to the fifth aspect, in the detection device 20 according to the fourth aspect, the learning unit 207 classifies the feature vectors extracted from each of the multiple normal images into multiple clusters, creates a learning image dataset by deleting feature vectors in order of increasing distance between feature vectors in each cluster, and learns a reference vector which is a low-dimensional representation of the feature vectors of the normal images included in the learning image dataset.

[0164] In this way, the learning unit 207 can create a training image dataset by reducing the number of similar normal images until the number of images can be processed by the memory capacity of the detection device 20. Furthermore, by removing similar normal images in this manner, the learning unit 207 can suppress the bias of normal images included in the training image dataset and learn a diverse range of normal images. This improves the accuracy of the evaluation model M2.

[0165] (6) According to the sixth aspect, in the detection device 20 according to the fifth aspect, the learning unit 207 reduces the dimensionality of the feature vectors extracted from each of the multiple normal images using a dimensionality reduction algorithm, and classifies the reduced dimensionality of the feature vectors into multiple clusters.

[0166] In this way, the learning unit 207 can reduce the dimensionality of feature vectors extracted from normal images before clustering, even if the dimensionality of the feature vectors is enormous. This prevents the learning unit 207 from reducing the number of vectors (number of images) that can be processed in the clustering and vector selection process for each cluster (steps S903 to S907). This improves the efficiency of selecting normal images to be used in the training image dataset.

[0167] (7) According to the seventh aspect, in the detection device 20 according to the fifth or sixth aspect, if the number of clusters obtained by classifying the feature vectors extracted from a new image which is a candidate for a normal image and the feature vectors extracted from each of the multiple normal images included in the learning image dataset into multiple second clusters is greater than the number of clusters obtained by classifying only the feature vectors extracted from each of the multiple normal images included in the learning image dataset into multiple first clusters, then the learning unit 207 creates a new learning image dataset by deleting feature vectors in each of the multiple second clusters in order of increasing distance between feature vectors.

[0168] In this way, the learning unit 207 can automatically determine whether the new normal image candidates include sufficiently diverse images (i.e., whether they may include images that are not similar to existing normal images) each time new normal image candidates are accumulated, and create a training image dataset. In other words, the learning unit 207 can create a diverse training image dataset suitable for training the evaluation model M2 by including newly acquired images as needed.

[0169] (8) According to the eighth aspect, in the detection device 20 according to any one of the fourth to seventh aspects, the learning unit 207 divides a plurality of normal images into a plurality of image groups such that the number of images is less than the upper limit of memory available when training the evaluation model M2, and creates a plurality of evaluation models M2 corresponding to each image group, and the evaluation unit 203 evaluates the abnormality score for each region of the evaluation image based on the minimum distance between the reference vector of each of the plurality of evaluation models and the feature vector extracted from the evaluation image.

[0170] In this way, the detection device 20 can learn more normal images without exceeding its memory capacity. This further improves the accuracy of detecting unsuitable materials in the waste pit 11. It also reduces the false detection rate. When unsuitable materials are detected in the waste pit 11, the operator needs to perform a visual inspection. Therefore, by reducing the false detection of unsuitable materials, the time required for the operator to check for unsuitable materials can be reduced, thus reducing the human cost of the operator.

[0171] (9) According to the ninth aspect, in the detection device 20 according to the eighth aspect, the learning unit 207 divides a plurality of normal images so that the number of image groups is minimized and the number of normal images contained in each of the plurality of image groups approaches equality.

[0172] The detection device 20 minimizes the increase in computation time and memory consumption during the evaluation phase by dividing normal images in such a way that the number of image groups is minimized. Furthermore, by adjusting the number of normal images in each image group to be approximately equal, the detection device 20 can prevent an imbalance in the number of normal images in each image group, which would reduce the accuracy of some evaluation models M2.

[0173] (10) According to the tenth aspect, in the detection device 20 according to the eighth or ninth aspect, the learning unit 207 creates and adds a new evaluation model M2 based on a newly added set of normal images when the difference between the total memory consumption of a set of evaluation models M2 and the upper limit of memory available to the evaluation unit 203 when performing evaluation using the evaluation models M2 is greater than the memory consumption of a single evaluation model M2.

[0174] In this way, the detection device 20 can learn from a large number of normal images newly acquired during the operation of the waste pit 11, thereby further improving the accuracy of detecting unsuitable materials. Furthermore, the detection device 20 adds the evaluation model M2 in a way that does not exceed the upper limit of available memory during the evaluation phase, making it possible to improve the accuracy of detecting unsuitable materials while suppressing the occurrence of memory errors during the evaluation phase.

[0175] (11) According to the eleventh embodiment, the detection system 1 comprises a detection device 20 according to any one of the first to third embodiments, and a learning device 30 equipped with a learning unit 307 that learns reference vectors obtained by quantizing feature vectors extracted from a normal image to create an evaluation model M2.

[0176] Thus, the detection device 20, which performs only the evaluation phase, and the learning device 30, which performs only the learning phase, may be configured using different computers. By doing so, the memory capacity of the detection device 20 and the learning device 30 can be appropriately set according to the memory consumption of the evaluation phase and the learning phase, respectively.

[0177] (12) According to the 12th aspect, in the detection system 1 according to the 11th aspect, the learning unit 307 classifies the feature vectors extracted from each of the multiple normal images into multiple clusters, creates a learning image dataset by deleting feature vectors in order of increasing distance between feature vectors in each cluster, and learns a reference vector which is a low-dimensional representation of the feature vectors of the normal images included in the learning image dataset.

[0178] In this way, the learning unit 307 can create a training image dataset by reducing the number of similar normal images until the number of images can be processed with the memory capacity of the learning device 30. Furthermore, by deleting similar normal images in this manner, the learning unit 307 can suppress the bias of normal images included in the training image dataset and learn a diverse range of normal images. This improves the accuracy of the evaluation model M2.

[0179] (13) According to the 13th aspect, in the detection system 1 according to the 12th aspect, the learning unit 307 reduces the dimensionality of the feature vectors extracted from each of the multiple normal images using a dimensionality reduction algorithm, and classifies the reduced dimensionality of the feature vectors into multiple clusters.

[0180] In this way, the learning unit 307 can reduce the dimensionality of feature vectors extracted from normal images before clustering, even if the dimensionality of the feature vectors is enormous. This prevents the learning unit 307 from reducing the number of vectors (number of images) that can be processed in the clustering and vector selection process for each cluster (steps S903 to S907). This improves the efficiency of selecting normal images to be used in the training image dataset.

[0181] (14) According to the 14th aspect, in the detection system 1 according to the 12th or 13th aspect, if the number of clusters obtained by classifying the feature vectors extracted from a new image which is a candidate for a normal image and the feature vectors extracted from each of the multiple normal images included in the learning image dataset into multiple second clusters is greater than the number of clusters obtained by classifying only the feature vectors extracted from each of the multiple normal images included in the learning image dataset into multiple first clusters, then the learning unit 307 creates a new learning image dataset by deleting feature vectors in order of increasing distance between feature vectors in each of the multiple second clusters.

[0182] In this way, the learning unit 307 can automatically determine whether the new normal image candidates include sufficiently diverse images (i.e., whether they may include images that are not similar to existing normal images) each time new normal image candidates are accumulated, and create a training image dataset. In other words, the learning unit 307 can create a diverse training image dataset suitable for training the evaluation model M2 by including newly acquired images as needed.

[0183] (15) According to the 15th embodiment, in the detection system 1 according to any one embodiment from the 11th to the 14th, the learning unit 307 divides a plurality of normal images into a plurality of image groups such that the number of images is less than the upper limit of memory available when training the evaluation model M2, and creates a plurality of evaluation models M2 corresponding to each image group, and the evaluation unit 203 evaluates the anomaly score for each region of the evaluation image based on the minimum distance between the reference vector of each of the plurality of evaluation models M2 and the feature vector extracted from the evaluation image.

[0184] In this way, the detection system 1 can learn more normal images without exceeding the memory capacity of the learning device 30. This allows the detection device 20 to further improve the accuracy of detecting unsuitable materials in the waste pit 11 and reduce the false detection rate.

[0185] (16) According to the sixteenth aspect, in the detection system 1 according to the fifteenth aspect, the learning unit 307 divides a plurality of normal images so that the number of image groups is minimized and the number of normal images contained in each of the plurality of image groups approaches equality.

[0186] The learning device 30 minimizes the increase in computation time and memory consumption during the evaluation phase of the detection device 20 by dividing normal images in such a way that the number of image groups is minimized. Furthermore, by adjusting the number of normal images in each image group to be roughly equal, the learning device 30 can prevent an imbalance in the number of normal images in each image group, which would reduce the accuracy of some evaluation models M2.

[0187] (17) According to the 17th aspect, in the detection system 1 according to the 15th or 16th aspect, the learning unit 307 creates and adds a new evaluation model M2 based on a newly added set of normal images when the difference between the total capacity of a set of evaluation models M2 and the upper limit of memory available to the evaluation unit 203 when performing evaluation using the evaluation models M2 is greater than the capacity of a single evaluation model M2.

[0188] In this way, the learning device 30 can further improve the detection accuracy of the detection device 20 by adding an evaluation model M2 that has been learned from a large number of normal images newly acquired during the operation of the waste pit 11. Furthermore, since the learning device 30 adds the evaluation model M2 in a manner that does not exceed the upper limit of memory available during the evaluation phase of the detection device 20, it is possible to improve the detection accuracy of unsuitable materials while suppressing the occurrence of memory errors during the evaluation phase.

[0189] (18) According to the 18th aspect, the detection method includes the steps of: acquiring an evaluation image taken inside the garbage pit 11; extracting feature vectors from the evaluation image using a pre-trained image recognition model M1; calculating the distance between the reference vector and the feature vector extracted from the evaluation image using an evaluation model M2 that has learned a reference vector obtained by reducing the dimensionality of the feature vector extracted from a normal image taken when the inside of the garbage pit 11 is in a normal state, and evaluating the anomaly score for each region of the evaluation image based on the calculated distance; and determining that there is an unsuitable object in a region when the area of ​​a region in which the anomaly score is equal to or greater than the anomaly threshold is equal to or greater than the upper limit area.

[0190] (19) According to the 19th aspect, the program causes the detection device 20 to perform the following steps: acquire an evaluation image taken inside the garbage pit 11; extract feature vectors from the evaluation image using a pre-trained image recognition model M1; calculate the distance between the reference vector and the feature vector extracted from the evaluation image using an evaluation model M2 that has learned a reference vector obtained by reducing the dimensionality of the feature vector extracted from a normal image taken when the inside of the garbage pit 11 is in a normal state, and evaluate the abnormality score for each region of the evaluation image based on the calculated distance; and determine that there is an unsuitable object in the region if the area of ​​the region in which the abnormality score is equal to or greater than the abnormal threshold is equal to or greater than the upper limit area.

[0191] According to the above embodiment, it is possible to detect abnormalities with high accuracy while reducing the computation time and memory consumption of the process for detecting unsuitable materials in the waste pit.

[0192] 1 Detection system 10 Camera 11 Garbage pit 20 Detection device 201 Acquisition unit 202 Extraction unit 203 Evaluation unit 204 Judgment unit 205 Control unit 206 Output unit 207 Learning unit 208 Storage unit 30 Learning device 301 Acquisition unit 302 Extraction unit 307 Learning unit 308 Storage unit M1 Image recognition model M2 Evaluation model

Claims

1. A detection device comprising: an acquisition unit that acquires evaluation images taken of the inside of a garbage pit; an extraction unit that extracts feature vectors from the evaluation images using a pre-trained image recognition model; an evaluation unit that calculates the distance between the reference vector and the feature vector extracted from the evaluation image using an evaluation model that has learned a reference vector obtained by reducing the dimensionality of the feature vector extracted from a normal image taken when the inside of the garbage pit is in a normal state, and evaluates the abnormality score for each region of the evaluation image based on the calculated distance; and a determination unit that determines that an unsuitable object exists in a region when the area of ​​the region in which the abnormality score is equal to or greater than an abnormal threshold is equal to or greater than an upper limit area.

2. The detection device according to claim 1, further comprising a control unit that removes the unsuitable substance when it is determined that such an unsuitable substance is present.

3. The detection device according to claim 1, wherein the evaluation unit identifies clusters of reference vectors that approximate the feature vectors extracted from the evaluation image, and calculates the distance between the reference vectors included in the identified clusters and the feature vectors extracted from the evaluation image.

4. The detection device according to any one of claims 1 to 3, further comprising a learning unit that learns the reference vector obtained by quantizing the feature vector extracted from the normal image to create the evaluation model.

5. The detection device according to claim 4, wherein the learning unit classifies the feature vectors extracted from each of the plurality of normal images into a plurality of clusters, creates a training image dataset by deleting feature vectors in order of increasing distance between feature vectors in each cluster, and learns the reference vector obtained by reducing the dimensionality of the feature vectors of the normal images included in the training image dataset.

6. The detection device according to claim 5, wherein the learning unit reduces the dimensionality of feature vectors extracted from each of the plurality of normal images using a dimensionality reduction algorithm, and classifies the reduced-dimensional feature vectors into a plurality of clusters.

7. The detection device according to claim 5, wherein the learning unit, when the number of clusters obtained by classifying the feature vectors extracted from a new image which is a candidate for a normal image and the feature vectors extracted from each of the multiple normal images included in the learning image dataset into a plurality of second clusters is greater than the number of clusters obtained by classifying only the feature vectors extracted from each of the plurality of normal images included in the learning image dataset into a plurality of first clusters, creates a new learning image dataset by deleting feature vectors in order of increasing distance between feature vectors in each of the plurality of second clusters.

8. The detection device according to claim 4, wherein the learning unit divides the plurality of normal images into a plurality of image groups such that the number of images is less than the upper limit of memory available for training the evaluation model, creates a plurality of evaluation models corresponding to each of the image groups, and the evaluation unit evaluates the abnormality score for each region of the evaluation image based on the minimum distance between the reference vector of each of the plurality of evaluation models and the feature vector extracted from the evaluation image.

9. The detection device according to claim 8, wherein the learning unit divides a plurality of normal images such that the number of image groups is minimized and the number of normal images contained in each of the plurality of image groups approaches equality.

10. The detection device according to claim 8, wherein the learning unit creates and adds a new evaluation model based on a newly added set of normal images when the difference between the total memory consumption of the set of evaluation models and the upper limit of memory available to the evaluation unit when performing evaluation using the evaluation models is greater than the memory consumption of a single evaluation model.

11. A detection system comprising: a detection device according to any one of claims 1 to 3; and a learning device that learns the reference vector obtained by quantizing the feature vector extracted from the normal image to create the evaluation model.

12. The detection system according to claim 11, wherein the learning unit classifies the feature vectors extracted from each of the plurality of normal images into a plurality of clusters, creates a training image dataset by deleting feature vectors in order of increasing distance between feature vectors in each cluster, and learns the reference vector obtained by reducing the dimensionality of the feature vectors of the normal images included in the training image dataset.

13. The detection system according to claim 12, wherein the learning unit reduces the dimensionality of feature vectors extracted from each of the multiple normal images using a dimensionality reduction algorithm, and classifies the reduced-dimensional feature vectors into multiple clusters.

14. The detection system according to claim 12, wherein the learning unit creates a new learning image dataset by deleting feature vectors in order of increasing distance between feature vectors in each of the multiple normal images included in the learning image dataset if the number of clusters obtained by classifying only the feature vectors extracted from each of the multiple normal images included in the learning image dataset into multiple first clusters is greater than the number of clusters obtained by classifying only the feature vectors extracted from a new image which is a candidate for a normal image into multiple second clusters.

15. The detection system according to claim 11, wherein the learning unit divides the plurality of normal images into a plurality of image groups such that the number of images is less than the upper limit of memory available when training the evaluation model, creates a plurality of evaluation models corresponding to each of the image groups, and the evaluation unit evaluates the abnormality score for each region of the evaluation image based on the minimum distance between the reference vector of each of the plurality of evaluation models and the feature vector extracted from the evaluation image.

16. The detection system according to claim 15, wherein the learning unit divides a plurality of normal images such that the number of image groups is minimized and the number of normal images contained in each of the plurality of image groups approaches equality.

17. The detection system according to claim 15, wherein the learning unit creates and adds a new evaluation model based on a newly added set of normal images when the difference between the total capacity of the set of evaluation models and the upper limit of memory available to the evaluation unit when performing evaluation using the evaluation models is greater than the capacity of a single evaluation model.

18. A detection method comprising: acquiring an evaluation image taken inside a garbage pit; extracting feature vectors from the evaluation image using a pre-trained image recognition model; calculating the distance between the reference vector and the feature vector extracted from the evaluation image using an evaluation model trained on a reference vector obtained by reducing the dimensionality of feature vectors extracted from a normal image taken when the inside of the garbage pit is in a normal state, and evaluating the anomaly score for each region of the evaluation image based on the calculated distance; and determining that an unsuitable object exists in a region if the area of ​​the region where the anomaly score is equal to or greater than an anomaly threshold is equal to or greater than an upper limit area.

19. A program that causes a detection device to perform the following steps: acquire an evaluation image taken inside a garbage pit; extract feature vectors from the evaluation image using a pre-trained image recognition model; calculate the distance between the reference vector and the feature vector extracted from the evaluation image using an evaluation model trained on a reference vector obtained by reducing the dimensionality of feature vectors extracted from a normal image taken when the inside of the garbage pit is in a normal state, and evaluate the abnormality score for each region of the evaluation image based on the calculated distance; and determine that an unsuitable object exists in a region if the area of ​​the region where the abnormality score is above an abnormality threshold is above an upper limit area.