Learning device, learning method, and program

The learning device efficiently collects teacher images by registering images with low similarity to abnormal state images as normal state images, addressing the challenge of collecting abnormal behavior images and improving estimation accuracy.

JP7694630B2Active Publication Date: 2025-06-18NEC CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023183772
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-10-26
Publication Date
2025-06-18
Estimated Expiration
2040-06-24

AI Technical Summary

Technical Problem

Existing techniques for generating estimation models for detecting abnormalities face challenges in efficiently collecting teacher images, particularly for abnormal behavior, which requires a large number of images.

Method used

A learning device and method that acquire images, calculate their similarity to pre-defined abnormal state images, register images with low similarity as normal state images, and generate estimation models using these images to discriminate between normal and abnormal states.

Benefits of technology

This approach enables efficient collection of teacher images, reduces user burden in preparing labeled images, and improves estimation accuracy by automatically accumulating normal state images and incrementally increasing abnormal state images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007694630000001
    Figure 0007694630000001
  • Figure 0007694630000002
    Figure 0007694630000002
  • Figure 0007694630000003
    Figure 0007694630000003
Patent Text Reader

Abstract

To enable efficient collection of teacher images to generate an estimation model for detecting abnormalities.SOLUTION: The present invention provides a learning device comprising an acquisition unit for acquiring images, a similarity degree computation unit for computing the degree of similarity between acquired images and first images representing pre-accumulated anomalous states, a registration unit for registering acquired images with degrees of similarity that are equal to or less than a first reference value as second images representing normal states, and a learning unit for generating an estimation model for determining between normal and abnormal states by means of machine learning using the first and second images.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a learning device, an estimation device, a learning method, and a program.

Background Art

[0002] Patent Document 1 discloses a technique for generating an estimation model that classifies an input image into a good image or a bad image by learning based on correct and incorrect teacher images. A good image is an image with a high degree of similarity to the correct teacher image, and a bad image is an image with a low degree of similarity to the correct teacher image. Patent Document 2 discloses a technique for defining abnormal behavior using a teacher image showing abnormal behavior and generating an estimation model for detecting the defined abnormal behavior.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the technique of generating an estimation model for detecting an abnormality, a technique for efficiently collecting teacher images is desired. Patent Document 1 does not disclose the problem and the solution means. In the case of the technique described in Patent Document 2, it is necessary to collect a large number of teacher images showing abnormal behavior. However, it is not easy to collect teacher images showing "abnormality". An object of the present invention is to provide a technique for efficiently collecting teacher images for generating an estimation model for detecting an abnormality.

Means for Solving the Problems

[0005] According to the present invention, an acquisition means for acquiring an image, Similarity calculation means for calculating the similarity between the acquired image and a first image indicating an abnormal state stored in advance; Registration means for registering the acquired image whose similarity is equal to or less than a first reference value as a second image indicating a normal state; Learning means for generating an estimation model for discriminating normal / abnormal by machine learning using the first image and the second image; There is provided a learning device having the above.

[0006] Further, according to the present invention, a computer acquires an image, calculates the similarity between the acquired image and a first image indicating an abnormal state stored in advance, registers the acquired image whose similarity is equal to or less than a first reference value as a second image indicating a normal state, and provides a learning method for generating an estimation model for discriminating normal / abnormal by machine learning using the first image and the second image.

[0007] Further, according to the present invention, a computer is caused to function as acquisition means for acquiring an image, similarity calculation means for calculating the similarity between the acquired image and a first image indicating an abnormal state stored in advance, registration means for registering the acquired image whose similarity is equal to or less than a first reference value as a second image indicating a normal state, learning means for generating an estimation model for discriminating normal / abnormal by machine learning using the first image and the second image, and there is provided a program for causing the above to function.

[0008] Further, according to the present invention, there is provided an estimation device for discriminating normal / abnormal using the estimation model generated by the learning device.

Advantages of the Invention

[0009] According to the present invention, it is possible to efficiently collect teacher images for generating an estimation model for detecting abnormalities.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Modes for Carrying Out the Invention

[0011] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In all the drawings, the same components are denoted by the same reference numerals, and the description will be omitted as appropriate.

[0012] <First Embodiment> The learning device of this embodiment (hereinafter, may be simply referred to as "learning device") generates an estimation model for discriminating whether the state indicated by the input image is normal or abnormal.

[0013] The object to be discriminated between normal and abnormal is, for example, a location (such as a park, a station, a facility, etc.). The normal state observed for most of the time is discriminated as normal, and a state different from the normal state is discriminated as abnormal. For example, a state where there is a person performing abnormal behavior, or a state where an object that always exists at that location has malfunctioned or moved, etc. is discriminated as abnormal. Abnormal behavior is behavior different from the behavior performed by most people observed in the image. Note that the object to be discriminated may also be other facilities such as factories, stores, facilities, offices, etc., or may be other things. In any case, the normal state observed for most of the time is discriminated as normal, and a state different from the normal state is discriminated as abnormal.

[0014] The learning device generates the above-mentioned estimation model by repeatedly executing the cycle shown in FIG. 1. As shown in FIG. 1, the learning device repeatedly executes a first image registration process S1, an image selection process S2, a learning process S3, an estimation process S4, a user confirmation process S5, and a second image registration process S6 in this order. Note that the order of the processes may be changed within the range where the same operational effects are achieved.

[0015] FIG. 2 shows an example of a functional block diagram of the learning device 10. As shown in the figure, the learning device 10 includes an acquisition unit 11, a similarity calculation unit 12, a registration unit 13, a learning unit 14, an estimation unit 15 during learning, a user confirmation unit 16, an image storage unit 17, and an estimation model storage unit 18. Each process shown in FIG. 1 is executed by these functional units.

[0016] FIG. 3 is a diagram showing the cycle of FIG. 1 in more detail. Using this figure, each process shown in FIG. 1 and the processes of each functional unit shown in FIG. 2 will be described.

[0017] "The first image registration process S1" The first image registration process S1 is a process of classifying and registering the image generated by the camera based on the similarity between the image generated by the camera and the image indicating an abnormal state registered in advance.

[0018] The first to third image group databases DB17-1 to 17-3, camera D14, similarity calculation S10, and registration S11 in FIG. 3 are related to the process. And the acquisition unit 11, similarity calculation unit 12, registration unit 13, and image storage unit 17 in FIG. 2 are related to the process. The first to third image group databases DB17-1 to 17-3 are realized by the image storage unit 17 in FIG. 2.

[0019] First, as a preliminary preparation for the process, labeled images with abnormal state labels are stored in the first image group database (database) 17-1. The user prepares several images indicating an abnormal state in advance, assigns an abnormal state label, and stores them in the first image group database 17-1. The images in the first image group database 17-1 accumulated in this way are highly reliable labeled images that have been confirmed by the user to indicate an abnormal state. Note that the number of images initially stored in the first image group database 17-1 may be about several tens to several hundreds, and a large number of images are not necessary. With this number, the user burden required for collecting labeled images is not large. When defining an abnormal state in advance and generating an estimation model for detecting the abnormal state, generally, it is necessary to prepare several thousand to several tens of thousands or more teacher images indicating the abnormal state. The first image group database 17-1 corresponds to the image storage unit 17 in FIG. 2. Hereinafter, the images indicating the abnormal state stored in the first image group database 17-1 are referred to as "first images".

[0020] The acquisition unit 11 acquires the images generated by the camera D14. The camera D14 may be a camera (such as a surveillance camera) that captures an object to be discriminated as normal / abnormal, or a camera that captures an object of the same type as the object to be discriminated. The camera D14 may capture a moving image, or may continuously capture still images at a frame interval longer than that of a moving image. In the figure, one camera D14 is shown, but a plurality of cameras D14 may be used.

[0021] The acquisition unit 11 may acquire the image generated by the camera D14 in real-time processing. In this case, the learning device 10 and the camera D14 are configured to be communicable with each other. Alternatively, the acquisition unit 11 may acquire the image generated by the camera D14 in batch processing. In this case, the image generated by the camera D14 is stored in the storage device of the camera D14 or any other arbitrary storage device, and the acquisition unit 11 acquires the stored image at an arbitrary timing.

[0022] Note that in this specification, "acquisition" means, based on user input or based on a program instruction, "the act of the own device going to retrieve data stored in another device or storage medium (active acquisition)", for example, making a request or inquiry to another device and receiving the response, accessing another device or storage medium and reading out the data, etc., and, based on user input or based on a program instruction, "the act of the own device inputting data output from another device (passive acquisition)", for example, receiving data distributed (or transmitted, push-notified, etc.), selecting and acquiring from the received data or information, and "the act of generating new data by editing (textualizing, rearranging data, extracting some data, changing the file format, etc.) the data and acquiring the new data", including at least any one of these.

[0023] The similarity calculation unit 12 calculates the similarity between the image acquired by the acquisition unit 11 (hereinafter referred to as "acquired image") and the first image indicating an abnormal state stored in the first image group DB17-1 in advance (S10 in FIG. 3). The similarity calculation unit 12 may calculate the similarity between each of the plurality of first images stored in the first image group DB17-1 and each acquired image. Alternatively, the similarity calculation unit 12 may calculate the similarity between one image (e.g., average image) generated based on the plurality of first images stored in the first image group DB17-1 and each acquired image.

[0024] In addition, various methods have been proposed for calculating the similarity between images. In this embodiment, any method can be adopted. For example, the similarity calculation unit 12 may detect an object from within an image and calculate the similarity of the detection result (the similarity of the number of detected objects, the similarity of the appearance of the detected objects, etc.). Further, the similarity calculation unit 12 may input each image into an estimation model that performs image analysis generated by deep learning, and calculate the similarity of the obtained image analysis results (the recognition result of the object shown in the image, the recognition result of the scene shown in the image, etc.). Further, the similarity calculation unit 12 may calculate the similarity of the colors and luminances appearing in the whole or local part of the image.

[0025] The registration unit 13 registers the acquired image whose similarity is equal to or less than the first reference value as a second image (image with a normal state label) indicating a normal state in the second image group DB (database) 17-2 (S11). When the similarity calculation unit 12 calculates the similarity between each of the plurality of first images stored in the first image group DB17-1 and each acquired image, the registration unit 13 registers the acquired image whose similarity with all of the plurality of first images is equal to or less than the first reference value as the second image in the second image group DB17-2.

[0026] Further, the registration unit 13 registers the acquired image whose similarity is equal to or greater than the second reference value as a third image (image with an abnormal state label) indicating an abnormal state in the third image group DB (database) 17-3 (S11). When the similarity calculation unit 12 calculates the similarity between each of the plurality of first images stored in the first image group DB17-1 and each acquired image, the registration unit 13 registers the acquired image whose similarity with at least one of the plurality of first images is equal to or greater than the second reference value as the third image in the third image group DB17-3.

[0027] In the third image group DB17-3, images determined to be similar to the first image by a computer at a predetermined level or higher in this way are registered as images indicating an abnormal state. In this regard, it is different from the first image group DB17-1 in which highly reliable first images confirmed by the user to indicate an abnormal state are stored.

[0028] The first reference value and the second reference value may be the same value or different values. However, by setting the first reference value and the second reference value to different values, making the first reference value a sufficiently small value, and making the second reference value a sufficiently large value, it is possible to suppress the inconvenience of registering an acquired image that exists in a gray zone (similarity is greater than the first reference value and less than the second reference value) and has neither a high nor a low similarity to the first image as the second image or the third image.

[0029] "Image Selection Process S2, Learning Process S3" The image selection process S2 is a process of selecting an image to be a teacher image from among the images stored in the first to third image groups DB17-1 to 17-3. The learning process S3 is a process of performing learning for each of the plurality of estimation models registered in the estimation model DB (database) 18-1 using the selected image as a teacher image.

[0030] The first to third image groups DB17-1 to 17-3, the estimation model DB18-1, the selection S12, and the learning S13 in FIG. 3 are related to the process. And the learning unit 14, the image storage unit 17, and the estimation model storage unit 18 in FIG. 2 are related to the process. The estimation model DB18-1 is realized by the estimation model storage unit 18 in FIG. 2.

[0031] First, information on a plurality of estimation models is stored in the estimation model DB18-1. All of the plurality of estimation models are models for discriminating whether the state indicated by the input image is normal or abnormal. The plurality of estimation models have mutually different learning and estimation algorithms. For example, the plurality of estimation models are generated by deep learning. In the present embodiment, for example, information on a plurality of estimation models learned and generated by a neural network, a Bayesian network, regression analysis, a support vector machine (SVM), a decision tree, a genetic algorithm, a nearest neighbor method classification, etc. is stored in the estimation model DB18-1.

[0032] The learning unit 14 selects at least a part from the images registered in the first to third image group databases 17-1 to 17-3 (S12 in FIG. 3), and generates an estimation model by machine learning using the selected images (S13 in FIG. 3).

[0033] There are various selection methods. For example, the learning unit 14 may randomly select a predetermined number of images from the entire first to third image group databases 17-1 to 17-3. Alternatively, the learning unit 14 may randomly select a first predetermined number of images from the first image group database 17-1, randomly select a second predetermined number of images from the second image group database 17-2, and randomly select a third predetermined number of images from the third image group database 17-3. The first to third predetermined numbers may be the same or different. That is, the ratio of the number of images selected from each of the first to third image group databases 17-1 to 17-3 (the ratio to the total number of selected images) may be the same or different.

[0034] Also, the learning unit 14 may select images for each estimation model. In this case, the above first to third predetermined numbers and the above ratio may be different for each estimation model.

[0035] After selecting the images, the learning unit 14 uses the selected first to third images as teacher images and performs learning for each of the plurality of estimation models registered in the estimation model database (database) 18-1. That is, the learning unit 14 generates an estimation model for discriminating normal / abnormal by machine learning (including the concept of deep learning) using the first to third images.

[0036] "Estimation process S4" The estimation process S4 is a process of inputting the acquired image into each of the plurality of estimation models registered in the estimation model database (database) 18-1 and discriminating the state indicated by the acquired image.

[0037] The estimation model DB18-1, camera D14, and estimation S14 in FIG. 3 are related to the said process. And the acquisition unit 11, estimation unit 15 during learning, and estimation model storage unit 18 in FIG. 2 are related to the said process.

[0038] The estimation unit 15 during learning inputs the acquired image into each of the plurality of estimation models stored in the estimation model storage unit 18, and discriminates the state (normal / abnormal) indicated by the acquired image. Note that the acquired image input into the estimation model in the said process is an acquired image not used at that time for the generation (learning) of the estimation model. For example, the estimation unit 15 during learning can perform the said discrimination using the acquired image before being stored in the image storage unit 17.

[0039] Note that the discrimination results for each of the plurality of estimation models may be stored in the storage device within the learning device 10.

[0040] "User confirmation process S5" The user confirmation process S5 is a process that outputs the discrimination result of the estimation process S4 to the user and receives the correct / incorrect input of the discrimination result from the user.

[0041] The display device D15, extraction S15, output S16, and correct / incorrect input S17 in FIG. 3 are related to the said process. And the user confirmation unit 16 in FIG. 2 is related to the said process.

[0042] The user confirmation unit 16 outputs the discrimination result by the estimation unit 15 during learning to the user (S16 in FIG. 3) and receives the correct / incorrect input of the discrimination result from the user (S17 in FIG. 3). For example, the user confirmation unit 16 outputs the acquired image and the discrimination result (normal state or abnormal state) and receives the correct / incorrect input of the discrimination result for the acquired image.

[0043] Executing the said process for all acquired images will increase the burden on the user. Therefore, the user confirmation unit 16 may extract some acquired images that satisfy a predetermined condition (S15 in FIG. 3) and perform only the output of the discrimination result (S16 in FIG. 3) and the reception of the correct / incorrect input (S17 in FIG. 3) for the extracted some acquired images.

[0044] Some of the acquired images for which the discrimination result is output and the correct / incorrect input is received may be, for example, any of the following.

[0045] · An acquired image determined to indicate an abnormal state in at least one estimation model. · An acquired image determined to indicate an abnormal state with a confidence level equal to or higher than a predetermined level in at least one estimation model. · An acquired image determined to indicate an abnormal state in a predetermined number or more of estimation models. · An acquired image determined to indicate an abnormal state with a confidence level equal to or higher than a predetermined level in a predetermined number or more of estimation models. · An acquired image determined to indicate an abnormal state in all estimation models. · An acquired image determined to indicate an abnormal state with a confidence level equal to or higher than a predetermined level in all estimation models.

[0046] Some of the acquired images for which the discrimination result is output and the correct / incorrect input is received may, in addition to any of the acquired images described above, include an acquired image randomly picked from among the acquired images that do not satisfy the above conditions (acquired images presumed to indicate a normal state).

[0047] The user confirmation unit 16 may output the discrimination result via an arbitrary output device such as a display or a projection device, and may receive the correct / incorrect input via an arbitrary input device such as a keyboard, a mouse, a touch panel, a physical button, or a microphone. Alternatively, the user confirmation unit 16 may transmit the discrimination result to a predetermined mobile terminal and acquire the content of the correct / incorrect input made to the mobile terminal from the mobile terminal. Alternatively, the user confirmation unit 16 may save the discrimination result in a browsable state on an arbitrary server from an arbitrary device. Then, the user confirmation unit 16 may acquire the content of the correct / incorrect input input from an arbitrary device and saved on the above server. Note that the examples illustrated here are merely examples and are not limited thereto.

[0048] "Second Image Registration Process S6" The second image registration process S6 is a process of registering, as the first image, the acquired image in which an abnormal state is input in the user confirmation process S5 into the first image group DB17-1.

[0049] The first image group DB17-1 and the registration S18 in FIG. 3 are related to this process. And the registration unit 13 and the image storage unit 17 in FIG. 2 are related to this process.

[0050] The registration unit 13 registers, as the first image, the acquired image in which an abnormal state is input in the correct / incorrect input received by the user confirmation unit 16 into the first image group DB17-1.

[0051] The acquired image in which an abnormal state is input corresponds to, for example, an acquired image whose discrimination result is "abnormal state" and whose correct / incorrect input is "correct", or an acquired image whose discrimination result is "normal state" and whose correct / incorrect input is "incorrect".

[0052] Here, a modified example of the learning device 10 of the present embodiment will be described. The learning device 10 may not have the third image group DB17-3. And the registration unit 13 may not execute a process of registering, as the second image, an acquired image whose similarity to the first image is equal to or less than the first reference value into the second image group DB17-2, and may not execute a process of registering, as the third image, an acquired image whose similarity to the first image is equal to or greater than the second reference value into the third image group DB17-3. In this case, images indicating a normal state will be accumulated by the process of the registration unit 13.

[0053] Next, an example of the hardware configuration of the learning device 10 will be described. Each functional unit of the learning device 10 is realized by an arbitrary combination of hardware and software centered around a CPU (Central Processing Unit) of an arbitrary computer, a memory, a program loaded into the memory, a storage unit such as a hard disk storing the program (in addition to the program stored in advance at the stage of shipping the device, a program downloaded from a storage medium such as a CD (Compact Disc) or a server on the Internet can also be stored), and a network connection interface. And it is understood by those skilled in the art that there are various modifications to the realization method and device.

[0054] FIG. 4 is a block diagram illustrating the hardware configuration of the learning device 10. As shown in FIG. 4, the learning device 10 includes a processor 1A, a memory 2A, an input / output interface 3A, a peripheral circuit 4A, and a bus 5A. The peripheral circuit 4A includes various modules. The learning device 10 may not have the peripheral circuit 4A. Note that the learning device 10 may be composed of a plurality of physically and / or logically divided devices, or may be composed of one physically and / or logically integrated device. When the learning device 10 is composed of a plurality of physically and / or logically divided devices, each of the plurality of devices can have the above hardware configuration.

[0055] Bus 5A is a data transmission path for the processor 1A, memory 2A, peripheral circuit 4A, and input / output interface 3A to transmit and receive data from each other. The processor 1A is an arithmetic processing device such as a CPU or a GPU (Graphics Processing Unit). The memory 2A is a memory such as a RAM (Random Access Memory) or a ROM (Read Only Memory). The input / output interface 3A includes an interface for acquiring information from an input device, an external device, an external server, an external sensor, a camera, etc., and an interface for outputting information to an output device, an external device, an external server, etc. The input device is, for example, a keyboard, a mouse, a microphone, a physical button, a touch panel, etc. The output device is, for example, a display, a speaker, a printer, a mailer, etc. The processor 1A can issue commands to each module and perform operations based on their operation results.

[0056] Next, the operation and effect of the learning device 10 will be described.

[0057] The learning device 10 of this embodiment generates an estimation model for discriminating normal / abnormal by machine learning using an image indicating a normal state and an image indicating an abnormal state as teacher images. In the estimation model, the normal state observed most of the time is discriminated as normal, and a state different from the normal state is discriminated as abnormal.

[0058] In such a case, even if an abnormal state that has not been defined in advance occurs, as long as the state is different from the normal state, it can be discriminated as an abnormal state. Therefore, it is possible to detect abnormal states without omission.

[0059] Also, when an abnormal state is defined in advance and an estimation model for detecting the abnormal state is generated, it is necessary to prepare a large number of teacher images indicating each abnormal state. However, it is not easy to prepare teacher images indicating abnormal states. In the case of this embodiment, compared with the case of generating an estimation model for detecting an abnormal state defined in advance, the number of "images indicating abnormal states" to be prepared is reduced. As a result, the burden on the user is reduced.

[0060] In addition, in the case of this embodiment, a large number of "images indicating a normal state" are required. However, usually, since many objects are in the "normal state", it is possible to easily collect "images indicating a normal state" from the images of such objects being photographed.

[0061] Also, in the case of this embodiment, based on the result of similarity determination between a small number of "images indicating an abnormal state" prepared in advance (meaning a number smaller than the number of images indicating an abnormal state required to generate an estimation model for detecting a predefined abnormal state) and the images generated by a surveillance camera or the like, "images indicating a normal state" can be automatically accumulated. For this reason, the burden on the user is reduced.

[0062] Also, in the case of this embodiment, the number of "images indicating an abnormal state" can be increased by the second image registration process S6. Since the number of "images indicating an abnormal state" can be increased in this way, the estimation accuracy of the obtained estimation model is improved.

[0063] Also, in the case of this embodiment, the number of "images indicating an abnormal state" can also be increased by the first image registration process S1. In this case, by setting the above-described second reference value to a sufficiently high value, more reliable "images indicating an abnormal state" can be increased. And with the increase in the "images indicating an abnormal state", an improvement in the estimation accuracy of the obtained estimation model is expected.

[0064] Also, in the case of this embodiment, the images indicating an abnormal state can be divided into and managed as "high-reliability first images confirmed by the user to indicate an abnormal state" and "third images determined by the computer to be similar to the first images at a predetermined level or higher". And only the first images can be the reference targets for the similarity calculation S10 in FIG. 3. In this way, by using only the high-reliability first images as the reference targets, the reliability of the process of classifying into a normal state / abnormal state based on the similarity between images (similarity calculation S10 and registration S11 in FIG. 3) is increased.

[0065] Also, in the case of this embodiment, a plurality of estimation models can be learned in parallel. Therefore, in an actual estimation scenario (estimation by the estimation device described in the following embodiments), it is possible to select and use an estimation model that can obtain a more preferable result from among them.

[0066] <Second Embodiment> FIG. 5 shows an example of a functional block diagram of the learning device 10 of this embodiment. Also, FIG. 6 shows a diagram that more specifically shows the cycle of FIG. 1. Comparing FIGS. 2 and 3 described in the first embodiment with FIGS. 5 and 6 showing the configuration of this embodiment, the learning device 10 of this embodiment is different in that it does not have the third image group DB 17-3, and the image storage unit 17 does not store the third image group.

[0067] In the first embodiment, images indicating an abnormal state were separately managed as "highly reliable first images confirmed by the user to indicate an abnormal state" and "third images determined by the computer to be similar to the first images by a predetermined level or more". However, the learning device 10 of this embodiment does not perform such management. That is, "images confirmed by the user to indicate an abnormal state and having a high reliability" and "images determined by the computer to be similar to the highly reliable images by a predetermined level or more" are collectively managed as "first images indicating an abnormal state". The "first images" of this embodiment are images indicating an abnormal state, and are a concept including the first images and the third images described in the first embodiment.

[0068] The registration unit 13 registers an acquired image whose similarity to the first image registered in the first image group DB 17-1 is equal to or higher than a second reference value as a first image in the first image group DB 17-1.

[0069] Other configurations of the learning device 10 of this embodiment are the same as those of the first embodiment.

[0070] According to the learning device 10 of the present embodiment described above, the same operational effects as those of the learning device 10 of the first embodiment are achieved. In addition, images indicating an abnormal state can be efficiently collected. Note that the reliability (reliability of indicating an abnormal state) of "high-reliability images confirmed by the user to indicate an abnormal state" and "images determined by the computer to be similar to the high-reliability images above a predetermined level" may be different. Mixing and managing images with different reliabilities may have an adverse effect on learning accuracy, estimation accuracy, etc. However, if the second reference value described above is set to a sufficiently high value, such inconveniences can be reduced.

[0071] <Third Embodiment> The estimation device of the present embodiment discriminates the state (normal / abnormal) indicated by an image using the estimation model generated by the learning device 10 of the first or second embodiment.

[0072] Since the estimation device of the present embodiment can collect sufficient and highly accurate teacher images by a characteristic method as described above and use the estimation model generated by learning based on the teacher images, high estimation accuracy can be obtained.

[0073] As described above, embodiments of the present invention have been described with reference to the drawings, but these are examples of the present invention, and various configurations other than the above can also be adopted.

[0074] In addition, in the plurality of flowcharts used in the above description, a plurality of steps (processes) are described in order, but the execution order of the steps executed in each embodiment is not limited to the described order. In each embodiment, the order of the illustrated steps can be changed within a range that does not interfere with the content. In addition, the above-described embodiments can be combined within a range where the contents do not conflict.

[0075] Some or all of the above embodiments can also be described as follows in the appended claims, but are not limited thereto. 1. Acquisition means for acquiring an image, Similarity calculation means for calculating the similarity between the acquired image and a first image indicating an abnormal state stored in advance; Registration means for registering the acquired image whose similarity is equal to or less than a first reference value as a second image indicating a normal state; Learning means for generating an estimation model for discriminating normal / abnormal by machine learning using the first image and the second image; A learning device having the above. 2. The registration means registers the acquired image whose similarity is equal to or more than a second reference value as a third image indicating an abnormal state, The learning means generates the estimation model by machine learning using the first image, the second image, and the third image. The learning device according to 1. 3. The registration means registers the acquired image whose similarity is equal to or more than a second reference value as the first image. The learning device according to 1. 4. The learning means selects a part from the registered images, and generates the estimation model by machine learning using the selected images. The learning device according to any one of 1 to 3. 5. Learning-time estimation means for discriminating the state indicated by the acquired image using the estimation model; User confirmation means for outputting the acquired image determined to indicate an abnormal state by the learning-time estimation means and receiving a correct / incorrect input from the user; The learning device further has the above, The registration means registers the acquired image input as indicating an abnormal state in the correct / incorrect input as the first image. The learning device according to any one of 1 to 4. 6. The learning means executes learning of each of the plurality of estimation models that learn with different algorithms, The learning-time estimation means discriminates the state indicated by the acquired image using each of the plurality of estimation models, and accumulates the discrimination results of each of the plurality of estimation models. The learning device according to any one of 1 to 5. 7. The acquisition means acquires an image generated by a surveillance camera. The learning device according to any one of 1 to 6. 8. A computer, Obtain an image, calculate the similarity between the obtained image and a first image indicating an abnormal state stored in advance, register the obtained image whose similarity is equal to or less than a first reference value as a second image indicating a normal state, A learning method for generating an estimation model for discriminating normal / abnormal by machine learning using the first image and the second image. 9. Cause a computer to acquisition means for acquiring an image, similarity calculation means for calculating the similarity between the obtained image and a first image indicating an abnormal state stored in advance, registration means for registering the obtained image whose similarity is equal to or less than a first reference value as a second image indicating a normal state, learning means for generating an estimation model for discriminating normal / abnormal by machine learning using the first image and the second image, A program that functions as. 10. An estimation device that discriminates normal / abnormal using the estimation model generated by the learning device according to any one of 1 to 7.

Explanation of symbols

[0076] 10 Learning device 11 Acquisition unit 12 Similarity calculation unit 13 Registration unit 14 Learning unit 15 Estimation unit during learning 16 User confirmation unit 17 Image storage unit 17-1 First image group DB 17-2 Second image group DB 17-3 Third image group DB 18 Estimation model storage unit 18-1 Estimation model DB D14 Camera D15 Display device

Claims

1. Learning means for generating a plurality of first estimation models by machine learning using a first image indicating an abnormal state and a second image indicating a normal state, Acquisition means for acquiring a classification determined for an imaging image in which an object is reflected among a plurality of classifications including a normal state and an abnormal state, the determination being made using at least one first estimation model among the plurality of first estimation models; Receiving means for receiving the result of a user's correct / incorrect input for the classification determined for each imaging image; having The learning means further generates a second estimation model for determining a normal state and an abnormal state for an image by machine learning using the imaging image classified as a normal state or an abnormal state by the correct / incorrect input together with the first image and the second image. Learning device.

2. The receiving means receives the result of the user's correct / incorrect input for the imaging image determined by the first estimation model to indicate an abnormal state with a reliability of a predetermined level or higher. The learning device according to claim 1.

3. The acquisition means acquires a classification determined for an imaging image in which an object is reflected among the plurality of classifications, the determination being made using the second estimation model. The learning device according to claim 1 or 2.

4. The abnormal state indicates a failure or defect of the object. The learning device according to any one of claims 1 to 3.

5. The first image is a pre-accumulated image, and the second image is an image determined to indicate a normal state with a similarity to the first image being equal to or less than a reference value. The learning device according to any one of claims 1 to 4.

6. The plurality of first estimation models include estimation models generated based on mutually different algorithms. The learning device according to any one of claims 1 to 5.

7. The learning means generates a plurality of the second estimation models by machine learning using the captured image classified as a normal state or an abnormal state by the correct / incorrect input together with the first image and the second image. The learning device according to any one of claims 1 to 6.

8. The plurality of the second estimation models include estimation models generated based on mutually different algorithms. The learning device according to claim 7.

9. The normal state and the abnormal state are classifications for the same type of object. The learning device according to any one of claims 1 to 8.

10. The first image indicates an abnormal state, and the second image indicates a normal state. The learning device according to any one of claims 1 to 9.

11. A computer An acquisition step of acquiring a classification determined for a captured image in which an object appears among a plurality of classifications including a normal state and an abnormal state, determined using at least one of the first estimation models obtained by machine learning using a first image indicating an abnormal state and a second image indicating a normal state; A reception step of receiving a result of a user's correct / incorrect input for the determined classification for each captured image; A learning step of generating a second estimation model for determining a normal state and an abnormal state for an image by machine learning using the captured image classified as a normal state or an abnormal state by the correct / incorrect input together with the first image and the second image; A learning method for executing.

12. In the reception step, the computer receives a result of the user's correct / incorrect input for the captured image determined by the first estimation model to indicate an abnormal state with a reliability equal to or higher than a predetermined level. The learning method according to claim 11.

13. In the acquisition step, the computer acquires a classification determined using the second estimation model for an imaging image in which an object is reflected and determined for one of the plurality of classifications. The learning method according to claim 11 or 12.

14. The abnormal state indicates a failure or defect of the object. The learning method according to any one of claims 11 to 13.

15. The first image is a pre-accumulated image, and the second image is an image determined to have a similarity with the first image below a reference value and to indicate a normal state. The learning method according to any one of claims 11 to 14.

16. The plurality of second estimation models include estimation models generated based on mutually different algorithms. The learning method according to any one of claims 11 to 15.

17. In the learning step, a plurality of the second estimation models are generated by machine learning using the imaging image classified as a normal state or an abnormal state by the correct / incorrect input together with the first image and the second image. The learning method according to any one of claims 11 to 16.

18. The plurality of second estimation models include estimation models generated based on mutually different algorithms. The learning method according to claim 17.

19. The normal state and the abnormal state are classifications for the same type of object. The learning method according to any one of claims 11 to 18.

20. The first image indicates an abnormal state, and the second image indicates a normal state. The learning method according to any one of claims 11 to 19.

21. A computer Acquisition means for obtaining a classification determined for a captured image in which an object is reflected, from among a plurality of classifications including a normal state and an abnormal state, using at least one of the plurality of first estimation models obtained by machine learning using a first image indicating an abnormal state and a second image indicating a normal state. Receiving means for receiving the result of a user's correct / incorrect input for the determined classification for each captured image. Learning means for generating a second estimation model for determining a normal state and an abnormal state for an image, by machine learning using the captured image classified as a normal state or an abnormal state by the correct / incorrect input, together with the first image and the second image. A program that functions as.

22. The receiving means receives the result of the user's correct / incorrect input for the captured image determined by the first estimation model to indicate an abnormal state with a reliability of a predetermined level or higher. The program according to claim 21.

23. The acquisition means acquires a classification determined for a captured image in which an object is reflected, from among the plurality of classifications, using the second estimation model. The program according to claim 21 or 22.

24. The program according to any one of claims 21 to 23, wherein the abnormal state indicates a failure or defect of the object.

25. The program according to any one of claims 21 to 24, wherein the first image is a pre-accumulated image, and the second image is an image determined to indicate a normal state with a similarity to the first image being equal to or less than a reference value.

26. The program according to any one of claims 21 to 25, wherein the plurality of first estimation models include estimation models generated based on mutually different algorithms.

27. The learning means generates a plurality of the second estimation models by machine learning using the captured image classified as a normal state or an abnormal state by the correct / incorrect input together with the first image and the second image, the program according to any one of claims 21 to 26.

28. The plurality of the second estimation models include estimation models generated based on mutually different algorithms, the program according to claim 27.

29. The normal state and the abnormal state are classifications for the same type of object, the program according to any one of claims 21 to 28.

30. The first image indicates an abnormal state, and the second image indicates a normal state, the program according to any one of claims 21 to 29.

Citation Information

Patent Citations

  • Method and device for defect inspection

    JP2010054346A

  • Teacher data creation support method, image classification method, teacher data creation support device and image classification device

    JP2016051429A

  • Abnormality detection system, information processing device, and abnormality detection method

    JP2019053384A

  • Image classifier and program

    JP2020024534A

  • Image identification device, image identification method, and image identification program

    JP2020035097A