Training data collection device and training data collection method

The learning data collection device optimizes image data collection by using a storage type setting unit, subject detection unit, and image collection unit to improve model accuracy and reduce costs, addressing user proficiency gaps and inefficient data management.

WO2025159212A1PCT designated stage expired Publication Date: 2025-07-31PANASONIC I PRO SENSING SOLUTIONS CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/005382
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-23
Filing Date
2025-02-18
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Users lacking proficiency in machine learning face challenges in determining the appropriate image data to collect for improving model accuracy, leading to high data management costs and inefficient data collection for training models.

Method used

A learning data collection device and method that includes a storage type setting unit for defining collection conditions, a subject detection unit for identifying specific subjects, and an image collection unit for selecting images that meet these conditions, thereby optimizing data collection for machine learning models.

Benefits of technology

The solution allows for efficient and cost-effective collection of necessary image data, reducing data management costs while improving model accuracy by focusing on false negatives and positives, and addressing sudden motion changes that affect detection accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025005382_31072025_PF_FP_ABST
    Figure JP2025005382_31072025_PF_FP_ABST
Patent Text Reader

Abstract

This training data collection device comprises a storage type configuration unit that receives a collection condition configuration operation for images to be used as training data, a subject detection unit to which images for processing by a trained model are input and which outputs subject detection results in which a specific subject has been detected, and an image collection unit that collects only images for processing for which the collection condition is satisfied on the basis of the subject detection results.
Need to check novelty before this filing date? Find Prior Art

Description

Learning data collection device and learning data collection method

[0001] The present invention relates to a training data collection device and a training data collection method.

[0002] Patent Document 1 discloses a learning data generation device that includes "an image acquisition unit that acquires an image, a moving object detection unit that detects a moving object in the image, a detection frame enclosing unit that adds a detection frame to the body area of ​​the detected moving object, an extraction unit that extracts a body area image, and a generation unit that generates learning data for a discrimination model from the body area image." (Abstract excerpt)

[0003] Japanese Patent Application Laid-Open No. 2022-013910

[0004] According to Patent Literature 1, a detection frame can be assigned to an individual by inputting image data into a discrimination model that has been trained by machine learning to extract the entire individual from an image of the individual. However, if the accuracy of the machine learning model is low, there is a risk that the individual will not be surrounded by the correct detection frame, or that the detection frame will surround a location other than the individual. Therefore, there is a need to improve the accuracy of the machine learning model. Therefore, while it is necessary to collect image data, users who are not familiar with machine learning find it difficult to accurately determine what type of image data should be collected and to what extent. Therefore, when accumulating image data to be used as training data, there is a problem in that data management costs increase if image data is stored indiscriminately.

[0005] The present invention has been made to solve the above-mentioned problems, and aims to provide a training data collection device and a training data collection method that more appropriately collect image data necessary for training a machine learning model while suppressing increases in data management costs.

[0006] To achieve the above object, the present invention employs the configuration described in the claims. As an example, the present invention provides a training data collection device that includes a save type setting unit that accepts a setting operation for collection conditions for images used as training data, an object detection unit that inputs processing target images to a trained model and outputs an object detection result that detects a specific object, and an image collection unit that collects only the processing target images that satisfy the collection conditions based on the object detection result.

[0007] The present invention also provides a training data collection method, which includes the steps of: inputting a processing target image into a trained model and outputting a subject detection result in which a specific subject is detected; and extracting only the processing target image that satisfies predetermined conditions for collecting images to be used as training data based on the subject detection result, and storing the extracted images in a storage medium.

[0008] According to the present invention, it is possible to provide a training data collection device and a training data collection method that more appropriately collect image data necessary for training a machine learning model while suppressing increases in data management costs. Objectives, configurations, and effects other than those described above will be made clear in the following embodiments.

[0009] FIG. 1 is a diagram illustrating an example of a system configuration of a training data collection device according to the present embodiment; FIG. 2 is a diagram illustrating an internal configuration of a training data collection device according to the present embodiment; FIG. 3 is a flowchart illustrating a processing flow of the training data collection device; FIG. 4 is a diagram illustrating an example of a storage type setting screen; FIG. 5 is a diagram illustrating an example of storage of false negative images; FIG. 6 is a diagram illustrating an example of a storage type setting screen when automatically collecting false positive images; FIG. 7 is a diagram illustrating an example of storage of false positive images; FIG. 8 is a diagram illustrating an example of a storage type setting screen when extracting and automatically collecting images with movement from among false negative images;

[0010] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. In all the drawings for explaining the embodiment, the same components are generally designated by the same reference numerals, and repeated description thereof will be omitted.

[0011] FIG. 1 is a diagram showing an example of the system configuration of a learning data collection device 1 according to this embodiment.

[0012] The learning data collection device 1 shown in FIG. 1 is configured by connecting a detection processing terminal 10 and a setting device 20 for communication.

[0013] The detection processing terminal 10 is a device equipped with a trained subject detection model that has undergone machine learning using images of a subject to be monitored as training data. The detection processing terminal 10 may be, for example, a so-called AI camera (AI: Artificial Intelligence) equipped with the trained subject detection model, or may be a server. Furthermore, the AI ​​camera used in the detection processing terminal 10 may be a fixed camera with a fixed angle of view, or a PTZ camera (PTZ is an abbreviation of Panoramac Tilt Zoom) that can change the camera's shooting direction in the pan and tilt directions and also has a zoom function.

[0014] The setting device 20 is a device for performing setting operations on the detection processing terminal 10. These setting operations include, for example, an operation for selecting the type of subject to be detected, an operation for selecting a trained subject detection model to be used in subject detection processing, and setting operations such as collection timing. The setting device 20 may be, for example, a personal computer, a tablet terminal, or a smartphone. When the detection processing terminal 10 is a camera, the setting device 20 is an operation terminal for the camera. When the detection processing terminal 10 is a server, the setting device 20 is an operation terminal for the server.

[0015] FIG. 2 is a diagram showing the internal configuration of the training data collection device according to this embodiment.

[0016] The detection processing terminal 10 includes a processor 11 , a memory 12 , an image storage unit 13 , a learning data storage unit 14 , a learned object detection model storage unit 15 , and a communication unit 19 .

[0017] The processor 11 may include, for example, a GPU (Graphics Processing Unit) suitable for AI processing, as well as a CPU (Central Processing Unit) that performs main control of the detection processing terminal 10.

[0018] The processor 11 includes a learning unit 111 , a moving object detection unit 112 , an object detection unit 113 , and an image collection unit 114 .

[0019] The learning unit 111 is a functional block in which the processor 11 inputs training data into an AI algorithm to perform machine learning and generate a trained object detection model. The AI ​​algorithm may be, for example, but is not limited to, an SVM (Support Vector Machine), deep learning, a CNN (Convolutional Neural Network), or an R-CNN (Region-Based Convolutional Neural Network). The trained object detection model generated by the learning unit 111 is stored in the trained object detection model storage unit 15.

[0020] The moving object detection unit 112 is a functional block that performs video analysis processing on an image to be detected (hereinafter referred to as a "processing target image") and detects a subject captured in the moving image. Note that the image includes both moving images and still images, but in the case of a moving image, it may be referred to as a "processing target image." Algorithms that can be used for the moving object detection processing include, for example, optical flow, template matching, block matching, background subtraction, and background estimation using a Gaussian mixture model, as appropriate.

[0021] The object detection unit 113 inputs the image to be processed into a trained object detection model, classifies the attributes of the detected object, and adds a detection frame.

[0022] The image collection unit 114 extracts only images to be processed that satisfy preset collection conditions based on the subject detection results output by the subject detection unit 113, and stores the extracted images in the learning data storage unit 14. Image collection is a general term that refers to the image extraction process and storage process.

[0023] The images collected by the image collection unit 114 are stored in the training data storage unit 14. The stored images are then annotated by adding information (labels) indicating the object type, and the learning unit 111 inputs the images to the trained object detection model, thereby becoming input images for retraining the machine learning. Here, "retraining the machine learning" means that a trained object detection model is once generated using training data, and then the images collected by the image collection unit 114 are input to the trained object detection model as training data for training the machine learning.

[0024] The memory 12 includes a RAM (Random Access Memory), a ROM (Read Only Memory), a flash memory, etc., and functions as a data storage area and a work area for the detection processing terminal 10 .

[0025] The image storage unit 13 is a storage medium that stores images to be processed, and may be configured using, for example, a hard disk drive (HDD) or a solid state drive (SSD). The images to be processed are, for example, videos captured by a camera.

[0026] The training data accumulation unit 14 is a storage medium that extracts images or frames constituting a video that do not produce the desired detection result as a result of inputting the image to be processed into the trained subject detection model, and stores them for use as new training data, and may be configured using, for example, an HDD or SSD.

[0027] The learned object detection model storage unit 15 is a storage medium that stores the learned object detection model generated by the learning unit 111 and may be configured using, for example, an HDD or SSD. The learned object detection model storage unit 15 may be configured to train multiple types of objects into a single learned object detection model so that detection (labeling) of different types of objects can be performed using a common learned object detection model. In this case, the decision to save may be made by referring only to the selected labels to be saved from the detection results of multiple objects. Alternatively, a learned object detection model may be generated for each type of object. For example, if the objects to be detected include dogs, monkeys, pheasants, and humans, the learned object detection model storage unit 15 stores a dog object detection model trained using dog images as training data, a monkey object detection model trained using monkey images as training data, a pheasant object detection model trained using pheasant images as training data, and a human object detection model trained using human images as training data. Then, the setting device 20 sets which learned object detection model to use for object detection processing.

[0028] The communication unit 19 is a communication interface for communicating with the setting device 20, and may be, for example, a wired LAN interface, a Wi-Fi (registered trademark) interface, or an LTE, 4G, or 5G interface.

[0029] The setting device 20 includes a processor 21 , a memory 22 , a monitor 23 , an input unit 24 , an output unit 25 , a storage medium 26 , and a communication unit 29 .

[0030] The processor 11 may include a CPU (Central Processing Unit) that performs main control of the setting device 20 .

[0031] The processor 11 includes a storage type setting unit 211 and a learning data setting unit 212 .

[0032] The storage type setting unit 211 accepts an operation to set the collection conditions for images to be used as learning data. More specifically, in accordance with an input operation from the user, the storage type setting unit 211 can set the conditions under which the subject detection unit 113 performs AI detection processing of a subject using a trained subject detection model, and the conditions under which the image collection unit 114 collects images based on the results of the AI ​​detection processing by the subject detection unit 113. Furthermore, the storage type setting unit 211 also accepts a setting for whether or not to use moving object detection processing in combination.

[0033] The learning data setting unit 212 accepts a setting operation for the type of trained object detection model to be used for object detection processing by the object detection unit 113. The object detection unit 113 reads the set object detection model from the trained object detection model storage unit 15 and uses it for object detection processing.

[0034] The memory 22 includes a RAM (Random Access Memory), a ROM (Read Only Memory), a flash memory, etc., and functions as a data storage area and a work area for the detection processing terminal 10 .

[0035] The monitor 23 is a display device using a liquid crystal panel, an organic EL panel, or the like.

[0036] The input unit 24 is a device that accepts user operations, and may be, for example, a keyboard, a mouse, a touch panel, or the like.

[0037] The output unit 25 is a device that outputs the processing results of the processor 21, and the monitor 23 is one aspect of the output unit. In addition, when audio data is output, a speaker, an output terminal, etc. correspond to the output unit 25.

[0038] The storage medium 26 is a storage device that stores data in a non-volatile manner, and may be a flash memory, HDD, SSD, or the like.

[0039] The communication unit 29 is a communication interface for communicating with the detection processing terminal 10, and may be, for example, a wired LAN interface, a Wi-Fi (registered trademark) interface, or an LTE, 4G, or 5G interface.

[0040] FIG. 3 is a flowchart showing the flow of processing by the learning data collection device.

[0041] When the user sets the save type on the setting device 20 (S01) and operates the "Start" button for automatic image saving (S02), the detection processing terminal 10 starts the detection processing. The detection processing continues to be executed until the user operates the "Stop" button on the setting device 20 (S03: NO).

[0042] The detection processing terminal 10 acquires the candidate image for storage (image to be processed) (S04) and stores it in the image storage unit 13. For example, if the detection processing terminal 10 is a surveillance camera, the candidate image for storage may be acquired by reading in an image generated by an imaging unit (CCD, CMOS, etc.). Also, if the detection processing terminal 10 is a server, the candidate image for storage may be acquired by receiving an image from a remote camera or by reading in an image stored in another storage medium.

[0043] The subject detection unit 113 of the detection processing terminal 10 reads out the image (image to be processed) that is a candidate for storage from the image storage unit 13, inputs the image that is a candidate for storage into the learned subject detection model stored in the learned subject detection model memory unit 15, and performs detection processing of a specific subject that corresponds to the storage type set in step S01 (S05).

[0044] If "non-detection" is selected in the save type setting in step S01 (S06: YES) and the specified subject is detected (S07: YES), the process returns to step S03 and the detection process continues until the "stop" button is operated. On the other hand, if the specified subject is not detected (S07: NO), the detection processing terminal 10 saves the captured image in the learning data storage unit 14 (S08).

[0045] Until the number of saved images reaches the designated number (S09: NO), the process returns to step S03 and continues the detection process until the "Stop" button is operated.

[0046] In step S06, if "non-detection" is not selected in the save type setting (S06: NO) and the specified subject is detected (S10: YES), the detection processing terminal 10 saves the captured image in the learning data storage unit 14 (S08). On the other hand, if the specified subject is not detected (S10: NO), the process returns to step S03 and the detection process continues until the "stop" button is operated.

[0047] If the user operates the "Stop" button (S03: YES), or if the number of images stored in the learning data storage unit 14 reaches the designated number (S09: YES), the detection process is stopped.

[0048] FIG. 4 is a diagram showing an example of a save type setting screen.

[0049] 4 is displayed on the monitor 23 of the setting device 20. When the user inputs each item on the storage type setting screen 30 using the input unit 24, the storage type setting unit 211 sets the input item as the storage type. The setting operation on the storage type setting screen 30 corresponds to the processing of step S01 in FIG.

[0050] As shown in Figure 4, the save type setting screen 30 includes, as an example, a "Save learning images" operation setting button 31, a "Save interval" setting button 32, a "Number of saves" setting button 33, a "Start" button 34, a progress bar 35, a "Save label" setting button 36, and a "Select" button 37.

[0051] The "Save learning images" operation setting button 31 can be set by the user by operating the pull-down button 311 to either "automatic save," which causes the detection processing terminal 10 to automatically collect images and store them in the learning data storage unit 14, or "manual save," which causes the user to manually store images.

[0052] The "save interval" setting button 32 allows the user to set the time interval at which the detection processing terminal 10 collects images by operating a pull-down button 321. The "save interval" can be set to, for example, "when no detection occurs," "when detection occurs," "1 minute," "2 minutes," etc.

[0053] The "Number of images to be saved" setting button 33 allows the user to set the maximum number of images to be saved collected by the detection processing terminal 10 by operating a pull-down button 331. The number of images to be saved set here corresponds to the "specified number of images" in step S09 in FIG. 3.

[0054] Furthermore, when each of the information buttons 312, 322, and 332 is operated, a pop-up is displayed, for example, explaining each setting item of the "Save learning images" operation setting button 31, the "Save interval" setting button 32, and the "Number of saves" setting button 33. The user can look at this explanation, set the desired collection conditions on the save type setting screen 30, and have the detection processing terminal 10 execute them.

[0055] The "Start" button 34 is a button for instructing the start of the detection process, and is the button operated in step S02 of FIG.

[0056] The progress bar 35 indicates the percentage of the detection process completed for the storage candidate images (processing target images) after the detection processing terminal 10 starts the detection process.

[0057] The "Save Label" setting button 36 is a button for inputting the processing target image into the trained object detection model to set the type of object to be detected. The save type setting screen 30 in FIG. 4 provides label buttons for the labels to be saved: a "Dog" button 361, a "Monkey" button 362, a "Pheasant" button 363, and a "Human" button 364. The user can set the object to be detected by activating one of the label buttons and operating the "Select" button 37. It is preferable to set only one save label to be selectable. This is because selecting multiple labels complicates the object detection process in the object detection unit 113, which may result in a decrease in detection accuracy or a momentary increase in processing load. An object corresponding to a save label set by pressing the "Save Label" setting button 36 is also referred to as a "specific object."

[0058] (Processing Example 1) Non-Detection: Case in which Images are Collected When a Subject is Not Detected (Collection of False Negative Images) FIG. 5 is a diagram showing an example of saving false negative images. When collecting false negative images, images are collected in which the subject corresponding to the selected label was captured (positive), but the subject was not detected. Therefore, on the save type setting screen 30, as shown in FIG. 4, the "Save Interval" setting button 32 is set to "Non-Detection," and after the determination of "Is the specified subject detected?" in step S07 of FIG. 3 is negative, the images are extracted and saved in step S08.

[0059] When "dog" 361 is selected as the saved label, object detection unit 113 reads from trained object detection model storage unit 15 a trained object detection model that has been trained using images of a dog as training data. Frames f1, f2, f3, and f4 that make up the video to be processed are then sequentially input into the read trained object detection model to label the subject. Once the dog has been labeled as "dog," an image with a detection frame 40 surrounding the dog is output as the object detection result. Note that adding detection frame 40 is not essential when collecting images automatically; any data that indicates whether or not a subject can be detected may be used in place of detection frame 40.

[0060] In FIG. 5, frames f1, f2, and f4 are provided with detection frames 40, but frame f3 is not provided with a detection frame.

[0061] Therefore, because the image collection unit 114 is set to collect images "during non-detection," frames f1, f2, and f4 are not stored as learning data (step S07: YES in FIG. 3 ), and frame f3 is stored in the learning data storage unit 14 (steps S07: NO, S08 in FIG. 3 ). Frame f3 is one of the frames constituting the video to be processed, and is also one of the frames constituting the video output as the object detection result. Since the video output as the object detection result includes data of the video to be processed even if a detection frame is attached to the frame constituting the video to be processed, the mode of collecting frames constituting the video of the object detection result is included in the mode of collecting frames constituting the video to be processed.

[0062] According to this example, it is possible to store only images that should have been detected as subjects but were not, i.e., false negative images. These false negative images are automatically collected, and the learning unit 111 uses them again as training data to input them into a trained subject detection model with a dog as the subject for machine learning, thereby improving the learning accuracy of the trained subject detection model. Furthermore, according to this example, because the detection processing terminal 10 automatically collects false negative images, it is possible to collect the necessary re-learning data even if the user does not have sufficient knowledge of what kind of images to collect for machine learning. In this case, only frames that are false negative images are extracted and stored from the images to be processed, thereby preventing the storage capacity of the collected images from becoming excessively large.

[0063] (Processing Example 2) Upon Detection: Case in which Images are Collected When a Subject is Falsely Detected (Collection of False-Positive Images) FIG. 6 is a diagram showing an example of a save type setting screen when automatically collecting false-positive images. FIG. 7 is a diagram showing an example of saving false-positive images. When collecting false-positive images, images in which a subject is detected despite the subject corresponding to the selected label not being captured are collected. Therefore, on the save type setting screen 30, as shown in FIG. 6, the "Save Interval" setting button 32 is set to "Upon Detection," and "Human" 364 is selected as the label to be saved. In the process of collecting false-positive images, after a positive determination is made in step S10 of FIG. 3 regarding "the specified subject has been detected," images are collected and saved in step S08.

[0064] The user uses the training data setting unit 212 to read an image of a dog from the trained object detection model storage unit 15, which has been trained using teacher data. The object detection unit 113 sequentially inputs frames f5, f6, f7, and f8 that make up the video to be processed into the trained object detection model, which has been machine-learned using images of dogs as subjects, and labels them as "human." Once the "human" labeling is complete, the unit outputs the video with a detection frame 40 surrounding the "human" added to the video.

[0065] 7, frames f5, f6, and f8 are not provided with a detection frame 40, but frame f7 is provided with a detection frame 40. In other words, in frame f7, the dog is detected as a person.

[0066] Therefore, since the image collection unit 114 is set to collect images "at the time of detection," frames f5, f6, and f8 are not stored as learning data (step S10: NO in Figure 3), and frame f7 is stored in the learning data storage unit 14 (step S10: YES, S08 in Figure 3).

[0067] According to this example, it is possible to store only images that should not have been detected as objects but were mistakenly detected, i.e., false positive images. These false positive images are automatically collected, and the learning unit 111 uses them again as training data to train the trained object detection model, thereby improving the learning accuracy of the trained object detection model. Furthermore, according to this example, because the detection processing terminal 10 automatically collects false positive images, necessary re-learning data can be collected even if the user does not have sufficient knowledge about what kind of images should be collected for machine learning. In this case, only frames that are false positive images are extracted and stored from the images to be processed, which prevents the storage capacity of the collected images from becoming excessively large.

[0068] (Processing Example 3) Non-detection: Case where images are collected when a subject is not detected and there is movement (collection of false negative images and images with movement) Fig. 8 is a diagram showing an example of a save type setting screen when images with movement are extracted and automatically collected from among false negative images. Fig. 9 is a diagram showing an example of saving images with movement from among false negative images.

[0069] In the processing example 3, as shown in FIG. 8, a "save only when motion is detected" button 381 and a "motion detection threshold" setting button 382 are provided on the save type setting screen 30a.

[0070] By checking the "Save only when motion is detected" button 381, only frames in which the moving object detection unit 112 has "determined there is motion" from the processing results of the subject detection unit 113 can be targeted for image collection.

[0071] In addition, a motion detection threshold for determining whether there is motion in the camera image can be set by inputting a value into the "Motion Detection Threshold" setting button 382. This allows the reference value for determining whether there is motion to be set. The lower the motion detection threshold, the less motion is required to determine there is motion, leading to an increase in the number of images collected. Therefore, it is desirable for the user to set the threshold appropriately, taking into account the balance between the data storage capacity of the detection processing terminal 10 and the accuracy of image collection.

[0072] In FIG. 9, frames f9 and f12 are assigned detection frames 40, but frames f10 and f11 are not assigned detection frames.

[0073] In the above-described processing example 1, frames f10 and f11 are the targets of image collection, but in this processing example, the moving object detection unit 112 performs moving object detection processing on the images to be processed to detect frames containing movement. In the example of Fig. 9, if frame f11 is the frame of interest, movement exceeding the motion detection threshold is detected for the subject captured in frame f11 relative to the subject captured in frame f10, the frame immediately preceding frame f10, and as a result, frame f11 is determined to contain "motion." A similar "motion presence determination" is made for frame f12.

[0074] The image collection unit 114 acquires the processing results of the moving object detection unit 112 and the processing results of the subject detection unit 113, and collects only frames f11 that are determined to have motion among the false negative images and stores them in the learning data storage unit 14.

[0075] According to Processing Example 3, when an image is detected as a false negative, by collecting only images that contain movement, it is possible to collect images that cannot be detected because the shape of the subject changes suddenly, such as immediately after the subject starts moving. This allows for re-learning using images in which the detection accuracy is likely to decrease due to sudden changes in movement in object detection processing using a trained object detection model, thereby improving object detection accuracy.

[0076] The aspect in which the image collection unit 114 determines images to be collected by combining both the processing results of the moving object detection unit 112 and the processing results of the subject detection unit 113 can also be applied to the case of false positive images. In other words, when a false positive image is detected, only images that show movement may be collected, so that images that are falsely detected due to a sudden change in the shape of the subject, such as immediately after the subject starts moving, can be collected specifically.

[0077] As described above, according to this embodiment, the setting tool assists the user in determining what image data to collect for a trained model that has accuracy issues, and the necessary images can be automatically saved.

[0078] In addition, since only images that satisfy the collection conditions are collected, it is possible to accumulate a minimum number of images without collecting a large number of images.Furthermore, since the images used in the next machine learning are narrowed down to a minimum number of images, the learning time is shortened.

[0079] The present invention is expected to have a particularly significant effect when the detection processing terminal is an edge AI camera, which tends to have smaller storage capacity and lower processing performance compared to a server.

[0080] The above-described embodiment is not intended to limit the present invention, and modifications that do not deviate from the spirit of the present invention are also included in the present invention. For example, the present invention may be used in combination with the following monitoring device. In such a case, modifications may be made by deleting unnecessary components.

[0081] For example, although the detection processing terminal 10 and the setting device 20 have been described above as separate devices, the detection processing terminal 10 and the setting device 20 may be configured as the same device, or may have a relationship between a local terminal and a cloud terminal.

[0082] In the above example, a single type of subject is captured in the image to be processed, and a single type of save type label is specified. However, in the case of an image to be processed that contains multiple subjects, images may be collected if there are differences between the images and a teacher model (high-performance model). Alternatively, the number of subjects may be specified, and images may be collected if there are differences.

[0083] As described above, this embodiment includes the following invention: (Supplementary Note 1) A training data collection device comprising: a save type setting unit that accepts a setting operation for collection conditions for images used as training data, an object detection unit that inputs processing target images to a trained model and outputs object detection results that detect a specific object, and an image collection unit that collects only the processing target images that satisfy the collection conditions based on the object detection results.

[0084] (Supplementary Note 2) The learning data collection device according to Supplementary Note 1, wherein the save type setting unit accepts a setting to collect only images in which the specific subject is not detected from the processing target image, or to collect only images in which the specific subject is detected from the processing target image, as the collection condition.

[0085] (Supplementary Note 3) The training data collection device according to Supplementary Note 1, wherein the processing target image is a video including a plurality of frames, the object detection unit sequentially inputs each frame to the trained model and outputs the object detection result for each frame, and the image collection unit collects only the frames that satisfy the collection condition.

[0086] (Supplementary Note 4) The learning data collection device according to Supplementary Note 3, further comprising a motion detection unit that executes motion detection processing on the processing target image to detect a moving subject, and the save type setting unit accepts a setting as the collection condition to collect only when, among frames in which the specific subject is not detected, the subject captured in a target frame is determined to be moving based on the subject captured in a frame immediately preceding the target frame.

[0087] (Supplementary Note 5) The training data collection device according to Supplementary Note 1, further comprising a learning unit that performs machine learning on the trained model, wherein the learning unit inputs images collected by the image collection unit into the trained model and performs machine learning again.

[0088] (Supplementary Note 6) The learning data collection device according to Supplementary Note 1, comprising: a detection processing terminal; and a setting device communicatively connected to the detection processing terminal; the setting device comprising the storage type setting unit; and the detection processing terminal comprising the subject detection unit and the image collection unit.

[0089] (Supplementary Note 7) The learning data collection device according to Supplementary Note 6, wherein the detection processing terminal is a camera, the setting device is an operation terminal of the detection processing terminal, and the processing target image is an image captured by the camera.

[0090] (Supplementary Note 8) The learning data collection device according to Supplementary Note 6, wherein the detection processing terminal is a server, the setting device is an operation terminal of the server, and the processing target image is a video image stored in the server.

[0091] (Supplementary Note 9) A training data collection method including the steps of: inputting a processing target image into a trained model and outputting a subject detection result in which a specific subject is detected; and extracting only the processing target image that satisfies a predetermined collection condition for images to be used as training data based on the subject detection result, and storing the extracted image in a storage medium.

[0092] 1: Learning data collection device, 10: Detection processing terminal, 11: Processor, 12: Memory, 13: Image storage unit, 14: Learning data storage unit, 15: Learned object detection model storage unit, 19: Communication unit, 20: Setting device, 21: Processor, 22: Memory, 23: Monitor, 24: Input unit, 25: Output unit, 26: Storage medium, 29: Communication unit, 30: Save type setting screen, 30a: Save type setting screen, 31: Learning image save operation setting button, 32: Save interval setting button, 33: Save number setting button, 34: Start button, 35: Progress bar, 36: Save label setting button, 37: Selection selection button, 40: detection frame, 111: learning unit, 112: motion detection unit, 113: subject detection unit, 114: image collection unit, 211: save type setting unit, 212: learning data setting unit, 311: pull-down button, 312: information button, 321: pull-down button, 322: information button, 331: pull-down button, 332: information button, 361: dog button, 362: monkey button, 363: pheasant button, 364: human button, 381: save only when motion is detected button, 381, 382: motion detection threshold setting button, f1 to f12: frames

Claims

1. A learning data collection device, comprising: a storage type setting unit that receives a setting operation for collection conditions of an image to be used as learning data; a subject detection unit that inputs a processing target image to a learned model and outputs a subject detection result of detecting a specific subject; and an image collection unit that collects only the processing target images that satisfy the collection conditions based on the subject detection result.

2. The learning data collection device according to claim 1, wherein the storage type setting unit receives a setting for collecting only images in which the specific subject is not detected from the processing target image or only images in which the specific subject is detected from the processing target image as the collection conditions.

3. The learning data collection device according to claim 1, wherein the processing target image is a video including a plurality of frames, the subject detection unit sequentially inputs each frame to the learned model and outputs the subject detection result for each frame, and the image collection unit collects only the frames that satisfy the collection conditions.

4. The learning data collection device according to claim 3, further comprising a moving object detection unit that performs a moving object detection process on the processing target image and detects a moving object as a subject, wherein the storage type setting unit receives a setting for collecting only when it is determined that a subject imaged in a target frame among the frames in which the specific subject is not detected has movement based on a subject imaged in a frame immediately before the target frame as the collection conditions.

5. The learning data collection device according to claim 1, further comprising a learning unit that performs machine learning of the learned model, wherein the learning unit inputs the images collected by the image collection unit to the learned model and performs machine learning again.

6. The learning data collection device according to claim 1, comprising a detection processing terminal and a setting device communicatively connected to the detection processing terminal, wherein the setting device includes the storage type setting unit, and the detection processing terminal includes the subject detection unit and the image collection unit.

7. The learning data collection device according to claim 6, wherein the detection processing terminal is a camera, the setting device is an operation terminal of the detection processing terminal, and the image to be processed is an image captured by the camera. Learning data collection device.

8. The learning data collection device according to claim 6, wherein the detection processing terminal is a server, the setting device is an operation terminal of the server, and the image to be processed is an image stored in the server. Learning data collection device.

9. A learning data collection method, comprising: a step in which a processor inputs a processing target image into a learned model and outputs a subject detection result in which a specific subject is detected; and based on the subject detection result, only the processing target image that satisfies the collection condition of the image to be used as predetermined learning data is extracted and stored in a storage medium. Learning data collection method.

Citation Information

Patent Citations

  • Image file generating apparatus, image file generating method, image management apparatus, and image management method

    JP2020123174A

  • Image search device and teacher data extraction method

    JP2020135494A

  • Information processor, detector, information processing method, and program

    JP2023120854A