Information processing system and information processing method
By displaying and capturing correct answer images with a blurred imaging device to generate accurate ground truth information, the method addresses the challenge of training machine learning models with blurred images, enhancing recognition accuracy and privacy protection.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-04-19
- Publication Date
- 2026-04-08
AI Technical Summary
Existing machine learning models trained with blurred images from multi-pinhole cameras face challenges in assigning accurate ground truth information, leading to decreased recognition accuracy and privacy concerns due to the difficulty in recognizing these images.
A method involving displaying correct answer images on a screen, capturing them with a blurred imaging device, and generating corresponding accurate ground truth information to create a dataset for training, ensuring privacy protection.
Improves the recognition accuracy of machine learning models while protecting subject privacy by generating accurate ground truth information based on blurred images.
Smart Images

Figure 0007842749000001 
Figure 0007842749000002 
Figure 0007842749000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to a technique for creating a dataset used for training a machine learning model.
Background Art
[0002] For example, Non-Patent Document 1 describes a method for creating a dataset for training a face detection model using a lensless camera by displaying an image captured by a normal camera on a display and capturing the image displayed on the display with the lensless camera.
[0003] However, since the training images captured by the camera that acquires the blurred images are images that are difficult for humans to recognize, it is difficult to assign accurate correct answer information to the captured training images. In addition, the inaccuracy of the correct answer information leads to a deterioration in the performance of the machine learning model to be trained. Therefore, even if a machine learning model is trained using a dataset including training images to which inaccurate correct answer information is assigned, it has been difficult to improve the recognition accuracy of the machine learning model while protecting the privacy of the subject.
Prior Art Documents
Non-Patent Documents
[0004]
Non-Patent Document 1
Summary of the Invention
[0005] The present disclosure has been made to solve the above problems, and an object thereof is to provide a technique capable of improving the recognition accuracy of a machine learning model while protecting the privacy of a subject.
[0006] The information processing system according to the present disclosure includes a correct answer information acquisition unit that acquires first correct answer information corresponding to a training image of a machine learning model from a first storage unit, a correct answer image display control unit that causes a display device to display a first correct answer image based on the first correct answer information, an imaging control unit that causes an imaging device that acquires a blurred image to image the first correct answer image displayed on the display device to acquire a second correct answer image, a generation unit that generates second correct answer information based on the second correct answer image, and a storage control unit that stores a data set including a pair of the training image and the second correct answer information in a second storage unit.
[0007] According to the present disclosure, it is possible to improve the recognition accuracy of a machine learning model while protecting the privacy of a subject.
Brief Description of Drawings
[0008] [Figure 1] It is a block diagram showing an example of the overall configuration of an imaging system according to a first embodiment of the present disclosure. [Figure 2] It is a diagram schematically showing the structure of a multi-pinhole camera which is an example of an imaging device. [Figure 3] It is a flowchart for explaining the data set creation process in the imaging control device according to the first embodiment of the present disclosure. [Figure 4] It is a diagram showing an example of a first training image displayed on a display device. [Figure 5] It is a diagram showing an example of a second training image obtained by imaging the first training image shown in FIG. 4 by an imaging device when the distance between the imaging device and the display device is L1. [Figure 6] It is a diagram showing an example of a second training image obtained by imaging the first training image shown in FIG. 4 by an imaging device when the distance between the imaging device and the display device is L2 (L1 < L2). [Figure 7] It is a figure in which the bounding box set for the first training image shown in FIG. 4 is superimposed on the second training image shown in FIG. 5. [Figure 8] It is a figure in which the bounding box set for the first training image shown in FIG. 4 is superimposed on the second training image shown in FIG. 6. [Figure 9] It is a figure showing an example of the first correct image displayed on the display device in the first embodiment. [Figure 10] It is a figure showing an example of the second correct image obtained by the imaging device imaging the first correct image shown in FIG. 9 when the distance between the imaging device and the display device is L1. [Figure 11] It is a figure showing an example of the second correct image obtained by the imaging device imaging the first correct image shown in FIG. 9 when the distance between the imaging device and the display device is L2 (L1 < L2). [Figure 12] It is a figure in which an external bounding box is superimposed on the second training image shown in FIG. 5. [Figure 13] It is a figure in which an external bounding box is superimposed on the second training image shown in FIG. 6. [Figure 14] It is a figure showing an example of the first correct image displayed on the display device in a modification of the first embodiment. [Figure 15] It is a figure showing an example of the second correct image obtained by the imaging device imaging the first correct image shown in FIG. 14 when the distance between the imaging device and the display device is L1. [Figure 16] It is a figure showing an example of the second correct image obtained by the imaging device imaging the first correct image shown in FIG. 14 when the distance between the imaging device and the display device is L2 (L1 < L2). [Figure 17] It is a figure showing an example of a rectangular external bounding box that circumscribes the second bounding box area shown in FIG. 15. [Figure 18] It is a figure showing an example of a rectangular external bounding box that circumscribes the second bounding box area shown in FIG. 16. [Figure 19] It is a diagram in which the circumscribed bounding box shown in FIG. 17 is superimposed on the second training image shown in FIG. 5. [Figure 20] It is a diagram in which the circumscribed bounding box shown in FIG. 18 is superimposed on the second training image shown in FIG. 6. [Figure 21] It is a block diagram showing an example of the overall configuration of the imaging system according to the second embodiment of the present disclosure. [Figure 22] It is a flowchart for explaining the dataset creation process in the imaging control device according to the second embodiment of the present disclosure. [Figure 23] It is a diagram showing an example of the first geometric image displayed on the display device in the second embodiment. [Figure 24] It is a diagram showing an example of the second geometric image obtained by the imaging device 4 imaging the first geometric image shown in FIG. 23. [Figure 25] It is a diagram showing another example of the first geometric image displayed on the display device. [Figure 26] It is a diagram showing an example of the conversion table in the second embodiment. [Figure 27] It is a schematic diagram for explaining the generation process of the second correct answer information in the second embodiment. [Figure 28] It is a diagram showing an example of the first geometric image displayed on the display device in the first modification of the second embodiment. [Figure 29] It is a diagram showing an example of the second geometric image obtained by the imaging device imaging the first geometric image shown in FIG. 28. [Figure 30] It is a diagram showing an example of the first geometric image including the first horizontal line displayed on the display device in the second modification of the second embodiment. [Figure 31] It is a diagram showing an example of the second geometric image obtained by the imaging device imaging the first geometric image shown in FIG. 30. [Figure 32] It is a diagram showing an example of the first geometric image including the first vertical line displayed on the display device in the second modification of the second embodiment. [Figure 33] This figure shows an example of a second geometric image obtained by the imaging device capturing the first geometric image shown in Figure 32. [Figure 34] This figure shows an example of the first correct answer information in a second modified example of the second embodiment. [Figure 35] This figure shows an example of the second correct answer information in a second modified example of the second embodiment. [Figure 36] This block diagram shows an example of the overall configuration of the imaging system according to the third embodiment of this disclosure. [Modes for carrying out the invention]
[0009] (Knowledge that forms the basis of this disclosure) In homes and indoors, various recognition technologies are important, such as recognizing the actions of people in the environment or recognizing the person operating a device. In recent years, a technology called deep learning has attracted attention for object recognition. Deep learning is a machine learning method that uses a multi-layered neural network, and by utilizing a large amount of training data, it is possible to achieve higher accuracy in recognition performance compared to conventional methods. Image information is particularly effective in such object recognition. Various methods have been proposed that significantly improve conventional object recognition capabilities by using a camera as an input device and performing deep learning with image information as input.
[0010] However, placing cameras inside homes and other public spaces presents the challenge of privacy violations if captured images are leaked externally due to hacking or other means. Therefore, measures are needed to protect the privacy of the subjects even if captured images are leaked externally.
[0011] For example, a multi-pinhole camera is used to obtain blurred images that are difficult for humans to visually perceive. Images captured by a multi-pinhole camera are intentionally blurred and difficult for humans to visually perceive due to the superposition of multiple images from different viewpoints, or because the subject image is difficult to focus on due to the absence of lenses. Therefore, images captured by a multi-pinhole camera are particularly suitable for building image recognition systems in environments where privacy protection is necessary, such as homes or indoors.
[0012] In this image recognition system, a multi-pinhole camera captures an image of the target area, and the captured image is input to a classifier. The classifier then uses a trained identification model to identify faces contained in the input image. In this way, because the target area is captured by a multi-pinhole camera, even if the captured image is leaked externally, the image is difficult for humans to visually recognize, thus protecting the privacy of the subject.
[0013] To train such a classifier, the imaging method described in Non-Patent Document 1 above creates a training dataset by displaying images captured by a regular camera on a display, and then capturing the images displayed on the display with a lensless camera. The identification task of the above image recognition system is not a detection process, but a classification process such as object identification or face recognition. When the identification task is a classification process, the correct information used for training is assigned to each image, so pre-prepared correct information can be used. On the other hand, when the identification task is a detection process such as object detection or region segmentation, correct information such as bounding boxes indicating the position of the object to be detected on the image needs to be assigned at the pixel level, not per image.
[0014] However, training images captured with multi-pinhole cameras or lensless cameras are difficult for humans to recognize, making it challenging to assign accurate ground truth information to these images. Furthermore, inaccuracies in the ground truth information lead to a decrease in the performance of the trained machine learning model. Therefore, even if a machine learning model is trained using a dataset containing training images with inaccurate ground truth information, it has been difficult to improve the recognition accuracy of the machine learning model while protecting the privacy of the subjects.
[0015] To address these challenges, the inventors devised an information processing method in which, during the stage of accumulating training datasets, not only training images but also correct images based on correct information are displayed on a screen, and correct information is generated based on images obtained by capturing the displayed correct images. This method allows for the acquisition of accurate correct information, and the inventors have found that it is possible to improve the recognition accuracy of machine learning models while protecting the privacy of subjects. This led to the present disclosure.
[0016] To solve the above problems, an information processing system according to one aspect of the present disclosure includes: a correct answer information acquisition unit that acquires first correct answer information corresponding to training images for a machine learning model from a first storage unit; a correct answer image display control unit that displays the first correct answer image based on the first correct answer information on a display device; an imaging control unit that causes an imaging device that acquires blurred images to capture the first correct answer image displayed on the display device and acquire a second correct answer image; a generation unit that generates second correct answer information based on the second correct answer image; and a storage control unit that stores a dataset including the training image and the second correct answer information in a second storage unit.
[0017] In this configuration, a first ground truth image based on first ground truth information corresponding to the training images for the machine learning model is displayed on the display device. Then, the first ground truth image displayed on the display device is captured by an imaging device that acquires blurred images, and a second ground truth image is obtained. Second ground truth information is generated based on the acquired second ground truth image. Then, a dataset containing the training images and the second ground truth information is stored in the second storage unit.
[0018] Therefore, by generating accurate second-hand information according to the degree of blurring of the imaging device and accumulating a dataset containing training images and pairs of second-hand information, it is possible to improve the recognition accuracy of machine learning models while protecting the privacy of the subjects.
[0019] Furthermore, in the above-described information processing system, the first correct answer image may include an object that indicates the position of the first correct answer information on the image.
[0020] With this configuration, the first ground truth image includes an object that indicates the position of the first ground truth information on the image. Since the position of the object to be detected on the image can be identified by the object, accurate second ground truth information can be generated, and the recognition accuracy of the machine learning model that performs object detection can be improved.
[0021] Furthermore, in the above-described information processing system, the object is a frame, the imaging control unit acquires a second ground truth image by superimposing multiple frames, and the generation unit selects one frame from the multiple frames based on the brightness or position of each of the multiple frames in the second ground truth image, and generates the ground truth information represented by the selected one frame as the second ground truth information.
[0022] With this configuration, the position of the object to be detected on the image can be determined by selecting one frame based on the brightness or position of each of the multiple superimposed frames. This allows for the generation of accurate second-hand information, thereby improving the recognition accuracy of machine learning models that perform object detection.
[0023] Furthermore, in the above-described information processing system, the object is a frame, the imaging control unit acquires the second ground truth image by superimposing a plurality of frames, and the generation unit identifies a circumscribing frame that circumscribes the plurality of frames from the second ground truth image, and generates the ground truth information represented by the identified circumscribing frame as the second ground truth information.
[0024] With this configuration, the position of the object to be detected on the image can be determined by the circumscribing frame that circumscribes the multiple superimposed frames, thereby enabling the generation of accurate second-order correct information and improving the recognition accuracy of machine learning models that perform object detection.
[0025] Furthermore, in the above-described information processing system, the object is a frame, the imaging control unit acquires a second ground truth image by superimposing multiple frames, the generation unit identifies a circumscribing frame that circumscribes the multiple frames from the second ground truth image, determines the center of the identified circumscribing frame as a reference position on the image, and generates ground truth information as second ground truth information, where the determined reference position is the center and the frame is of the same size as the frame.
[0026] With this configuration, a circumscribing frame that circumscribes multiple superimposed frames is identified, the center of the identified circumscribing frame is determined as the reference position on the image, and the position of the object to be detected on the image can be determined using a frame of the same size as the frame with the determined reference position as the center. This allows for the generation of accurate second ground truth information, improving the recognition accuracy of machine learning models that perform object detection.
[0027] Furthermore, in the above-described information processing system, the object is a first region, the imaging control unit acquires a second ground truth image by superimposing a plurality of first regions, and the generation unit determines a reference position on the image from the second ground truth image based on the brightness of the second region in which the plurality of first regions are superimposed, and generates ground truth information as the second ground truth information, which is represented by a region of the same size as the first region centered on the determined reference position.
[0028] With this configuration, a reference position on the image is determined based on the brightness of a second region in which multiple first regions are superimposed. The position of the object to be detected on the image can be identified using a region of the same size as the first region, centered on the determined reference position. This allows for the generation of accurate second ground truth information, improving the recognition accuracy of machine learning models that perform object detection.
[0029] Furthermore, in the above-described information processing system, the object is a first region, the imaging control unit acquires a second ground truth image by superimposing a plurality of first regions, and the generation unit identifies a circumscribing frame from the second ground truth image that circumscribes the second region in which the plurality of first regions are superimposed, and generates the ground truth information represented by the identified circumscribing frame as the second ground truth information.
[0030] With this configuration, the position of the object to be detected on the image can be determined by the circumscribing frame that circumscribes the second region, which is formed by the superposition of multiple first regions. This allows for the generation of accurate second ground truth information, thereby improving the recognition accuracy of machine learning models that perform object detection.
[0031] Furthermore, in the above-described information processing system, the object is a first region, the imaging control unit acquires a second ground truth image by superimposing a plurality of first regions, the generation unit identifies a circumscribing frame from the second ground truth image that circumscribes the second region in which the plurality of first regions are superimposed, determines the center of the identified circumscribing frame as a reference position on the image, and generates ground truth information as the second ground truth information, which is represented by a region of the same size as the first region with the determined reference position as its center.
[0032] With this configuration, a circumscribing frame is identified that circumscribes a second region formed by the superposition of multiple first regions. The center of this identified circumscribing frame is determined as the reference position on the image. The position of the object to be detected on the image can then be determined using a region of the same size as the first region, centered on the determined reference position. This allows for the generation of accurate second ground truth information, thereby improving the recognition accuracy of machine learning models that perform object detection.
[0033] Furthermore, in the above-described information processing system, the generation unit may identify a region encompassing multiple objects from the second correct image and generate the correct information represented by the identified region as the second correct information.
[0034] With this configuration, the position of the object to be detected on the image can be determined from the region encompassing multiple objects in the second ground truth image, thereby enabling the generation of accurate second ground truth information and improving the recognition accuracy of the machine learning model that performs object detection.
[0035] Furthermore, in the above-described information processing system, the training image stored in the first storage unit is a first training image without blur acquired by an imaging device different from the imaging device, and the system further comprises an image acquisition unit that acquires the first training image from the first storage unit, and a training image display control unit that displays the first training image on the display device, wherein the imaging control unit causes the imaging device to image the first training image displayed on the display device to acquire a second training image, and the storage control unit may store a dataset including the pair of the second training image and the second correct answer information in the second storage unit.
[0036] This configuration allows for the accumulation of a dataset containing pairs of second training images corresponding to the degree of blurring of the imaging device, and accurate second ground truth information corresponding to the degree of blurring of the imaging device.
[0037] Furthermore, the above-described information processing system may further include a training unit that trains the machine learning model using a dataset containing the training image and second correct answer information stored in the second memory unit.
[0038] With this configuration, the machine learning model is trained using a dataset that includes pairs of second training images and second ground truth information stored in the second memory unit. This improves the recognition ability of the machine learning model to recognize subjects from captured images with varying degrees of blur depending on the distance to the subject.
[0039] Furthermore, this disclosure can be implemented not only as an information processing system having the characteristic configuration described above, but also as an information processing method that performs characteristic processing corresponding to the characteristic configuration of the information processing system. Therefore, the same effects as the above-described information processing system can be achieved in the following other embodiments.
[0040] An information processing method according to another aspect of the present disclosure involves a computer acquiring first correct information corresponding to training images for a machine learning model from a first storage unit, displaying the first correct image based on the first correct information on a display device, having an imaging device that acquires blurred images capture the first correct image displayed on the display device to acquire a second correct image, generating second correct information based on the second correct image, and storing a dataset including the training image and the second correct information in a second storage unit.
[0041] An information processing system according to another aspect of the present disclosure includes: an acquisition unit that acquires first correct information corresponding to training images for a machine learning model from a first storage unit; a geometric image display control unit that displays a first geometric image on a display device; an imaging control unit that causes an imaging device that acquires blurred images to capture the first geometric image displayed on the display device and acquire a second geometric image; a conversion table generation unit that generates a conversion table for converting the position of the first geometric image to the position of the second geometric image; a conversion unit that converts the first correct information to second correct information using the conversion table; and a storage control unit that stores a dataset including the training image and the second correct information in a second storage unit.
[0042] In this configuration, the first geometric image displayed on the display device is captured by an imaging device that acquires blurred images, thereby acquiring a second geometric image. A conversion table is generated to convert the position of the first geometric image to the position of the second geometric image. Then, using the conversion table, the first ground truth information corresponding to the training images of the machine learning model is converted into second ground truth information, and a dataset containing the pairs of training images and second ground truth information is stored in the second storage unit.
[0043] Therefore, by generating accurate second-hand information according to the degree of blurring of the imaging device and accumulating a dataset containing training images and pairs of second-hand information, it is possible to improve the recognition accuracy of machine learning models while protecting the privacy of the subjects.
[0044] Furthermore, in the above-described information processing system, the first geometric image may include a first dot positioned at a predetermined location on the image, and the conversion table generation unit may identify the positions of a plurality of second dots from the second geometric image and generate a conversion table for converting the position of the first dot to the positions of the identified plurality of second dots.
[0045] In this configuration, the first geometric image includes a first dot placed at a predetermined position on the image. The positions of multiple second dots are identified from the second geometric image, and a conversion table is generated to convert the position of the first dot to the positions of the identified multiple second dots. Using the conversion table, for example, the position of the frame represented by the first ground truth information is converted to a position of the frame corresponding to the degree of blur of the imaging device, and the converted frame position is generated as the second ground truth information. Therefore, accurate second ground truth information corresponding to the degree of blur of the imaging device can be generated.
[0046] Furthermore, in the above-described information processing system, the first geometric image includes a first horizontal line and a first vertical line positioned at predetermined locations on the image, and the conversion table generation unit may identify the positions of a plurality of second horizontal lines and a plurality of second vertical lines from the second geometric image, convert the position of the first horizontal line to the positions of the identified plurality of second horizontal lines, and generate a conversion table for converting the position of the first vertical line to the positions of the identified plurality of second vertical lines.
[0047] In this configuration, the first geometric image includes a first horizontal line and a first vertical line positioned at predetermined locations on the image. The positions of multiple second horizontal lines and multiple second vertical lines are identified from the second geometric image, and a conversion table is generated for converting the position of the first horizontal line to the positions of the identified multiple second horizontal lines, and for converting the position of the first vertical line to the positions of the identified multiple second vertical lines. Using the conversion table, for example, the position of the frame represented by the first ground truth information is converted to the position of the frame according to the degree of blur of the imaging device, and the converted position of the frame is generated as the second ground truth information. Therefore, accurate second ground truth information according to the degree of blur of the imaging device can be generated.
[0048] Embodiments of this disclosure will be described below with reference to the attached drawings. Note that the following embodiments are merely examples of the disclosure and do not limit the technical scope of this disclosure.
[0049] (First Embodiment) Figure 1 is a block diagram showing an example of the overall configuration of the imaging system 1 according to the first embodiment of this disclosure.
[0050] The imaging system 1 comprises an imaging control device 2, a display device 3, and an imaging device 4.
[0051] The display device 3 is, for example, a liquid crystal display device or an organic EL (Electro-Luminescence) display device. The display device 3 is controlled by the imaging control device 2 and displays the image output from the imaging control device 2. The display device 3 may also be a projector that projects images onto a screen.
[0052] The imaging device 4 is a computational imaging camera such as a lensless camera, a coded aperture camera, a multi-pinhole camera, a lensless multi-pinhole camera, or a light field camera. The imaging device 4 acquires a blurred image by imaging.
[0053] The imaging device 4 is positioned to capture the display screen of the display device 3. In this first embodiment, the imaging device 4 is a lensless multi-pinhole camera in which a mask having a mask pattern with multiple pinholes is positioned to cover the light-receiving surface of the image sensor. In other words, the mask pattern can be said to be positioned between the subject and the light-receiving surface.
[0054] Unlike a normal camera that captures a normal image without blur, the imaging device 4 captures a computational image, which is a blurred image. A computational image is an image in which the subject cannot be recognized by a human eye due to the intentionally created blur.
[0055] Figure 2 is a schematic diagram showing the structure of a multi-pinhole camera 200, which is an example of an imaging device 4. Figure 2 is a top view of the multi-pinhole camera 200.
[0056] The multi-pinhole camera 200 shown in Figure 2 comprises a multi-pinhole mask 201 and an image sensor 202 such as a CMOS. The multi-pinhole camera 200 does not have a lens. The multi-pinhole mask 201 is positioned at a certain distance from the light-receiving surface of the image sensor 202. The multi-pinhole mask 201 has a plurality of pinholes 2011, 2012 arranged randomly or at equal intervals. The plurality of pinholes 2011, 2012 are also called a multi-pinhole. The image sensor 202 acquires an image by capturing the image displayed on the display device 3 through each pinhole 2011, 2012. The image acquired through the pinholes is also called a pinhole image.
[0057] The pinhole image of the subject differs depending on the position and size of each pinhole 2011, 2012. Therefore, the image sensor 202 acquires a superimposed image in which multiple pinhole images are slightly shifted and overlapping (multiple images). The positional relationship of the multiple pinholes 2011, 2012 affects the positional relationship of the multiple pinhole images projected onto the image sensor 202 (i.e., the degree of superposition of the multiple images). The size of the pinholes 2011, 2012 affects the degree of blurring of the pinhole images.
[0058] By using the multi-pinhole mask 201, it is possible to acquire multiple pinhole images with different positions and degrees of blur by superimposing them. In other words, it is possible to acquire computationally captured images in which multiple images and blur are intentionally created. As a result, the captured image becomes a multiple-image and blurred image, and the privacy of the subject can be protected by this blur.
[0059] Furthermore, by changing the number of pinholes, their positions, and their sizes, images with different degrees of blurring can be obtained. In other words, the multi-pinhole mask 201 may have a structure that allows the user to easily attach and detach it. Multiple types of multi-pinhole masks 201 with different mask patterns may be prepared in advance. The multi-pinhole mask 201 may be freely replaced by the user according to the mask pattern of the multi-pinhole camera used during image recognition.
[0060] In addition to replacing the multi-pinhole mask 201, the following various methods can be used to modify the multi-pinhole mask 201. For example, the multi-pinhole mask 201 may be rotatably mounted in front of the image sensor 202 and may be rotated arbitrarily by the user. Alternatively, the multi-pinhole mask 201 may be created by the user drilling holes at any point on a plate mounted in front of the image sensor 202. Alternatively, the multi-pinhole mask 201 may be a liquid crystal mask utilizing a spatial light modulator or the like. A predetermined number of pinholes may be formed at predetermined locations by arbitrarily setting the transmittance of each position within the multi-pinhole mask 201. Furthermore, the multi-pinhole mask 201 may be molded using an expandable material such as rubber. The user may also change the position and size of the pinholes by physically deforming the multi-pinhole mask 201 by applying an external force.
[0061] The multi-pinhole camera 200 is also used in image recognition using a pre-trained machine learning model. Images captured by the multi-pinhole camera 200 are collected as training data. The collected training data is used to train the machine learning model.
[0062] Furthermore, although Figure 2 shows two pinholes 2011 and 2012 arranged horizontally, the disclosure is not limited thereto, and the multi-pinhole camera 200 may have three or more pinholes. Also, the two pinholes 2011 and 2012 may be arranged vertically.
[0063] The imaging control device 2 is specifically composed of a microprocessor, RAM (Random Access Memory), ROM (Read Only Memory), and a hard disk, which are not shown in the diagram. The RAM, ROM, or hard disk stores computer programs, and the functions of the imaging control device 2 are realized when the microprocessor operates according to the computer programs.
[0064] The imaging control device 2 comprises a first storage unit 21, a second storage unit 22, an image acquisition unit 23, a training image display control unit 24, a correct answer information acquisition unit 25, a correct answer image display control unit 26, an imaging control unit 27, a correct answer information generation unit 28, and a storage control unit 29.
[0065] The first storage unit 21 stores the first training images for the machine learning model and the first correct answer information corresponding to the first training images. The first storage unit 21 stores multiple first training images captured by a normal camera and the first correct answer information (annotation information) corresponding to each of the multiple first training images. The first training images are images that include the subject to be recognized by the machine learning model. The first training images are unblurred images acquired by an imaging device different from the imaging device 4.
[0066] The first ground truth information differs for each identification task. For example, if the identification task is object detection, the first ground truth information is the bounding box representing the area occupied by the detected object on the image. Also, for example, if the identification task is object recognition, the first ground truth information is the classification result. Also, for example, if the identification task is image region segmentation, the first ground truth information is the region information for each pixel. The first training image and the first ground truth information stored in the first storage unit 21 are the same information used in machine learning of classifiers that normally use cameras.
[0067] The image acquisition unit 23 acquires the first training image for the machine learning model from the first storage unit 21. The image acquisition unit 23 outputs the first training image acquired from the first storage unit 21 to the training image display control unit 24.
[0068] The training image display control unit 24 causes the first training image to be displayed on the display device 3. The display device 3 displays the first training image in accordance with the instructions from the training image display control unit 24.
[0069] The correct answer information acquisition unit 25 acquires the first correct answer information corresponding to the first training image (training image) of the machine learning model from the first storage unit 21. The correct answer information acquisition unit 25 acquires the first correct answer information corresponding to the first training image acquired by the image acquisition unit 23. The correct answer information acquisition unit 25 outputs the first correct answer information acquired from the first storage unit 21 to the correct image display control unit 26.
[0070] The correct answer image display control unit 26 causes the display device 3 to display the first correct answer image based on the first correct answer information. The display device 3 displays the first correct answer image in accordance with the instructions from the correct answer image display control unit 26.
[0071] The first ground truth image includes an object that indicates the position of the first ground truth information on the image. For example, the object is a frame containing the object to be detected. The frame is rectangular and is also called a bounding box. The first ground truth information includes the position information of the bounding box on the image. That is, the first ground truth information includes the x-coordinate of the top-left vertex of the bounding box, the y-coordinate of the top-left vertex of the bounding box, the width of the bounding box, and the height of the bounding box. Note that the position information is not limited to the above, and any information that makes it possible to identify the position of the bounding box on the image is acceptable. In addition, the first ground truth information may include not only the position information of the bounding box but also the name (class name) of the object to be detected within the bounding box.
[0072] The imaging control unit 27 causes the imaging device 4 to capture the first training image displayed on the display device 3 to acquire the second training image. When the first training image is displayed on the display device 3, the imaging control unit 27 causes the imaging device 4 to capture the first training image to acquire the second training image. The imaging control unit 27 outputs the acquired second training image to the storage control unit 29.
[0073] Furthermore, the imaging control unit 27 causes the imaging device 4 to capture the first correct image displayed on the display device 3 to acquire a second correct image. When the first correct image is displayed on the display device 3, the imaging control unit 27 causes the imaging device 4 to capture the first correct image to acquire a second correct image. The imaging control unit 27 acquires a second correct image by superimposing multiple bounding boxes (frames). The imaging control unit 27 outputs the acquired second correct image to the correct information generation unit 28.
[0074] The ground truth information generation unit 28 generates second ground truth information based on the second ground truth image acquired by the imaging control unit 27. The ground truth information generation unit 28 identifies the circumscribing bounding boxes that circumscribe multiple bounding boxes (frames) from the second ground truth image, and generates the ground truth information represented by the identified circumscribing bounding boxes as the second ground truth information.
[0075] The memory control unit 29 stores a dataset in the second storage unit 22 that includes pairs of training images and second correct answer information. The memory control unit 29 stores a dataset in the second storage unit 22 that includes pairs of second training images obtained by imaging by the imaging device 4 and second correct answer information generated by the correct answer information generation unit 28.
[0076] The second memory unit 22 stores a dataset containing pairs of second training images and correct answer information.
[0077] Next, the dataset creation process in the imaging control device 2 according to the first embodiment of this disclosure will be described.
[0078] Figure 3 is a flowchart illustrating the dataset creation process in the imaging control device 2 according to the first embodiment of this disclosure.
[0079] First, the image acquisition unit 23 acquires a first training image from the first storage unit 21 (step S101). The image acquisition unit 23 then acquires a first training image that has not been captured from among the multiple first training images stored in the first storage unit 21.
[0080] Next, the training image display control unit 24 displays the first training image acquired by the image acquisition unit 23 on the display device 3 (step S102). The training image display control unit 24 instructs the display device 3 on the display position and size of the first training image. At this time, the training image display control unit 24 instructs the display device 3 on the display position and size of the first training image so that the image acquired by the imaging device 4 is the same size as the first training image.
[0081] Next, the imaging control unit 27 causes the imaging device 4 to capture the first training image displayed on the display device 3 to acquire a second training image (step S103). The imaging device 4 captures the image so that the display device 3 is within its field of view.
[0082] Next, the correct answer information acquisition unit 25 acquires the first correct answer information corresponding to the displayed first training image from the first storage unit 21 (step S104).
[0083] Next, the correct answer image display control unit 26 displays the first correct answer image on the display device 3 based on the first correct answer information acquired by the correct answer information acquisition unit 25 (step S105). At this time, if the position of the bounding box is represented in the same coordinate system as the first training image, the correct answer image display control unit 26 only needs to display the first correct answer image on the display device 3 with the bounding box drawn at the same position as the bounding box on the first training image. Details of the process for displaying the first correct answer image will be described later.
[0084] In this first embodiment, the first training image and the first correct answer image are displayed on the same display device 3, but this disclosure is not limited thereto, and the first training image and the first correct answer image may be displayed on different display devices.
[0085] Next, the imaging control unit 27 instructs the imaging device 4 to capture the first correct image displayed on the display device 3 to acquire the second correct image (step S106). The imaging device 4 captures the image so that the display device 3 is within its field of view.
[0086] Next, the correct answer information generation unit 28 generates second correct answer information based on the second correct answer image acquired by the imaging control unit 27 (step S107).
[0087] The imaging device 4 can acquire a second correct answer image for the blurred second training image acquired in step S103. When the imaging device 4 is a multi-pinhole camera, one subject is imaged as a multiple image by passing through the multi-pinhole mask 201. The parallax, which is the displacement amount of this multiple image, changes depending on the distance between the subject and the imaging device 4.
[0088] Here, referring to FIGS. 4 to 6, the parallax depending on the distance in the multi-pinhole camera will be described.
[0089] FIG. 4 is a diagram showing an example of the first training image displayed on the display device 3, FIG. 5 is a diagram showing an example of the second training image obtained by the imaging device 4 imaging the first training image shown in FIG. 4 when the distance between the imaging device 4 and the display device 3 is L1, and FIG. 6 is a diagram showing an example of the second training image obtained by the imaging device 4 imaging the first training image shown in FIG. 4 when the distance between the imaging device 4 and the display device 3 is L2 (L1 < L2).
[0090] In the first training image 51 shown in FIG. 4, a person 301 and a television 302 are shown. Also, a bounding box 321 is set for the person 301 which is the detection target. The bounding box 321 is not included in the first training image 51 displayed on the display device 3.
[0091] The second training image 52 shown in FIG. 5 is obtained by the imaging device 4 having two pinholes 2011 and 2012 imaging the first training image 51 shown in FIG. 4 when the distance between the imaging device 4 and the display device 3 is L1. Also, the second training image 53 shown in FIG. 6 is obtained by the imaging device 4 having two pinholes 2011 and 2012 imaging the first training image 51 shown in FIG. 4 when the distance between the imaging device 4 and the display device 3 is L2. However, L1 is shorter than L2.
[0092] In the second training image 52 shown in Figure 5, person 303 is person 301 on the first training image 51, captured by passing through a pinhole 2011 on the optical axis; television 304 is television 302 on the first training image 51, captured by passing through a pinhole 2011 on the optical axis; person 305 is person 301 on the first training image 51, captured by passing through a pinhole 2012 that is not on the optical axis; and television 306 is television 302 on the first training image 51, captured by passing through a pinhole 2012 that is not on the optical axis.
[0093] Furthermore, in the second training image 53 shown in Figure 6, person 307 is person 301 on the first training image 51, which was captured by passing through the pinhole 2011 on the optical axis; television 308 is television 302 on the first training image 51, which was captured by passing through the pinhole 2011 on the optical axis; person 309 is person 301 on the first training image 51, which was captured by passing through the pinhole 2012 which is not on the optical axis; and television 310 is television 302 on the first training image 51, which was captured by passing through the pinhole 2012 which is not on the optical axis.
[0094] Thus, the second training image obtained from the imaging device 4, which is a multi-pinhole camera, is an image in which multiple subject images are superimposed. The position and size of people 303, 307 and televisions 304, 308, which are captured by passing through the pinhole 2011 on the optical axis, do not change in the captured image. On the other hand, the position of people 305, 309 and televisions 306, 310, which are captured by passing through the pinhole 2012 which is not on the optical axis, changes in the captured image depending on the distance between the subject and the imaging device 4. The greater the distance between the subject and the imaging device 4, the smaller the amount of parallax. In other words, if the distance between the imaging device 4 and the display device 3 changes, a second training image with a changed amount of parallax is obtained.
[0095] FIG. 7 is a diagram in which the bounding box 321 set for the first training image 51 shown in FIG. 4 is superimposed on the second training image 52 shown in FIG. 5, and FIG. 8 is a diagram in which the bounding box 321 set for the first training image 51 shown in FIG. 4 is superimposed on the second training image 53 shown in FIG. 6.
[0096] The bounding box 322 shown in FIG. 7 and the bounding box 323 shown in FIG. 8 are superimposed at the same position as the bounding box 321 set for the first training image 51 shown in FIG. 4. The bounding box 322 and the bounding box 323 include the person 303 and the person 307 imaged through the pinhole 2011 on the optical axis. Therefore, the bounding box 322 is the correct information of the person 303 in the second training image 52, and the bounding box 323 is the correct information of the person 307 in the second training image 53.
[0097] On the other hand, the bounding box 322 and the bounding box 323 do not include the person 305 and the person 309 imaged through the pinhole 2012 not on the optical axis. Therefore, it can be seen that the bounding box 322 and the bounding box 323 are not the correct information of the person 303 and the person 307 due to the influence of parallax.
[0098] FIG. 9 is a diagram showing an example of the first correct image displayed on the display device 3 in the first embodiment. FIG. 10 is a diagram showing an example of the second correct image obtained by the imaging device 4 imaging the first correct image shown in FIG. 9 when the distance between the imaging device 4 and the display device 3 is L1, and FIG. 11 is a diagram showing an example of the second correct image obtained by the imaging device 4 imaging the first correct image shown in FIG. 9 when the distance between the imaging device 4 and the display device 3 is L2 (L1 < L2).
[0099] The first correct answer image 61 shown in Figure 9 includes a bounding box 331, which is an example of an object. The bounding box 331 is a rectangular frame that surrounds the person 301 in the first training image 51 shown in Figure 4 as the detection target. The bounding box 331 is displayed on the first correct answer image 61, which is the same size as the first training image 51 displayed on the display device 3.
[0100] Furthermore, in the second correct image 62 shown in Figure 10, bounding box 332 is bounding box 331 captured by passing through pinhole 2011 on the optical axis, and bounding box 333 is bounding box 331 captured by passing through pinhole 2012 which is not on the optical axis.
[0101] Furthermore, in the second correct image 63 shown in Figure 11, bounding box 334 is bounding box 331 captured by passing through pinhole 2011 on the optical axis, and bounding box 335 is bounding box 331 captured by passing through pinhole 2012 which is not on the optical axis.
[0102] Thus, the second ground truth images 62 and 63 obtained from the imaging device 4, which is a multi-pinhole camera, are images in which multiple images (bounding boxes) are superimposed. The position and size of the bounding boxes 332 and 334 captured by passing through the pinhole 2011 on the optical axis do not change on the second ground truth image. On the other hand, the position of the bounding boxes 333 and 335 captured by passing through the pinhole 2012, which is not on the optical axis, changes on the second ground truth image depending on the distance between the subject and the imaging device 4. The greater the distance between the subject and the imaging device 4, the smaller the parallax. In other words, if the distance between the imaging device 4 and the display device 3 changes, bounding boxes with a changed amount of parallax are obtained.
[0103] The ground truth information generation unit 28 detects the boundaries of multiple bounding boxes from the acquired second ground truth image by performing brightness-based binarization, edge detection, or filtering on the second ground truth image. The ground truth information generation unit 28 then identifies rectangular circumscribed bounding boxes that circumscribe multiple bounding boxes in the second ground truth image and generates the ground truth information represented by the identified circumscribed bounding boxes as the second ground truth information. Alternatively, the ground truth information generation unit 28 may identify rectangular inscribed bounding boxes that circumscribe multiple bounding boxes in the second ground truth image and generate the ground truth information represented by the identified inscribed bounding boxes as the second ground truth information.
[0104] Next, we will explain a rectangular bounding box that circumscribes multiple bounding boxes using Figures 12 and 13.
[0105] Figure 12 shows the second training image 52 shown in Figure 5 with the bounding box 336 superimposed, and Figure 13 shows the second training image shown in Figure 6 with the bounding box 337 superimposed.
[0106] The circumscribing bounding box 336 shown in Figure 12 is a frame that circumscribes the two bounding boxes 332 and 333 shown in Figure 10, and the circumscribing bounding box 337 shown in Figure 13 is a frame that circumscribes the two bounding boxes 334 and 335 shown in Figure 11.
[0107] The ground truth information generation unit 28 identifies a rectangular bounding box 336 that circumscribes multiple bounding boxes 332, 333. The bounding box 336 is ground truth information that includes a person 303 captured by passing through a pinhole 2011 on the optical axis and a person 305 captured by passing through a pinhole 2012 that is not on the optical axis. The ground truth information generation unit 28 generates the ground truth information represented by the identified bounding box 336 as second ground truth information.
[0108] Furthermore, the ground truth information generation unit 28 identifies a rectangular circumscribed bounding box 337 that circumscribes the multiple bounding boxes 334 and 335. The circumscribed bounding box 337 is ground truth information that includes a person 307 captured by passing through a pinhole 2011 on the optical axis and a person 309 captured by passing through a pinhole 2012 that is not on the optical axis. The ground truth information generation unit 28 generates the ground truth information represented by the identified circumscribed bounding box 337 as second ground truth information.
[0109] Thus, the imaging system 1 of this first embodiment acquires a second training image and a second ground truth image captured by the imaging device 4, and generates second ground truth information corresponding to the second training image based on the second ground truth image, thereby acquiring second ground truth information that matches the parallax. Therefore, a more accurate dataset can be constructed, and the recognition accuracy of the machine learning model can be improved.
[0110] Returning to Figure 3, the memory control unit 29 then stores a dataset in the second memory unit 22 that includes pairs of second training images acquired by the imaging device 4 and second correct answer information generated by the correct answer information generation unit 28 (step S108). As a result, the second memory unit 22 stores a dataset that includes pairs of second training images acquired by the imaging device 4 and their correct answer information.
[0111] Next, the image acquisition unit 23 determines whether the imaging device 4 has captured all of the first training images stored in the first storage unit 21 (step S109). If it is determined that the imaging device 4 has not captured all of the first training images (NO in step S109), the process returns to step S101, and the image acquisition unit 23 acquires the uncaptured first training images from among the multiple first training images stored in the first storage unit 21. Then, the training image display control unit 24 displays the first training images on the display device 3, and the imaging control unit 27 causes the imaging device 4 to capture the displayed first training images. The correct answer information acquisition unit 25 acquires the first correct answer information corresponding to the first training images. Then, the correct answer image display control unit 26 displays the first correct answer image based on the first correct answer information on the display device 3, and the imaging control unit 27 causes the imaging device 4 to capture the displayed first correct answer image.
[0112] On the other hand, if it is determined that the imaging device 4 has captured all of the first training images (YES in step S109), the process ends.
[0113] In this first embodiment, the correct answer information generation unit 28 identifies the circumscribing bounding boxes that circumscribe multiple bounding boxes from the second correct answer image and generates the correct answer information represented by the identified circumscribing bounding boxes as the second correct answer information. However, this disclosure is not limited to this. The correct answer information generation unit 28 may also select one bounding box from among multiple bounding boxes based on the brightness or position of each of the multiple bounding boxes (frames) in the second correct answer image, and generate the correct answer information represented by the selected bounding box as the second correct answer information.
[0114] More specifically, the ground truth information generation unit 28 may calculate the average brightness of each of the multiple bounding boxes from the second ground truth image, select the bounding box with the highest calculated average brightness, and generate the ground truth information represented by the selected bounding box as the second ground truth information.
[0115] Further, the correct answer information generation unit 28 may specify an outer bounding box that circumscribes a plurality of bounding boxes from the second correct answer image, select one bounding box having a center point closest to the center point of the specified outer bounding box, and generate the correct answer information represented by the selected one bounding box as the second correct answer information. Note that the center point of the bounding box is the center of gravity of the rectangular bounding box and is the intersection of the two diagonal lines of the rectangular bounding box.
[0116] Further, the correct answer information generation unit 28 may specify an outer bounding box that circumscribes a plurality of bounding boxes from the second correct answer image, determine the center in the specified outer bounding box as a reference position on the image, and generate correct answer information having the same size as the bounding box with the determined reference position as the center as the second correct answer information.
[0117] In the above description, the correct answer image display control unit 26 displays the first correct answer image including the bounding box represented by the rectangular frame, but may also display the first correct answer image including the bounding box represented by the rectangular area. Hereinafter, a modification example of the first embodiment in which the first correct answer image including the rectangular bounding box area is displayed will be described.
[0118] FIG. 14 is a diagram showing an example of the first correct answer image displayed on the display device 3 in a modification example of the first embodiment. FIG. 15 is a diagram showing an example of the second correct answer image obtained by the imaging device 4 imaging the first correct answer image shown in FIG. 14 when the distance between the imaging device 4 and the display device 3 is L1, and FIG. 16 is a diagram showing an example of the second correct answer image obtained by the imaging device 4 imaging the first correct answer image shown in FIG. 14 when the distance between the imaging device 4 and the display device 3 is L2 (L1 < L2).
[0119] The first correct answer image 71 shown in Figure 14 includes a first bounding box region 341, which is an example of an object. The first bounding box region 341 is a rectangular area that includes the person 301 of the first training image 51 shown in Figure 4 as the detection target. The first bounding box region 341 is an area where the bounding box of the first correct answer information is filled with a predetermined brightness or predetermined color. The first bounding box region 341 is displayed on the first correct answer image 71, which is the same size as the first training image 51 displayed on the display device 3.
[0120] Furthermore, in the second correct image 72 shown in Figure 15, the first bounding box region 342 is the first bounding box region 341 captured by passing through the pinhole 2011 on the optical axis, and the first bounding box region 343 is the first bounding box region 341 captured by passing through the pinhole 2012 which is not on the optical axis.
[0121] Furthermore, in the second correct image 73 shown in Figure 16, the first bounding box region 344 is the first bounding box region 341 captured by passing through the pinhole 2011 on the optical axis, and the first bounding box region 345 is the first bounding box region 341 captured by passing through the pinhole 2012 which is not on the optical axis.
[0122] In a modified version of the first embodiment, the correct answer information generation unit 28 may identify the circumscribed frame of a second bounding box region in which a plurality of first bounding box regions are superimposed from the second correct answer image, and generate the correct answer information represented by the identified circumscribed frame as the second correct answer information.
[0123] Thus, the second ground truth images 72 and 73 obtained from the imaging device 4, which is a multi-pinhole camera, are images in which multiple images are superimposed. The position and size of the first bounding box regions 342 and 344, which are captured by passing through the pinhole 2011 on the optical axis, do not change on the second ground truth image. On the other hand, the position of the first bounding box regions 343 and 345, which are captured by passing through the pinhole 2012, which is not on the optical axis, changes on the second ground truth image depending on the distance between the subject and the imaging device 4. The greater the distance between the subject and the imaging device 4, the smaller the parallax. In other words, if the distance between the imaging device 4 and the display device 3 changes, bounding box regions with a changed amount of parallax are obtained.
[0124] The ground truth information generation unit 28 performs edge detection processing on the acquired second ground truth image 72 to detect the boundary of the second bounding box region 346, which is formed by the superposition of multiple first bounding box regions 342, 343, from the second ground truth image 72. The ground truth information generation unit 28 then identifies a rectangular circumscribing bounding box that circumscribes the second bounding box region 346 and generates the ground truth information represented by the identified circumscribing bounding box as the second ground truth information.
[0125] Furthermore, the correct answer information generation unit 28 performs edge detection processing on the acquired second correct answer image 73 to detect the boundary of the second bounding box region 347, which is formed by the superposition of multiple first bounding box regions 344, 345, from the second correct answer image 73. The correct answer information generation unit 28 then identifies a rectangular circumscribing bounding box that circumscribes the second bounding box region 347 and generates the correct answer information represented by the identified circumscribing bounding box as the second correct answer information.
[0126] The correct answer information generation unit 28 may detect the boundary of the second bounding box region 346,347 by performing, for example, a binarization process instead of an edge detection process. Alternatively, the correct answer information generation unit 28 may identify a rectangular inscribed bounding box inscribed within the second bounding box region 346,347 and generate the correct answer information represented by the identified inscribed bounding box as the second correct answer information.
[0127] Figure 17 shows an example of a rectangular bounding box circumscribing the second bounding box region 346 shown in Figure 15, and Figure 18 shows an example of a rectangular bounding box circumscribing the second bounding box region 347 shown in Figure 16.
[0128] The correct answer information generation unit 28 identifies the circumscribed bounding box 348 shown in Figure 17 by detecting the edges of the second bounding box region 346 shown in Figure 15. Furthermore, the correct answer information generation unit 28 identifies the circumscribed bounding box 349 shown in Figure 18 by detecting the edges of the second bounding box region 347 shown in Figure 16.
[0129] Figure 19 shows the second training image shown in Figure 5 with the bounding box 348 shown in Figure 17 superimposed on it, and Figure 20 shows the second training image shown in Figure 6 with the bounding box 349 shown in Figure 18 superimposed on it.
[0130] The bounding box 348 shown in Figure 19 is ground truth information that includes a person 303 captured by passing through a pinhole 2011 on the optical axis and a person 305 captured by passing through a pinhole 2012 that is not on the optical axis. The ground truth information generation unit 28 generates the ground truth information represented by the identified bounding box 348 as second ground truth information.
[0131] Furthermore, the bounding box 349 shown in Figure 20 is correct information that includes the person 307 captured by passing through the pinhole 2011 on the optical axis and the person 309 captured by passing through the pinhole 2012 which is not on the optical axis. The correct information generation unit 28 generates the correct information represented by the identified bounding box 349 as second correct information.
[0132] In a modified version of this first embodiment, the ground truth information generation unit 28 identifies an external bounding box that circumscribes a second bounding box region in which multiple first bounding box regions are superimposed, from the second ground truth image, and generates the ground truth information represented by the identified external bounding box as the second ground truth information. However, this disclosure is not limited to this. The ground truth information generation unit 28 may also determine a reference position on the image based on the brightness in the second bounding box region in which multiple first bounding box regions are superimposed, from the second ground truth image, and generate ground truth information of the same size as the first bounding box region centered on the determined reference position as the second ground truth information.
[0133] More specifically, the ground truth information generation unit 28 may determine the pixel with the maximum brightness value within the second bounding box region of the second ground truth image as the reference position, and generate ground truth information of the same size as the first bounding box region, centered on the determined reference position, as the second ground truth information.
[0134] Alternatively, the correct answer information generation unit 28 may identify a bounding box that circumscribes the second bounding box region, which is formed by the superposition of multiple first bounding box regions, from the second correct answer image, determine the center of the identified bounding box as a reference position on the image, and generate correct answer information of the same size as the first bounding box region as the second correct answer information, with the determined reference position as the center.
[0135] Furthermore, for example, if the pinholes are far apart, the second bounding box region may be imaged as multiple regions rather than as a single region. In this case, the ground truth information generation unit 28 may identify a region encompassing multiple bounding box regions (objects) from the second ground truth image and generate the ground truth information represented by the identified region as the second ground truth information.
[0136] In the above description, an example was given in which the identification task is object detection and the first ground truth information is a bounding box. However, the first ground truth information used by the imaging system 1 of this first embodiment is not limited to this. For example, the identification task may be pixel-by-pixel segmentation, such as semantic segmentation. In this case, the ground truth image display control unit 26 can sequentially display the first ground truth images based on the first ground truth information for each class. For example, if semantic segmentation is a task of classifying an outdoor image into three classes: roads, sky, and buildings, the ground truth image display control unit 26 may first display a first ground truth image in which pixels corresponding to roads are represented in white and pixels corresponding to parts other than roads are represented in black, and have the imaging device 4 capture it. Next, the ground truth image display control unit 26 may display a first ground truth image in which pixels corresponding to the sky are represented in white and pixels corresponding to parts other than the sky are represented in black, and have the imaging device 4 capture it. Finally, the ground truth image display control unit 26 may display a first ground truth image in which pixels corresponding to buildings are represented in white and pixels corresponding to parts other than buildings are represented in black, and have the imaging device 4 capture the image. In this way, a second ground truth image corresponding to each class can be obtained.
[0137] As described above, in this first embodiment, a first ground truth image based on first ground truth information corresponding to the training images for the machine learning model is displayed on the display device 3. Then, the first ground truth image displayed on the display device 3 is captured by the imaging device 4 that acquires blurred images, and a second ground truth image is acquired. Second ground truth information is generated based on the acquired second ground truth image. Then, a dataset including the training image and the second ground truth information is stored in the second storage unit 22.
[0138] Therefore, by generating accurate second-hand correct information corresponding to the degree of blur in the imaging device 4, and accumulating a dataset containing training images and pairs of second-hand correct information, it is possible to improve the recognition accuracy of the machine learning model while protecting the privacy of the subjects.
[0139] Furthermore, since the imaging system 1 of this first embodiment can acquire correct information that matches the parallax, it is possible to construct a more accurate dataset and improve the recognition accuracy of the machine learning model.
[0140] (Second Embodiment) In the first embodiment, the imaging control device 2 displays a first correct answer image on the display device 3 based on first correct answer information corresponding to a first training image, has the imaging device 4 capture the displayed first correct answer image to obtain a second correct answer image, and generates second correct answer information based on the acquired second correct answer image. In contrast, the imaging control device in the second embodiment displays a first geometric image on the display device 3, has the imaging device 4 capture the displayed first geometric image to obtain a second geometric image, generates a conversion table for converting the first geometric image to the second geometric image, and uses the generated conversion table to convert the first correct answer information to the second correct answer information.
[0141] Figure 21 is a block diagram showing an example of the overall configuration of imaging system 1A according to the second embodiment of this disclosure. In Figure 21, the same reference numerals are used for the same components as in Figure 1, and their descriptions are omitted.
[0142] The imaging system 1A comprises an imaging control device 2A, a display device 3, and an imaging device 4.
[0143] The imaging control device 2A includes a first storage unit 21, a second storage unit 22, an image acquisition unit 23, a training image display control unit 24, a correct answer information acquisition unit 25A, an imaging control unit 27A, a storage control unit 29A, a geometric image display control unit 30, a conversion table generation unit 31, a third storage unit 32, and a correct answer information conversion unit 33.
[0144] The geometric image display control unit 30 causes the display device 3 to display a first geometric image. The first geometric image includes a first dot placed at a predetermined position on the image.
[0145] The imaging control unit 27A causes the imaging device 4 to capture the first geometric image displayed on the display device 3, thereby acquiring a second geometric image. The second geometric image includes multiple second dots superimposed on the first dot as it is captured.
[0146] The conversion table generation unit 31 generates a conversion table for converting the positions of the first geometric image to the positions of the second geometric image. The conversion table generation unit 31 identifies the positions of multiple second dots from the second geometric image and generates a conversion table for converting the positions of the first dots to the positions of the identified multiple second dots.
[0147] The third storage unit 32 stores the conversion table generated by the conversion table generation unit 31.
[0148] The correct answer information acquisition unit 25A acquires first correct answer information corresponding to the first training image acquired by the image acquisition unit 23. The correct answer information acquisition unit 25A outputs the first correct answer information acquired from the first storage unit 21 to the correct answer information conversion unit 33.
[0149] The correct answer information conversion unit 33 uses the conversion table stored in the third storage unit 32 to convert the first correct answer information acquired by the correct answer information acquisition unit 25A into second correct answer information. The second correct answer information is the correct answer information corresponding to the second training image.
[0150] The memory control unit 29A stores a dataset in the second memory unit 22 that includes pairs of second training images obtained by imaging by the imaging device 4 and second correct answer information converted by the correct answer information conversion unit 33.
[0151] Next, the dataset creation process in the imaging control device 2A according to the second embodiment of this disclosure will be described.
[0152] Figure 22 is a flowchart illustrating the dataset creation process in the imaging control device 2A according to the second embodiment of this disclosure.
[0153] First, the geometric image display control unit 30 causes the first geometric image to be displayed on the display device 3 (step S201). This is because, if the display image on the display device 3 is represented by coordinates (u,v) (0≦u≦N,0≦v≦M), the geometric image display control unit 30 displays a first geometric image in which only the pixel at coordinate (u,v) is represented in white, and the other pixels are represented in black.
[0154] Figure 23 shows an example of a first geometric image displayed on the display device 3 in the second embodiment.
[0155] As shown in Figure 23, the geometric image display control unit 30 displays a first geometric image 81 in which only the first dot 401 corresponding to the coordinate (u1,v1), which is one point among multiple pixels, is represented in white. Pixels other than the first dot 401 are black. First, the geometric image display control unit 30 displays a first geometric image 81 in which only the first dot 401 corresponding to the top-left pixel is represented in white. Note that the first dot 401 is not composed of a single pixel, but may be composed of multiple pixels such as four pixels or nine pixels.
[0156] Returning to Figure 22, the next step is for the imaging control unit 27A to cause the imaging device 4 to capture the first geometric image displayed on the display device 3 and acquire a second geometric image (step S202). When the first geometric image is displayed on the display device 3, the imaging control unit 27A causes the imaging device 4 to capture the first geometric image and acquire a second geometric image. The imaging control unit 27A outputs the acquired second geometric image to the conversion table generation unit 31. The imaging control unit 27A takes images so that the display device 3 is in the field of view.
[0157] Next, the conversion table generation unit 31 generates a conversion table for converting the first geometric image into the second geometric image (step S203). The conversion table generation unit 31 stores the generated conversion table in the third storage unit 32.
[0158] Figure 24 shows an example of a second geometric image obtained by the imaging device 4 capturing the first geometric image shown in Figure 23.
[0159] In the second geometric image 82 shown in Figure 24, the second dot 411 is the first dot 401 captured by passing through the pinhole 2011 on the optical axis, and the second dot 412 is the first dot 401 captured by passing through the pinhole 2012 which is not on the optical axis. In other words, the first dot 401 on the first geometric image 81 is affected by the parallax from the imaging device 4 and is transformed into multiple second dots 411, 412 on the second geometric image 82.
[0160] The conversion table generation unit 31 identifies the pixel positions (x,y) of multiple second dots 411 and 412 from the second geometric image 82. The conversion table generation unit 31 then generates a conversion table that converts the pixel position (u,v) of the first dot 401 in the first geometric image 81 to the pixel positions (x,y) of multiple second dots 411 and 412 in the second geometric image 82. As shown in Figure 24, the first dot 401 at coordinates (u1,v1) displayed on the display device 3 is imaged as two bright points: the second dot 411 at coordinates (x1_1,y1_1) and the second dot 412 at coordinates (x1_2,y1_2).
[0161] As a result, the conversion table generation unit 31 can generate a conversion table for converting the coordinates (u,v) of the first correct answer information stored in the first storage unit 21 to the coordinates (x,y) of the second correct answer information captured by the imaging device 4.
[0162] In the second geometric image 82, the second dots 411 and 412 may each consist of a single dot or multiple dots. If the pinhole size is sufficiently large, the second dots 411 and 412 will be multiple dots with a spread. In this case, the conversion table generation unit 31 may identify the positions of all of these dots as the positions of the second dots 411 and 412. Alternatively, the conversion table generation unit 31 may identify the positions of the second dots 411 and 412, for example, one dot in each spread dot that corresponds to the local maximum brightness value with the highest brightness value.
[0163] Next, the conversion table generation unit 31 determines whether the conversion table is complete or not (step S204). The conversion table generation unit 31 determines that the conversion table is complete if all pixels constituting the display image are displayed and all pixels have been captured. The conversion table generation unit 31 determines that the conversion table is not complete if not all pixels constituting the display image are displayed and not all pixels have been captured.
[0164] If it is determined that the conversion table is not complete (NO in step S204), the process returns to step S201. The geometric image display control unit 30 then displays a first geometric image on the display device 3, in which the first dots corresponding to the pixels that are not being displayed are represented in white from among the multiple pixels that make up the display image.
[0165] Figure 25 shows another example of the first geometric image displayed on the display device 3.
[0166] For example, the geometric image display control unit 30 displays white dots on the display device 3 one pixel at a time, starting from the top left pixel of the displayed image. As shown in Figure 25, the geometric image display control unit 30 displays a first geometric image 81 including a first dot 401 corresponding to coordinates (u1, v1), and then displays a first geometric image 81A in which only the first dot 402 corresponding to coordinates (u2, v2) is shown in white, and all pixels other than the first dot 402 are shown in black. The position of the first dot 402 is different from the position of the first dot 401.
[0167] On the other hand, if it is determined that the conversion table is complete (YES in step S204), the image acquisition unit 23 acquires the first training image from the first storage unit 21 (step S205).
[0168] Note that the processes in steps S205 to S207 shown in Figure 22 are the same as the processes in steps S101 to S103 shown in Figure 3, so their explanation will be omitted.
[0169] Figure 26 shows an example of a conversion table in the second embodiment. The conversion table shown in Figure 26 associates the coordinates (u,v) of each pixel in the display image shown on the display device 3 with the coordinates (x,y) of the converted pixels in the captured image acquired by the imaging device 4.
[0170] For example, the first dot (pixel) at coordinates (u1,v1) in the displayed image is converted into two second dots (pixels) at coordinates (x1_1,y1_1) and (x1_2,y1_2) in the captured image.
[0171] Returning to Figure 22, the next step is for the correct answer information acquisition unit 25A to acquire the first correct answer information corresponding to the displayed first training image from the first storage unit 21 (step S208).
[0172] Next, the correct answer information conversion unit 33 uses the conversion table generated by the conversion table generation unit 31 to convert the first correct answer information acquired by the correct answer information acquisition unit 25A into second correct answer information corresponding to the second training image captured by the imaging device 4 (step S209). The correct answer information conversion unit 33 uses the conversion table shown in Figure 26 to convert the coordinates of the four vertices of the bounding box, which is the first correct answer information corresponding to the first training image acquired from the first storage unit 21, into the coordinates of multiple points. Then, the correct answer information conversion unit 33 identifies a rectangular bounding box that circumscribes the multiple conversion points. The correct answer information conversion unit 33 generates the correct answer information represented by the identified bounding box as second correct answer information corresponding to the second training image captured by the imaging device 4.
[0173] Figure 27 is a schematic diagram illustrating the generation process of the second correct answer information in this second embodiment. Here, the coordinates of the four vertices of the bounding box, which is the first correct answer information corresponding to the first training image acquired from the first storage unit 21, are (u1,v1), (u2,v2), (u3,v3), and (u4,v4). From the conversion table shown in Figure 26, the vertex at coordinate (u1,v1) is converted to points 411 and 412 at coordinates (x1_1,y1_1) and (x1_2,y1_2) on the second training image. Also, the vertex at coordinate (u2,v2) is converted to points 413 and 414 at coordinates (x2_1,y2_1) and (x2_2,y2_2) on the second training image. Furthermore, the vertex at coordinate (u3,v3) is transformed into points 415 and 416 at coordinates (x3_1,y3_1) and (x3_2,y3_2) on the second training image. Similarly, the vertex at coordinate (u4,v4) is transformed into points 417 and 418 at coordinates (x4_1,y4_1) and (x4_2,y4_2) on the second training image. At this time, an external bounding box 420 that circumscribes points 411 to 418 is generated as second correct answer information corresponding to the second training image captured by the imaging device 4.
[0174] Returning to Figure 22, the memory control unit 29A then stores a dataset in the second memory unit 22 that includes a pair of the second training image obtained by imaging by the imaging device 4 and the second correct answer information converted by the correct answer information conversion unit 33 (step S210).
[0175] Next, the image acquisition unit 23 determines whether the imaging device 4 has captured all of the first training images stored in the first storage unit 21 (step S211). If it is determined that the imaging device 4 has not captured all of the first training images (NO in step S211), the process returns to step S205, and the image acquisition unit 23 acquires the uncaptured first training images from among the multiple first training images stored in the first storage unit 21. Then, the training image display control unit 24 displays the first training images on the display device 3, and the imaging control unit 27A causes the imaging device 4 to capture the displayed first training images. The correct answer information acquisition unit 25A acquires the first correct answer information corresponding to the first training images. The correct answer information conversion unit 33 uses the conversion table generated by the conversion table generation unit 31 to convert the first correct answer information acquired by the correct answer information acquisition unit 25A into second correct answer information corresponding to the second training images captured by the imaging device 4.
[0176] On the other hand, if it is determined that the imaging device 4 has captured all of the first training images (YES in step S211), the process ends.
[0177] As described above, the imaging system 1A of this second embodiment can acquire correct information that matches the parallax, making it possible to construct a more accurate dataset and improve the recognition accuracy of the machine learning model.
[0178] In the above description, the geometric image display control unit 30 displays one first dot while shifting it by one pixel at a time, but it may also display multiple first dots while shifting them by one pixel at a time. Using Figures 28 and 29, the process of displaying multiple first dots simultaneously on the display device 3 in the first modified example of the second embodiment will be described.
[0179] Figure 28 shows an example of a first geometric image displayed on the display device 3 in a first modified example of the second embodiment.
[0180] As shown in Figure 28, the geometric image display control unit 30 displays a first geometric image 81B in which only the first dots 403-406 corresponding to four points among the multiple pixels, namely coordinates (u11,v11), (u12,v12), (u13,v13), and (u14,v14), are represented in white. Pixels other than the first dots 403-406 are black. The geometric image display control unit 30 displays a first geometric image 81B in which only the first dots 403-406 corresponding to the four pixels are represented in white.
[0181] Figure 29 shows an example of a second geometric image obtained by the imaging device 4 capturing the first geometric image shown in Figure 28.
[0182] As shown in Figure 29, the four first dots 403 to 406 displayed on the display device 3 are imaged as multiple bright spots having local maximum brightness values. Based on the acquired second geometric image 82B, the conversion table generation unit 31 generates a conversion table that converts the pixel positions (u,v) of the first dots 403 to 406 in the first geometric image 81B to the pixel positions (x,y) of multiple second dots in the second geometric image 82B. The conversion table generation unit 31 divides the second geometric image 82B into multiple regions to correspond to the bright spots displayed on the display device 3, and calculates the local maximum value of the bright spots for each region. The conversion table generation unit 31 may calculate the coordinates of the bright spot with the local maximum value as the pixel positions (x,y) of multiple second dots in the second geometric image 82B that correspond to the pixel positions (u,v) of the first dots 403 to 406 in the first geometric image 81B.
[0183] Specifically, the conversion table generation unit 31 divides the second geometric image 82B into multiple regions 431 to 434, corresponding to the four bright spots, or first dots 403 to 406, displayed by the display device 3. At this time, the conversion table generation unit 31 divides the second geometric image 82B into a number of regions corresponding to the number of first dots displayed by the display device 3, using a clustering method such as the k-means method. The conversion table generation unit 31 then identifies the coordinates corresponding to the local maximum brightness in region 431 as the pixel position of the second dot in the second geometric image 82B, corresponding to the first dot 403 in the first geometric image 81B. The conversion table generation unit 31 also identifies the coordinates corresponding to the local maximum brightness in region 432 as the pixel position of the second dot in the second geometric image 82B, corresponding to the first dot 404 in the first geometric image 81B. Furthermore, the conversion table generation unit 31 identifies the coordinates corresponding to the local maximum value of brightness in region 433 as the pixel position of the second dot of the second geometric image 82B corresponding to the first dot 405 of the first geometric image 81B. Also, the conversion table generation unit 31 identifies the coordinates corresponding to the local maximum value of brightness in region 434 as the pixel position of the second dot of the second geometric image 82B corresponding to the first dot 406 of the first geometric image 81B.
[0184] Furthermore, the geometric image display control unit 30 may display multiple first dots of different colors, rather than displaying multiple first dots of only white. This type of processing is effective for simultaneously displaying multiple first dots on the display device 3. In other words, by making the colors of the multiple first dots displayed on the display device 3 different, it is possible to identify which first dot corresponds to each of the multiple second dots on the second geometric image 82B captured by the imaging device 4.
[0185] Furthermore, in this second embodiment, a first geometric image including a first dot is displayed, the positions of a plurality of second dots are identified from a second geometric image obtained by imaging the first geometric image, and a conversion table is created for converting the position of the first dot to the positions of the plurality of second dots, but this disclosure is not particularly limited thereto. In a second modification of this second embodiment, the first geometric image may include a first horizontal line and a first vertical line arranged at predetermined positions on the image. The conversion table generation unit 31 may then identify the positions of a plurality of second horizontal lines and a plurality of second vertical lines from the second geometric image, convert the position of the first horizontal line to the positions of the identified plurality of second horizontal lines, and generate a conversion table for converting the position of the first vertical line to the positions of the identified plurality of second vertical lines.
[0186] Figure 30 shows an example of a first geometric image including a first horizontal line displayed on the display device 3 in a second modified example of the second embodiment.
[0187] As shown in Figure 30, the geometric image display control unit 30 displays a first geometric image 91A in which only the first horizontal line 501, which corresponds to the horizontal pixel row among the multiple pixels constituting the display image, is shown in white. Pixels other than the first horizontal line 501 are shown in black. First, the geometric image display control unit 30 displays a first geometric image 91A in which only the first horizontal line 501, which corresponds to the uppermost pixel row, is shown in white. Note that the width of the first horizontal line 501 may be multiple pixels, not just one.
[0188] Figure 31 shows an example of a second geometric image obtained by the imaging device 4 capturing the first geometric image shown in Figure 30.
[0189] In the second geometric image 92A shown in Figure 31, the second horizontal line 511 is the first horizontal line 501 captured by passing through the pinhole 2011 on the optical axis, and the second horizontal line 512 is the first horizontal line 501 captured by passing through the pinhole 2012 which is not on the optical axis. In other words, the first horizontal line 501 on the first geometric image 91A is affected by parallax from the imaging device 4 and is transformed into multiple second horizontal lines 511, 512 on the second geometric image 92A.
[0190] The conversion table generation unit 31 identifies the pixel rows (positions) of multiple second horizontal lines 511 and 512 from the second geometric image 92A. The pixel rows of the second horizontal lines 511 and 512 are represented by the coordinates of the leftmost pixel in the horizontal direction and the coordinates of the rightmost pixel in the horizontal direction. The conversion table generation unit 31 then generates a conversion table that converts the pixel row of the first horizontal line 501 in the first geometric image 91A into the pixel rows of multiple second horizontal lines 511 and 512 in the second geometric image 92A. As shown in Figure 31, the first horizontal line 501 displayed on the display device 3 is imaged as two straight lines of the second horizontal lines 511 and 512.
[0191] The geometric image display control unit 30 displays the uppermost first horizontal line 501, and then sequentially displays the first horizontal lines by shifting them downwards one pixel at a time. The imaging control unit 27A acquires the second geometric image 92A by having the imaging device 4 capture the first geometric image 91A each time the first horizontal line is displayed. The conversion table generation unit 31 identifies the positions of multiple second horizontal lines from the second geometric image 92A each time the second geometric image 92A is acquired. The conversion table generation unit 31 generates a conversion table for converting the positions of the first horizontal lines to the positions of the identified multiple second horizontal lines.
[0192] Figure 32 shows an example of a first geometric image including a first vertical line displayed on the display device 3 in a second modified example of the second embodiment.
[0193] As shown in Figure 32, the geometric image display control unit 30 displays a first geometric image 91B in which only the first vertical line 521 corresponding to the vertical pixel row among the multiple pixels is represented in white. Pixels other than the first vertical line 521 are black. First, the geometric image display control unit 30 displays a first geometric image 91B in which only the first vertical line 521 corresponding to the leftmost pixel row is represented in white. Note that the width of the first vertical line 521 may be multiple pixels, not just one.
[0194] Figure 33 shows an example of a second geometric image obtained by the imaging device 4 capturing the first geometric image shown in Figure 32.
[0195] In the second geometric image 92B shown in Figure 33, the second vertical line 531 is the first vertical line 521 captured by passing through the pinhole 2011 on the optical axis, and the second vertical line 532 is the first vertical line 521 captured by passing through the pinhole 2012 which is not on the optical axis. In other words, the first vertical line 521 on the first geometric image 91B is affected by parallax from the imaging device 4 and is transformed into multiple second vertical lines 531, 532 on the second geometric image 92B.
[0196] The conversion table generation unit 31 identifies the pixel rows (positions) of multiple second vertical lines 531 and 532 from the second geometric image 92B. The pixel rows of the second vertical lines 531 and 532 are represented by the coordinates of the pixel at the upper end in the vertical direction and the coordinates of the pixel at the lower end in the vertical direction. The conversion table generation unit 31 then generates a conversion table that converts the pixel row of the first vertical line 521 in the first geometric image 91B into the pixel rows of multiple second vertical lines 531 and 532 in the second geometric image 92B. As shown in Figure 33, the first vertical line 521 displayed on the display device 3 is imaged as two straight lines of the second vertical lines 531 and 532.
[0197] The geometric image display control unit 30 displays the first vertical line 521 at the left edge, and then sequentially displays the first vertical lines by shifting them to the right one pixel at a time. The imaging control unit 27A acquires the second geometric image 92B by having the imaging device 4 capture the first geometric image 91B each time the first vertical line is displayed. The conversion table generation unit 31 identifies the positions of multiple second vertical lines from the second geometric image 92A each time the second geometric image 92B is acquired. The conversion table generation unit 31 generates a conversion table for converting the positions of the first vertical lines to the positions of the identified multiple second vertical lines.
[0198] Of course, the pixel rows of the second horizontal line or second vertical line may be represented not by the coordinates of the leftmost pixel in the horizontal direction and the coordinates of the rightmost pixel in the horizontal direction, or by the coordinates of the topmost pixel in the vertical direction and the coordinates of the bottommost pixel in the vertical direction, but by the coordinates of all pixels on the second horizontal line or second vertical line.
[0199] The above describes the conversion table generation process in the second modified example of this second embodiment.
[0200] Next, the process for generating the second correct answer information in the second modified example of this second embodiment will be described.
[0201] Figure 34 shows an example of the first correct answer information in a second modified example of the second embodiment, and Figure 35 shows an example of the second correct answer information in a second modified example of the second embodiment.
[0202] As shown in Figure 34, the rectangular bounding box 551 is the first correct answer information corresponding to the first training image. Each side of the bounding box 551 lies on two first horizontal lines 601, 602 and two first vertical lines 603, 604. As shown in Figure 35, the first horizontal line 601 is transformed into two second horizontal lines 611, 612, the first horizontal line 602 is transformed into two second horizontal lines 613, 614, the first vertical line 603 is transformed into two second vertical lines 615, 616, and the first vertical line 604 is transformed into two second vertical lines 617, 618.
[0203] In other words, the correct answer information conversion unit 33 uses a conversion table to convert the four lines of the bounding box 551, which is the first correct answer information corresponding to the first training image acquired from the first storage unit 21, into multiple lines. The correct answer information conversion unit 33 then identifies a rectangular circumscribing bounding box 563 that circumscribes the bounding box 561 enclosed by the second horizontal lines 611, 613 and the second vertical lines 615, 617, and the bounding box 562 enclosed by the second horizontal lines 612, 614 and the second vertical lines 616, 618. The correct answer information conversion unit 33 generates the correct answer information represented by the identified circumscribing bounding box 563 as the second correct answer information corresponding to the second training image captured by the imaging device 4.
[0204] In the second modification of this second embodiment, the geometric image display control unit 30 displays one first horizontal line while shifting it by one pixel row at a time, and displays one first vertical line while shifting it by one pixel row at a time, but this disclosure is not limited thereto. The geometric image display control unit 30 may display multiple first horizontal lines while shifting them by one pixel row at a time. The imaging control unit 27A may acquire a second geometric image by having the imaging device 4 capture a first geometric image each time multiple first horizontal lines are displayed. The conversion table generation unit 31 may identify the positions of multiple second horizontal lines from the second geometric image each time a second geometric image is acquired. The conversion table generation unit 31 may generate a conversion table for converting the positions of multiple first horizontal lines to the positions of the identified multiple second horizontal lines.
[0205] Furthermore, the geometric image display control unit 30 may display multiple first vertical lines while shifting them by one pixel row at a time. The imaging control unit 27A may acquire a second geometric image by having the imaging device 4 capture the first geometric image each time multiple first vertical lines are displayed. The conversion table generation unit 31 may identify the positions of multiple second vertical lines from the second geometric image each time a second geometric image is acquired. The conversion table generation unit 31 may generate a conversion table for converting the positions of multiple first vertical lines to the positions of the identified multiple second vertical lines.
[0206] Furthermore, the geometric image display control unit 30 may simultaneously display at least one first horizontal line and at least one first vertical line while shifting them by one pixel row at a time. The imaging control unit 27A may acquire a second geometric image by having the imaging device 4 capture the first geometric image each time at least one first horizontal line and at least one first vertical line are displayed. Each time the second geometric image is acquired, the conversion table generation unit 31 generates a plurality of second horizontal lines and a plurality of second vertical lines from the second geometric image. vertical The position of the lines may be identified. The conversion table generation unit 31 identifies the positions of at least one first horizontal line and at least one first vertical line, and the positions of the identified plurality of second horizontal lines and plurality of second lines. vertical A conversion table may be generated to convert to line positions.
[0207] Furthermore, in this second embodiment, the geometric image display control unit 30 may display a first dot based on the first correct answer information. For example, the geometric image display control unit 30 may display at least one first dot while shifting it by one pixel at a time along the outline of the bounding box, which is the first correct answer information. The imaging control unit 27A may acquire a second geometric image by causing the imaging device 4 to capture the first geometric image each time at least one first dot is displayed. The conversion table generation unit 31 may identify the positions of a plurality of second dots from the second geometric image each time the second geometric image is acquired. The conversion table generation unit 31 may generate a conversion table for converting the position of at least one first dot to the positions of the identified plurality of second dots.
[0208] In this case, the geometric image display control unit 30 may refer to the conversion table before displaying at least one first dot included in the outline of the bounding box, and determine whether the position of at least one first dot and the positions of multiple second dots are already associated in the conversion table. If the position of at least one first dot and the positions of multiple second dots are already associated in the conversion table, the geometric image display control unit 30 does not need to display the at least one first dot. On the other hand, if the position of at least one first dot and the positions of multiple second dots are not associated in the conversion table, the geometric image display control unit 30 may display the at least one first dot.
[0209] Furthermore, in this second embodiment, the geometric image display control unit 30 may display a first horizontal line or a first vertical line based on the first correct answer information. For example, the geometric image display control unit 30 may display a first horizontal line corresponding to the upper or lower edge of the bounding box outline, which is the first correct answer information, or it may display a first vertical line corresponding to the left or right edge of the bounding box outline. The imaging control unit 27A may acquire a second geometric image by causing the imaging device 4 to capture the first geometric image each time a first horizontal line or a first vertical line is displayed. The conversion table generation unit 31 may identify the positions of a plurality of second horizontal lines or a plurality of second vertical lines from the second geometric image each time a second geometric image is acquired. The conversion table generation unit 31 may generate a conversion table for converting the positions of the first horizontal line or first vertical line to the positions of the identified plurality of second horizontal lines or a plurality of second vertical lines.
[0210] In this case, the geometric image display control unit 30 may refer to the conversion table before displaying the first horizontal line or first vertical line included in the outline of the bounding box, and determine whether the position of the first horizontal line or first vertical line is already associated with the positions of multiple second horizontal lines or multiple second vertical lines in the conversion table. If the position of the first horizontal line or first vertical line is already associated with the positions of multiple second horizontal lines or multiple second vertical lines in the conversion table, the geometric image display control unit 30 does not need to display the first horizontal line or first vertical line. On the other hand, if the position of the first horizontal line or first vertical line is not associated with the positions of multiple second horizontal lines or multiple second vertical lines in the conversion table, the geometric image display control unit 30 may display the first horizontal line or first vertical line.
[0211] (Third embodiment) In the first and second embodiments, a dataset containing pairs of second training images and second ground truth information is stored in the second memory unit, while in the third embodiment, a machine learning model is trained using the dataset containing pairs of second training images and second ground truth information stored in the second memory unit.
[0212] Figure 36 is a block diagram showing an example of the overall configuration of the imaging system 1B according to the third embodiment of this disclosure. In Figure 36, the same reference numerals are used for the same components as in Figure 1, and their descriptions are omitted.
[0213] The imaging system 1B comprises an imaging control device 2B, a display device 3, and an imaging device 4.
[0214] The imaging control device 2B includes a first storage unit 21, a second storage unit 22, an image acquisition unit 23, a training image display control unit 24, a correct answer information acquisition unit 25, a correct answer image display control unit 26, an imaging control unit 27, a correct answer information generation unit 28, a storage control unit 29, a training unit 34, and a model storage unit 35.
[0215] The training unit 34 trains a machine learning model using a dataset containing pairs of second training images and second correct answer information stored in the second memory unit 22. In this third embodiment, the machine learning model applied to the classifier is a machine learning model using a neural network such as deep learning, but other machine learning models may also be used. For example, the machine learning model may be a machine learning model using random forest or genetic programming.
[0216] Machine learning in the training unit 34 is implemented, for example, by backpropagation (BP) in deep learning. Specifically, the training unit 34 inputs a second training image to the machine learning model and obtains the recognition result output by the machine learning model. The training unit 34 then adjusts the machine learning model so that the recognition result becomes the second ground truth information. The training unit 34 improves the recognition accuracy of the machine learning model by repeating the adjustment of the machine learning model for multiple sets (e.g., thousands of sets) of different second training images and second ground truth information.
[0217] The model memory unit 35 stores a trained machine learning model. This machine learning model is also an image recognition model used for image recognition.
[0218] In this third embodiment, the imaging control device 2B includes a training unit 34 and a model storage unit 35. However, the disclosure is not limited thereto, and an external computer connected to the imaging control device 2B via a network may also include a training unit 34 and a model storage unit 35. In this case, the imaging control device 2B may further include a communication unit for transmitting a dataset to the external computer. Alternatively, an external computer connected to the imaging control device 2B via a network may also include a model storage unit 35. In this case, the imaging control device 2B may further include a communication unit for transmitting a trained machine learning model to the external computer.
[0219] In this third embodiment, the imaging system 1B can use the depth information of the subject, which is included in the disparity information, as training data, and is therefore effective in improving the recognition ability of the machine learning model. For example, the machine learning model can recognize that a small object in the image is a subject located at a distance, and can prevent it from being mistaken for dust or ignored. As a result, the machine learning model constructed by machine learning using the second training image can improve its recognition performance.
[0220] In this third embodiment, the imaging system 1B displays a first training image stored in the first storage unit 21, and acquires a second training image from the imaging device 4 when the first training image is captured. The imaging system 1B also displays a first correct answer image based on first correct answer information stored in the first storage unit 21, and acquires a second correct answer image from the imaging device 4 when the first correct answer image is captured. The imaging system 1B then generates second correct answer information based on the second correct answer image, stores a dataset containing the pair of the second training image and the second correct answer information in the second storage unit 22, and uses the stored dataset for training.
[0221] As described above, the imaging system 1B is effective not only for training and optimizing machine learning parameters, but also for optimizing the device parameters of the imaging device 4. When a multi-pinhole camera is used as the imaging device 4, the recognition performance and privacy protection performance of the imaging device 4 depend on device parameters such as the size of each pinhole, the shape of each pinhole, the arrangement of each pinhole, and the number of pinholes. Therefore, in order to realize an optimal recognition system, it is necessary to optimize not only the machine learning parameters, but also the device parameters of the imaging device 4, such as the size of each pinhole, the shape of each pinhole, the arrangement of each pinhole, and the number of pinholes. In this third embodiment, the imaging system 1B can select device parameters that have a high recognition rate and high privacy protection performance as the optimal device parameters by training and evaluating the second training images obtained when the device parameters of the imaging device 4 are changed.
[0222] In each of the above embodiments, each component may be implemented by dedicated hardware or by executing a software program suitable for each component. Each component may also be implemented by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory. Furthermore, the program may be executed by another independent computer system by recording and transferring the program to a recording medium, or by transferring the program via a network.
[0223] Some or all of the functions of the apparatus according to the embodiments of this disclosure are typically implemented as an integrated circuit, or LSI (Large Scale Integration). These may be individually integrated onto a single chip, or some or all of them may be integrated onto a single chip. Furthermore, the integration is not limited to LSIs, but may also be implemented using dedicated circuits or general-purpose processors. An FPGA (Field Programmable Gate Array) that can be programmed after LSI manufacturing, or a reconfigurable processor that can reconfigure the connections and settings of circuit cells inside the LSI may also be used.
[0224] Furthermore, some or all of the functions of the apparatus according to the embodiments of this disclosure may be realized by a processor such as a CPU executing a program.
[0225] Furthermore, all figures used above are illustrative examples provided to illustrate this disclosure, and this disclosure is not limited to these illustrative figures.
[0226] Furthermore, the order in which the steps shown in the flowchart above are performed is illustrative for the purpose of specifically illustrating this disclosure, and other orders are acceptable as long as similar effects are achieved. Also, some of the steps above may be performed simultaneously (in parallel) with other steps. [Industrial applicability]
[0227] The technology disclosed herein is useful as a technique for creating datasets used to train machine learning models, as it can improve the recognition accuracy of machine learning models while protecting the privacy of the subjects.
Claims
1. An image acquisition unit that acquires a first training image without blurring for a machine learning model from a first storage unit, A training image display control unit that displays the first training image on a display device, A first imaging control unit obtains a second training image by having an imaging device that acquires a blurred image whose parallax changes according to the distance to the subject capture the first training image displayed on the display device, A correct answer information acquisition unit acquires first correct answer information corresponding to the first training image from the first storage unit, A correct answer image display control unit that displays a first correct answer image based on the first correct answer information on the display device, A second imaging control unit which causes the imaging device to capture the first correct answer image displayed on the display device in order to acquire a second correct answer image, A generation unit that generates second correct information based on the second correct image, which is adjusted to match the parallax of the second training image, A storage control unit stores a dataset containing the pair of the second training image and the second correct answer information in a second storage unit, An information processing system equipped with the following features.
2. The first correct image includes an object that indicates the position of the first correct information on the image, The information processing system according to claim 1.
3. The aforementioned object is a frame, The imaging device acquires a superimposed image in which the single subject is shifted and overlapped by imaging a single subject. The second imaging control unit obtains a second ground truth image by superimposing multiple frames, by causing the imaging device to capture the first ground truth image including the frame. The generation unit selects one frame from the plurality of frames based on the brightness or position of each of the plurality of frames in the second correct answer image, and generates the correct answer information represented by the selected frame as the second correct answer information. The information processing system according to claim 2.
4. The aforementioned object is a frame, The imaging device acquires a superimposed image in which the single subject is shifted and overlapped by imaging a single subject. The second imaging control unit obtains a second ground truth image by superimposing multiple frames, by causing the imaging device to capture the first ground truth image including the frame. The generation unit identifies the circumscribed frames that circumscribe the plurality of frames from the second correct answer image, and generates the correct answer information represented by the identified circumscribed frames as the second correct answer information. The information processing system according to claim 2.
5. The aforementioned object is a frame, The imaging device acquires a superimposed image in which the single subject is shifted and overlapped by imaging a single subject. The second imaging control unit obtains a second ground truth image by superimposing multiple frames, by causing the imaging device to capture the first ground truth image including the frame. The generation unit identifies a circumscribed frame that circumscribes the plurality of frames from the second ground truth image, determines the center of the identified circumscribed frame as a reference position on the image, and generates ground truth information as the second ground truth information, which is represented by a frame of the same size as the frame, with the determined reference position as the center. The information processing system according to claim 2.
6. The aforementioned object is a first planar figure, The imaging device acquires a superimposed image in which the single subject is shifted and overlapped by imaging a single subject. The second imaging control unit obtains a second ground truth image by superimposing a plurality of first ground truth images by causing the imaging device to capture the first ground truth image including the first planar figure, The generation unit determines a reference position on the image based on the brightness of the second planar figure in which the plurality of first planar figures are superimposed from the second correct answer image, and generates correct answer information as the second correct answer information, which is represented by a planar figure of the same size as the first planar figure, centered on the determined reference position. The information processing system according to claim 2.
7. The aforementioned object is a first planar figure, The imaging device acquires a superimposed image in which the single subject is shifted and overlapped by imaging a single subject. The second imaging control unit obtains a second ground truth image by superimposing a plurality of first ground truth images by causing the imaging device to capture the first ground truth image including the first planar figure, The generation unit identifies a circumscribed frame that circumscribes the second planar figure in which the plurality of first planar figures are superimposed, from the second correct answer image, and generates the correct answer information represented by the identified circumscribed frame as the second correct answer information. The information processing system according to claim 2.
8. The aforementioned object is a first planar figure, The imaging device acquires a superimposed image in which the single subject is shifted and overlapped by imaging a single subject. The second imaging control unit obtains a second ground truth image by superimposing a plurality of first ground truth images by causing the imaging device to capture the first ground truth image including the first planar figure, The generation unit identifies a circumscribing frame that circumscribes the second planar figure in which the plurality of first planar figures are superimposed, determines the center of the identified circumscribing frame as a reference position on the image, and generates the second correct information as correct information represented by a planar figure of the same size as the first planar figure, with the determined reference position as the center. The information processing system according to claim 2.
9. The generation unit identifies a region containing multiple objects from the second ground truth image and generates the ground truth information represented by the identified region as the second ground truth information. The information processing system according to claim 2.
10. The system further comprises a training unit that trains the machine learning model using a dataset containing pairs of the second training images and the second correct answer information stored in the second memory unit. An information processing system according to any one of claims 1 to 9.
11. Computers The first training image, free from blur, is obtained from the first memory unit for the machine learning model. The first training image is displayed on the display device, The first training image displayed on the display device is captured by an imaging device that acquires a blurred image whose parallax changes according to the distance to the subject, thereby acquiring a second training image. The first correct answer information corresponding to the first training image is obtained from the first storage unit, The first correct answer image based on the first correct answer information is displayed on the display device. The first correct image displayed on the display device is captured by the imaging device to obtain a second correct image. Based on the second correct image, second correct information is generated that matches the parallax of the second training image. A dataset containing the pair of the second training image and the second correct answer information is stored in the second memory unit. Information processing methods.
Citation Information
Patent Citations
Mark recognizing device and mark recognizing method
JP1996122267A
Learning device, method for learning, and program
JP2019200769A
Identification system, identification device, method for identification, and program
JP2019200772A