Information processing system and information processing method

By adjusting display sizes and capturing multiple blurred images with varying blur levels, the system addresses inconsistent blur in training datasets, improving recognition accuracy and protecting privacy in machine learning models.

JP7842741B2Active Publication Date: 2026-04-08PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-04-19
Publication Date
2026-04-08

AI Technical Summary

Technical Problem

Existing methods for creating training datasets for machine learning models using lensless cameras result in inconsistent blur levels due to varying distances between the display and camera, affecting recognition accuracy while compromising subject privacy.

Method used

An information processing system that adjusts the display size of training images based on the distance between the display and imaging device, capturing multiple blurred images with varying degrees of blur to create a diverse dataset, ensuring accurate recognition while protecting privacy.

Benefits of technology

Improves the recognition accuracy of machine learning models by incorporating training images with varying degrees of blur, thereby enhancing privacy protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007842741000001
    Figure 0007842741000001
  • Figure 0007842741000002
    Figure 0007842741000002
  • Figure 0007842741000003
    Figure 0007842741000003
Patent Text Reader

Abstract

An imaging control device comprising: an image acquiring unit for acquiring a first training image for a machine learning model from a first storage unit; a distance acquiring unit for acquiring a distance between a display device and an imaging device for acquiring a blurred image; a display control unit for causing the first training image to be displayed on the display device on the basis of the distance; an imaging control unit for causing the first training image to be imaged by the imaging device to acquire a second training image; and a storage control unit for storing a data set including a set of the second training image and correct information in a second storage unit. The display control unit, when the distance is changed, changes the display size of the first training image so that the size of the second training image obtained by the imaging by the imaging device can be maintained. The imaging control unit, when the distance is changed, causes the first training image to be imaged by the imaging device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a technique for creating a dataset used for training a machine learning model.

Background Art

[0002] For example, Non-Patent Document 1 describes a method for creating a dataset for training a face detection model using a lensless camera by displaying an image captured by a normal camera on a display and capturing the image displayed on the display with the lensless camera.

[0003] However, when a dataset is created by capturing an image displayed on a display as in the prior art, a training dataset including training images with various degrees of blur according to the distance between the display and the camera is not created. Therefore, it has been difficult to improve the recognition accuracy of a machine learning model while protecting the privacy of a subject.

Prior Art Documents

Non-Patent Documents

[0004]

Non-Patent Document 1

Summary of the Invention

[0005] The present disclosure has been made to solve the above problems, and an object thereof is to provide a technique capable of improving the recognition accuracy of a machine learning model while protecting the privacy of a subject.

[0006] The information processing system according to this disclosure includes: an image acquisition unit that acquires a first training image for a machine learning model from a first storage unit; a distance acquisition unit that acquires the distance between a display device and an imaging device that acquires a blurred image by imaging; a display control unit that displays the first training image on the display device based on the distance; an imaging control unit that causes the imaging device to image the first training image displayed on the display device and acquire a second training image; and a storage control unit that stores a dataset in a second storage unit that includes the second training image obtained by imaging by the imaging device and correct answer information corresponding to the first training image. The display control unit changes the display size of the first training image when the distance is changed so that the size of the second training image obtained by imaging by the imaging device is maintained, and displays the first training image with the changed display size on the display device. The imaging control unit causes the imaging device to image the first training image when the distance is changed.

[0007] According to this disclosure, it is possible to improve the recognition accuracy of machine learning models while protecting the privacy of the subjects. [Brief explanation of the drawing]

[0008] [Figure 1] This block diagram shows an example of the overall configuration of the imaging system according to the first embodiment of this disclosure. [Figure 2] This diagram schematically shows the structure of a multi-pinhole camera, which is an example of an imaging device. [Figure 3] This is a flowchart illustrating the dataset creation process in the imaging control device according to the first embodiment of this disclosure. [Figure 4] This is a schematic diagram illustrating the positional relationship between the display device and the imaging device in the first embodiment. [Figure 5] This is a schematic diagram illustrating the images captured by passing through two pinholes in a multi-pinhole camera. [Figure 6]This is a diagram showing an example of an image captured by an imaging device when the distance between the imaging device and the display device is L1. [Figure 7] This is a diagram showing an example of an image captured by an imaging device when the distance between the imaging device and the display device is L2 (L1 < L2). [Figure 8] This is a diagram showing an example of a first training image displayed on a display device. [Figure 9] This is a diagram showing an example of a second training image obtained by the imaging device capturing the first training image shown in FIG. 8 when the distance between the imaging device and the display device is L1. [Figure 10] This is a diagram showing an example of a second training image obtained by the imaging device capturing the first training image shown in FIG. 8 when the distance between the imaging device and the display device is L2 (L1 < L2). [Figure 11] This is a flowchart for explaining the dataset creation process in an imaging control device according to a modification of the first embodiment of the present disclosure. [Figure 12] This is a block diagram showing an example of the overall configuration of an imaging system according to the second embodiment of the present disclosure. [Figure 13] This is a flowchart for explaining the dataset creation process in an imaging control device according to the second embodiment of the present disclosure. [Figure 14] This is a block diagram showing an example of the overall configuration of an imaging system according to the third embodiment of the present disclosure.

Embodiments for Carrying Out the Invention

[0009] (Findings on which the present disclosure is based) In homes and indoors, various recognition technologies are important, such as recognizing the actions of people in the environment or recognizing the person operating a device. In recent years, a technology called deep learning has attracted attention for object recognition. Deep learning is a machine learning method that uses a multi-layered neural network, and by utilizing a large amount of training data, it is possible to achieve higher accuracy in recognition performance compared to conventional methods. Image information is particularly effective in such object recognition. Various methods have been proposed that significantly improve conventional object recognition capabilities by using a camera as an input device and performing deep learning with image information as input.

[0010] However, placing cameras inside homes and other public spaces presents the challenge of privacy violations if captured images are leaked externally due to hacking or other means. Therefore, measures are needed to protect the privacy of the subjects even if captured images are leaked externally.

[0011] For example, a multi-pinhole camera is used to obtain blurred images that are difficult for humans to visually perceive. Images captured by a multi-pinhole camera are intentionally blurred and difficult for humans to visually perceive due to the superposition of multiple images from different viewpoints, or because the subject image is difficult to focus on due to the absence of lenses. Therefore, images captured by a multi-pinhole camera are particularly suitable for building image recognition systems in environments where privacy protection is necessary, such as homes or indoors.

[0012] In this image recognition system, a multi-pinhole camera captures an image of the target area, and the captured image is input to a classifier. The classifier then uses a trained identification model to identify faces contained in the input image. In this way, because the target area is captured by a multi-pinhole camera, even if the captured image is leaked externally, the image is difficult for humans to visually recognize, thus protecting the privacy of the subject.

[0013] To train such a classifier, the imaging method described in Non-Patent Document 1 above creates a training dataset by displaying an image captured by a conventional camera on a display, and then capturing the image displayed on the display with a lensless camera. However, while multi-pinhole cameras and lensless cameras can acquire images with different degrees of blur due to the degree of superposition of multiple images depending on the distance to the subject, conventional imaging methods can only acquire one blurred image from a single image displayed on the display, corresponding to the distance between the camera and the display. The degree of blur in the image actually used for image recognition may differ from the degree of blur in the image used for training. Therefore, when a machine learning model is trained using images captured by conventional imaging methods, it has been difficult to improve the recognition accuracy of the machine learning model.

[0014] Therefore, the inventors devised an information processing method in which, during the stage of accumulating training datasets, images displayed on a display are captured while changing the distance between the display and the camera. This method allows for the accumulation of training datasets containing training images with various degrees of blur depending on the distance between the display and the camera. Based on this finding, the inventors conceived the present disclosure, realizing that it is possible to improve the recognition accuracy of machine learning models while protecting the privacy of subjects.

[0015] In order to solve the above problems, an information processing system according to an aspect of the present disclosure includes an image acquisition unit that acquires a first training image of a machine learning model from a first storage unit, a display device, a distance acquisition unit that acquires the distance between the imaging device that acquires a blurred image by imaging, and a display control unit that causes the first training image to be displayed on the display device based on the distance, an imaging control unit that causes the imaging device to image the first training image displayed on the display device to acquire a second training image, and a storage control unit that stores, in a second storage unit, a data set including a pair of the second training image obtained by imaging of the imaging device and correct answer information corresponding to the first training image. The display control unit changes the display size of the first training image so that the size of the second training image obtained by imaging of the imaging device is maintained when the distance is changed, causes the first training image with the changed display size to be displayed on the display device, and the imaging control unit causes the imaging device to image the first training image when the distance is changed.

[0016] According to this configuration, the first training image of the machine learning model is displayed on the display device, and the first training image displayed on the display device is imaged by the imaging device to acquire a second training image. The imaging device acquires a blurred image by imaging. A data set including a pair of the second training image obtained by imaging of the imaging device and correct answer information is stored in the second storage unit. Then, when the distance between the display device and the imaging device is changed, the display size of the first training image is changed so that the size of the second training image obtained by imaging of the imaging device is maintained, and the first training image with the changed display size is displayed on the display device. Then, when the distance is changed, the first training image is imaged by the imaging device.

[0017] Therefore, it is possible to accumulate a data set including a pair of training images with various degrees of blur corresponding to the distance between the display device and the imaging device and correct answer information, and it is possible to improve the recognition accuracy of the machine learning model while protecting the privacy of the subject.

[0018] Furthermore, the above-described information processing system may further include a change instruction unit that instructs a change in the distance between the display device and the imaging device.

[0019] With this configuration, a change in the distance between the display device and the imaging device is instructed, allowing the distance between the display device and the imaging device to be changed to a predetermined distance, and training images with a desired degree of blur corresponding to the changed distance can be acquired.

[0020] Furthermore, in the above-described information processing system, the change instruction unit may instruct the change of distance multiple times, the display control unit may change the display size of the first training image each time the distance is changed multiple times so as to maintain the size of the second training image obtained by imaging with the imaging device, and display the first training image with the changed display size on the display device, the imaging control unit may have the imaging device capture the first training image each time the distance is changed multiple times to obtain multiple second training images, and the storage control unit may store a dataset in the second storage unit that includes each of the multiple second training images obtained by imaging with the imaging device and the correct answer information.

[0021] With this configuration, the distance between the display device and the imaging device is changed multiple times, and multiple second training images with different degrees of blur are acquired for a single first training image. As a result, a dataset can be accumulated that includes pairs of multiple training images with different degrees of blur and the correct answer information.

[0022] Furthermore, in the above-described information processing system, the change instruction unit may instruct a moving device that moves at least one of the imaging device and the display device to move at least one of the imaging device and the display device.

[0023] This configuration allows the distance between the imaging device and the display device to be changed automatically, reducing the need for the user to move at least one of the imaging device or the display device.

[0024] Furthermore, in the above-described information processing system, the change instruction unit may move at least one of the imaging device and the display device in the optical axis direction of the imaging device.

[0025] With this configuration, at least one of the imaging device and the display device is moved in the optical axis direction of the imaging device, so the distance between the display device and the imaging device in the optical axis direction can be changed.

[0026] Furthermore, in the above-described information processing system, the change instruction unit may move at least one of the imaging device and the display device in a direction intersecting the optical axis direction of the imaging device.

[0027] With this configuration, if the second training image is affected by vignetting due to the optical system of the imaging device, at least one of the imaging device and the display device is moved in a direction intersecting the optical axis of the imaging device, so that multiple training images affected by different vignetting effects can be obtained for a single subject.

[0028] Furthermore, in the above-described information processing system, the distance acquisition unit may acquire the distance from a distance measuring device that measures the distance between the display device and the imaging device.

[0029] With this configuration, even if the user manually changes the distance between the display device and the imaging device, the distance is obtained from the distance measuring device that measures the distance between the display device and the imaging device, so the display size of the first training image can be changed according to the measured distance.

[0030] Furthermore, in the above-described information processing system, the first training image stored in the first storage unit may be a clear image acquired by an imaging device different from the imaging device.

[0031] With this configuration, a second training image with blur can be obtained by capturing a first training image, which is an image without blur, using the imaging device.

[0032] Furthermore, in the above-described information processing system, the display control unit may change the display size of the first training image in proportion to the distance between the display device and the imaging device.

[0033] With this configuration, the display size of the first training image is changed in proportion to the distance between the display device and the imaging device. Therefore, if the distance between the display device and the imaging device is known, the display size of the first training image, which is proportional to that distance, can be easily determined.

[0034] Furthermore, the above-described information processing system may further include a training unit that trains the machine learning model using a dataset containing pairs of the second training images and the correct answer information stored in the second memory unit.

[0035] With this configuration, the machine learning model is trained using a dataset containing pairs of second training images and correct information stored in the second memory unit. This improves the recognition ability of the machine learning model to recognize subjects from captured images with varying degrees of blur depending on the distance to the subject.

[0036] Furthermore, in the above-described information processing system, the correct answer information corresponding to the first training image is stored in the first storage unit, and the storage control unit may acquire the correct answer information from the first storage unit.

[0037] This configuration allows for the accumulation of datasets containing training images with varying degrees of blur depending on the distance between the display device and the imaging device, along with ground truth information. This improves the recognition accuracy of machine learning models while protecting the privacy of the subjects.

[0038] Furthermore, in the above-described information processing system, the correct answer information corresponding to the first training image is stored in the first storage unit, and the image acquisition unit may acquire the correct answer information from the first storage unit and output the acquired correct answer information to the storage control unit.

[0039] This configuration allows for the accumulation of datasets containing training images with varying degrees of blur depending on the distance between the display device and the imaging device, along with ground truth information. This improves the recognition accuracy of machine learning models while protecting the privacy of the subjects.

[0040] Furthermore, this disclosure can be implemented not only as an information processing system having the characteristic configuration described above, but also as an information processing method that performs characteristic processing corresponding to the characteristic configuration of the information processing system. Therefore, the same effects as the above-described information processing system can be achieved in the following other embodiments.

[0041] An information processing method according to another aspect of the present disclosure involves a computer acquiring a first training image of a machine learning model from a first storage unit, acquiring the distance between a display device and an imaging device that acquires a blurred image by imaging, displaying the first training image on the display device based on the distance, acquiring a second training image by having the imaging device image the first training image displayed on the display device, storing a dataset in a second storage unit that includes the second training image obtained by imaging by the imaging device and correct answer information corresponding to the first training image, changing the display size of the first training image in the display of the first training image when the distance is changed so that the size of the second training image obtained by imaging by the imaging device is maintained, displaying the first training image with the changed display size on the display device, and having the imaging device image the first training image in the acquisition of the second training image when the distance is changed.

[0042] Embodiments of this disclosure will be described below with reference to the attached drawings. Note that the following embodiments are merely examples of the disclosure and do not limit the technical scope of this disclosure.

[0043] (First Embodiment) Figure 1 is a block diagram showing an example of the overall configuration of the imaging system 1 according to the first embodiment of this disclosure.

[0044] The imaging system 1 comprises an imaging control device 2, a moving device 3, a display device 4, and an imaging device 5.

[0045] The display device 4 is, for example, a liquid crystal display device or an organic EL (Electro-Luminescence) display device. The display device 4 is controlled by the imaging control device 2 and displays the image output from the imaging control device 2. The display device 4 may also be a projector that projects images onto a screen.

[0046] The imaging device 5 is a computational imaging camera such as a lensless camera, a coded aperture camera, a multi-pinhole camera, a lensless multi-pinhole camera, or a light field camera. The imaging device 5 acquires a blurred image by imaging.

[0047] The imaging device 5 is positioned to capture the display screen of the display device 4. In this first embodiment, the imaging device 5 is a lensless multi-pinhole camera in which a mask having a mask pattern with multiple pinholes is positioned to cover the light-receiving surface of the image sensor. In other words, the mask pattern can be said to be positioned between the subject and the light-receiving surface.

[0048] Unlike a normal camera that captures a normal image without blur, the imaging device 5 captures a computational image, which is a blurred image. A computational image is an image in which the subject cannot be recognized by a human eye due to the intentionally created blur.

[0049] Figure 2 is a schematic diagram showing the structure of a multi-pinhole camera 200, which is an example of an imaging device 5. Figure 2 is a top view of the multi-pinhole camera 200.

[0050] The multi-pinhole camera 200 shown in Figure 2 comprises a multi-pinhole mask 201 and an image sensor 202 such as a CMOS. The multi-pinhole camera 200 does not have a lens. The multi-pinhole mask 201 is positioned at a certain distance from the light-receiving surface of the image sensor 202. The multi-pinhole mask 201 has a plurality of pinholes 2011, 2012 arranged randomly or at equal intervals. The plurality of pinholes 2011, 2012 are also called a multi-pinhole. The image sensor 202 acquires an image by capturing the image displayed on the display device 4 through each pinhole 2011, 2012. The image acquired through the pinholes is also called a pinhole image.

[0051] The pinhole image of the subject differs depending on the position and size of each pinhole 2011, 2012. Therefore, the image sensor 202 acquires a superimposed image in which multiple pinhole images are slightly shifted and overlapping (multiple images). The positional relationship of the multiple pinholes 2011, 2012 affects the positional relationship of the multiple pinhole images projected onto the image sensor 202 (i.e., the degree of superposition of the multiple images). The size of the pinholes 2011, 2012 affects the degree of blurring of the pinhole images.

[0052] By using the multi-pinhole mask 201, it is possible to acquire multiple pinhole images with different positions and degrees of blur by superimposing them. In other words, it is possible to acquire computationally captured images in which multiple images and blur are intentionally created. As a result, the captured image becomes a multiple-image and blurred image, and the privacy of the subject can be protected by this blur.

[0053] Furthermore, by changing the number of pinholes, their positions, and their sizes, images with different degrees of blurring can be obtained. In other words, the multi-pinhole mask 201 may have a structure that allows the user to easily attach and detach it. Multiple types of multi-pinhole masks 201 with different mask patterns may be prepared in advance. The multi-pinhole mask 201 may be freely replaced by the user according to the mask pattern of the multi-pinhole camera used during image recognition.

[0054] In addition to replacing the multi-pinhole mask 201, the following various methods can be used to modify the multi-pinhole mask 201. For example, the multi-pinhole mask 201 may be rotatably mounted in front of the image sensor 202 and may be rotated arbitrarily by the user. Alternatively, the multi-pinhole mask 201 may be created by the user drilling holes at any point on a plate mounted in front of the image sensor 202. Alternatively, the multi-pinhole mask 201 may be a liquid crystal mask utilizing a spatial light modulator or the like. A predetermined number of pinholes may be formed at predetermined locations by arbitrarily setting the transmittance of each position within the multi-pinhole mask 201. Furthermore, the multi-pinhole mask 201 may be molded using an expandable material such as rubber. The user may also change the position and size of the pinholes by physically deforming the multi-pinhole mask 201 by applying an external force.

[0055] The multi-pinhole camera 200 is also used in image recognition using a pre-trained machine learning model. Images captured by the multi-pinhole camera 200 are collected as training data. The collected training data is used to train the machine learning model.

[0056] Furthermore, although Figure 2 shows two pinholes 2011 and 2012 arranged horizontally, the disclosure is not limited thereto, and the multi-pinhole camera 200 may have three or more pinholes. Also, the two pinholes 2011 and 2012 may be arranged vertically.

[0057] The moving device 3 is a drive device such as a motor. The moving device 3 is, for example, a rail with a linear motor and is equipped with an imaging device 5. The moving device 3 moves the imaging device 5 by driving the linear motor. The position of the display device 4 is fixed. The moving device 3 moves the imaging device 5 either in a direction toward the display device 4 or in a direction toward the display device 4. In this case, the moving device 3 moves the imaging device 5 along the optical axis of the imaging device 5. As a result, as will be described later, the imaging device 5 can acquire an image in which the position of the pinhole image corresponding to the pinhole located on the optical axis of the imaging device 5 is fixed. Of course, the moving device 3 only needs to be capable of moving the imaging device 5, and could be, for example, a motor that moves wheels, or an aerial vehicle such as a drone.

[0058] The imaging control device 2 is specifically composed of a microprocessor, RAM (Random Access Memory), ROM (Read Only Memory), and a hard disk, which are not shown in the diagram. The RAM, ROM, or hard disk stores computer programs, and the functions of the imaging control device 2 are realized when the microprocessor operates according to the computer programs.

[0059] The imaging control device 2 comprises a first storage unit 21, a second storage unit 22, a third storage unit 23, a distance acquisition unit 24, a movement instruction unit 25, an image acquisition unit 26, a display control unit 27, an imaging control unit 28, and a storage control unit 29.

[0060] The first storage unit 21 stores the first training images for the machine learning model and the correct answer information corresponding to the first training images. The first storage unit 21 stores multiple first training images captured by a normal camera and the correct answer information (annotation information) corresponding to each of the multiple first training images. The first training images are images that include the subject to be recognized by the machine learning model. The first training images are unblurred images acquired by an imaging device different from the imaging device 5.

[0061] The correct answer information differs for each identification task. For example, if the identification task is object detection, the correct answer information is the bounding box representing the area occupied by the detected object in the image. If the identification task is object recognition, the correct answer information is the classification result. If the identification task is image region segmentation, the correct answer information is the region information for each pixel. The first training image and correct answer information stored in the first storage unit 21 are the same information used in machine learning for classifiers that normally use cameras.

[0062] The third storage unit 23 stores the distance between the display device 4 and the imaging device 5 in association with the display size of the first training image to be displayed on the display device 4. The third storage unit 23 stores multiple distances in association with the display size for each of the multiple distances. For example, the multiple distances may include a first distance, which is the closest distance between the display device 4 and the imaging device 5; a second distance, which is longer than the first distance; and a third distance, which is longer than the second distance. The first distance may be, for example, 1 meter, the second distance, for example, 2 meters, and the third distance, for example, 3 meters.

[0063] The distance acquisition unit 24 acquires the distance between the display device 4 and the imaging device 5. When the distance between the display device 4 and the imaging device 5 is changed, the distance acquisition unit 24 acquires the distance between the display device 4 and the imaging device 5 from the third storage unit 23. The distance acquisition unit 24 acquires one of the multiple distances between the display device 4 and the imaging device 5 stored in the third storage unit 23. The distance acquisition unit 24 outputs the acquired distance to the movement instruction unit 25.

[0064] For example, if the third memory unit 23 stores the first distance, the second distance, and the third distance (first distance < second distance < third distance), the distance acquisition unit 24 first acquires the third distance from the third memory unit 23 and outputs it to the movement instruction unit 25. Then, when the acquisition of the first training image at the third distance is completed, the distance acquisition unit 24 acquires the second distance from the third memory unit 23 and outputs it to the movement instruction unit 25. Then, when the acquisition of the first training image at the second distance is completed, the distance acquisition unit 24 acquires the first distance from the third memory unit 23 and outputs it to the movement instruction unit 25.

[0065] The movement instruction unit 25 instructs the movement instruction unit 25 to change the distance between the display device 4 and the imaging device 5. The movement instruction unit 25 instructs the movement device 3 to move the imaging device 5 according to the distance between the display device 4 and the imaging device 5 acquired by the distance acquisition unit 24. The reference position of the imaging device 5 is predetermined. For example, when the distance acquisition unit 24 acquires a third distance, the movement instruction unit 25 instructs the movement device 3 to move the imaging device 5 to the reference position corresponding to the third distance. Then, when the distance acquisition unit 24 acquires a second distance, the movement instruction unit 25 instructs the movement device 3 to move the imaging device 5 from the reference position by the difference between the third distance and the second distance. The movement instruction unit 25 also instructs the movement instruction unit 25 to change the distance between the display device 4 and the imaging device 5 multiple times.

[0066] The moving device 3 moves the imaging device 5 based on instructions from the moving instruction unit 25.

[0067] The image acquisition unit 26 acquires the first training image from the first storage unit 21. The image acquisition unit 26 outputs the first training image acquired from the first storage unit 21 to the display control unit 27.

[0068] The display control unit 27 displays the first training image on the display device 4 based on the distance between the display device 4 and the imaging device 5. When the distance between the display device 4 and the imaging device 5 is changed, the display control unit 27 changes the display size of the first training image so that the size of the second training image obtained by imaging by the imaging device 5 is maintained, and displays the first training image with the changed display size on the display device 4. The display control unit 27 changes the display size of the first training image in proportion to the distance between the display device 4 and the imaging device 5. The display size corresponding to the distance between the display device 4 and the imaging device 5 is stored in advance in the third storage unit 23. Therefore, the display control unit 27 obtains the display size corresponding to the distance between the display device 4 and the imaging device 5 from the third storage unit 23 and changes the display size of the first training image to the obtained display size.

[0069] Furthermore, each time the distance is changed multiple times by the movement instruction unit 25, the display control unit 27 changes the display size of the first training image so that the size of the second training image obtained by imaging by the imaging device 5 is maintained, and displays the first training image with the changed display size on the display device 4.

[0070] The imaging control unit 28 causes the imaging device 5 to capture the first training image displayed on the display device 4, thereby acquiring a second training image. When the distance between the display device 4 and the imaging device 5 changes, the imaging control unit 28 causes the imaging device 5 to capture the first training image. The imaging control unit 28 outputs the acquired second training image to the storage control unit 29. In addition, each time the distance is changed multiple times by the movement instruction unit 25, the imaging control unit 28 causes the imaging device 5 to capture the first training image, thereby acquiring multiple second training images.

[0071] The memory control unit 29 acquires the correct answer information corresponding to the first training image from the first storage unit 21, and stores a dataset in the second storage unit 22 that includes pairs of the second training image obtained by imaging with the imaging device 5 and the correct answer information corresponding to the first training image. The memory control unit 29 stores a dataset in the second storage unit 22 that includes pairs of the second training image obtained by imaging with the imaging device 5 and the correct answer information acquired from the first storage unit 21. In addition, the memory control unit 29 stores a dataset in the second storage unit 22 that includes pairs of each of the multiple second training images obtained by imaging with the imaging device 5 and the correct answer information.

[0072] The second memory unit 22 stores a dataset containing pairs of second training images and correct answer information.

[0073] Next, the dataset creation process in the imaging control device 2 according to the first embodiment of this disclosure will be described.

[0074] Figure 3 is a flowchart illustrating the dataset creation process in the imaging control device 2 according to the first embodiment of this disclosure.

[0075] First, the distance acquisition unit 24 acquires the distance between the display device 4 and the imaging device 5 in order to change the distance between the display device 4 and the imaging device 5 (step S101). The distance acquisition unit 24 outputs the acquired distance to the movement instruction unit 25.

[0076] Next, the movement instruction unit 25 instructs the movement device 3 to move the imaging device 5 to a predetermined imaging position corresponding to the distance acquired by the distance acquisition unit 24 (step S102). Here, the imaging position is a position where the imaging device 5 captures the first training image, which is set in advance to change the distance between the imaging device 5 and the display device 4. When the movement device 3 receives an instruction from the movement instruction unit 25, it moves the imaging device 5 to the predetermined imaging position.

[0077] Next, the image acquisition unit 26 acquires a first training image from the first storage unit 21 (step S103). The image acquisition unit 26 acquires a first training image that has not been captured from among the multiple first training images stored in the first storage unit 21.

[0078] Next, when the distance between the display device 4 and the imaging device 5 is changed, the display control unit 27 changes the display size of the first training image so that the size of the second training image obtained by imaging by the imaging device 5 is maintained (step S104). At this time, the display control unit 27 obtains a display size corresponding to the distance between the display device 4 and the imaging device 5 from the third storage unit 23 and changes the display size of the first training image to the obtained display size.

[0079] Next, the display control unit 27 causes the display size of the first training image to be changed to be displayed on the display device 4 (step S105).

[0080] The display control unit 27 changes the display size of the first training image according to the distance between the display device 4 and the imaging device 5.

[0081] Here, we will explain how the display size of the first training image is changed by the display control unit 27.

[0082] Figure 4 is a schematic diagram illustrating the positional relationship between the display device 4 and the imaging device 5 in the first embodiment. Figure 4 is a top view of the display device 4 and the imaging device 5. Here, distance L is the distance between the multi-pinhole mask 201 of the imaging device 5 and the display screen of the display device 4, and display size W is the display size of the first training image displayed on the display device 4. The imaging device 5 has the same configuration as the lensless multi-pinhole camera 200 shown in Figure 2, with a pinhole 2011 located on the optical axis of the imaging device 5 and a pinhole 2012 located adjacent to pinhole 2011.

[0083] When the distance between the imaging device 5 and the display device 4 is L1, the display size is W1, and when the distance between the imaging device 5 and the display device 4 is L2, the display size is W2, the display control unit 27 determines the display size of the first training image displayed on the display device 4 so as to satisfy the following relationship. Thereby, the first training image on which the second training image of the same size is acquired is displayed on the display device 4 regardless of the change in the distance between the imaging device 5 and the display device 4.

[0084] L1:W1 = L2:W2 That is, the display size is determined so that the distance between the imaging device 5 and the display device 4 and the display size of the first training image displayed on the display device 4 are in a proportional relationship. Thereby, the size and position of the image captured by passing through the pinhole 2011 on the optical axis by the imaging device 5 are constant regardless of the distance L between the imaging device 5 and the display device 4. On the other hand, the imaging range of the image captured by passing through the pinhole 2012 that is not on the optical axis by the imaging device 5 changes depending on the distance L between the imaging device 5 and the display device 4. That is, depending on the distance L between the imaging device 5 and the display device 4, each of the pinholes 2011 and 2012 generates a different parallax. And the parallax information includes depth information.

[0085] FIG. 5 is a schematic diagram for explaining an image captured by passing through two pinholes in a multi-pinhole camera. FIG. 6 is a diagram showing an example of an image captured by the imaging device 5 when the distance between the imaging device 5 and the display device 4 is L1. FIG. 7 is a diagram showing an example of an image captured by the imaging device 5 when the distance between the imaging device 5 and the display device 4 is L2 (L1 < L2). FIG. 5 is a view of the display device 4 and the imaging device 5 seen from above.

[0086] In Figure 5, the same reference numerals are used for components identical to those in Figure 4, and their explanations are omitted. In Figure 5, the display device 4 displays a first training image consisting of the numbers "123456" arranged horizontally. As mentioned above, the display size of the first training image 41 is W1 when the distance between the imaging device 5 and the display device 4 is L1, and the display size of the first training image 42 is W2 when the distance between the imaging device 5 and the display device 4 is L2. Distance L1 is shorter than distance L2, and display size W1 is smaller than display size W2. Note that display sizes W1 and W2 represent the width of the first training images 41 and 42. The aspect ratio of the first training images 41 and 42 is predetermined, for example, 3:2, 4:3, or 16:9. The display size is changed while maintaining the aspect ratio of the first training images 41 and 42.

[0087] The further the distance between the imaging device 5 and the display device 4, the larger the display size of the first training image becomes, and the closer the distance between the imaging device 5 and the display device 4, the smaller the display size of the first training image becomes.

[0088] In Figure 6, Image 2021 represents an image captured by passing through pinhole 2011 on the optical axis when the distance between the imaging device 5 and the display device 4 is L1, and Image 2022 represents an image captured by passing through pinhole 2012 which is not on the optical axis when the distance between the imaging device 5 and the display device 4 is L1. Furthermore, in Figure 7, Image 2023 represents an image captured by passing through pinhole 2011 on the optical axis when the distance between the imaging device 5 and the display device 4 is L2, and Image 2024 represents an image captured by passing through pinhole 2012 which is not on the optical axis when the distance between the imaging device 5 and the display device 4 is L2.

[0089] As shown in Figures 6 and 7, the position and size of images 2021 and 2023, which are captured by passing through the pinhole 2011 on the optical axis, do not change regardless of the distance between the imaging device 5 and the display device 4. In this way, by adjusting the display size and position of the first training images 41 and 42 displayed on the display device 4, the size of the image captured by passing through the pinhole 2011 in the multi-pinhole camera is kept constant.

[0090] On the other hand, as shown in Figures 6 and 7, the size of images 2022 and 2024, which are captured by passing through the pinhole 2012 that is not on the optical axis, does not change, but the imaging range of images 2022 and 2024 changes relative to each other. Images 2022 and 2024 exhibit distance-dependent parallax when compared to images 2021 and 2023 on the left. Here, parallax refers to the amount of shift between multiple images on the image. Therefore, it can be seen that the parallax information includes distance information between the subject and the imaging device 5.

[0091] When the distance between the imaging device 5 and the display device 4 is changed, the display control unit 27 changes the display size of the first training image displayed on the display device 4. At this time, the display control unit 27 changes the display size of the first training image so that the size of the subject image on the second training image does not change even when the distance between the imaging device 5 and the display device 4 is changed. As a result, a second training image is obtained in which multiple subject images are superimposed with a shift amount corresponding to the distance between the imaging device 5 and the display device 4.

[0092] Next, we will explain distance-dependent parallax in multi-pinhole cameras using Figures 8 to 10.

[0093] FIG. 8 is a diagram showing an example of the first training image displayed on the display device 4, FIG. 9 is a diagram showing an example of the second training image obtained by the imaging device 5 imaging the first training image shown in FIG. 8 when the distance between the imaging device 5 and the display device 4 is L1, and FIG. 10 is a diagram showing an example of the second training image obtained by the imaging device 5 imaging the first training image shown in FIG. 8 when the distance between the imaging device 5 and the display device 4 is L2 (L1 < L2).

[0094] In the first training image 51 shown in FIG. 8, a person 301 and a TV 302 are shown. The second training image 52 shown in FIG. 9 is obtained by the imaging device 5 having two pinholes 2011, 2012 imaging the first training image 51 shown in FIG. 8 when the distance between the imaging device 5 and the display device 4 is L1. Also, the second training image 53 shown in FIG. 10 is obtained by the imaging device 5 having two pinholes 2011, 2012 imaging the first training image 51 shown in FIG. 8 when the distance between the imaging device 5 and the display device 4 is L2. However, L1 is shorter than L2.

[0095] In the second training image 52 shown in FIG. 9, the person 303 is the person 301 on the first training image 51 imaged through the pinhole 2011 on the optical axis, the TV 304 is the TV 302 on the first training image 51 imaged through the pinhole 2011 on the optical axis, the person 305 is the person 301 on the first training image 51 imaged through the pinhole 2012 not on the optical axis, and the TV 306 is the TV 302 on the first training image 51 imaged through the pinhole 2012 not on the optical axis.

[0096] Furthermore, in the second training image 53 shown in Figure 10, person 307 is person 301 on the first training image 51, which was captured by passing through the pinhole 2011 on the optical axis; television 308 is television 302 on the first training image 51, which was captured by passing through the pinhole 2011 on the optical axis; person 309 is person 301 on the first training image 51, which was captured by passing through the pinhole 2012 which is not on the optical axis; and television 310 is television 302 on the first training image 51, which was captured by passing through the pinhole 2012 which is not on the optical axis.

[0097] Thus, the second training image obtained from the imaging device 5, which is a multi-pinhole camera, is an image in which multiple subject images are superimposed. The position and size of people 303, 307 and televisions 304, 308, which are captured by passing through the pinhole 2011 on the optical axis, do not change in the captured image. On the other hand, the positions of people 305, 309 and televisions 306, 310, which are captured by passing through the pinhole 2012 which is not on the optical axis, change in the captured image, and the amount of parallax decreases as the distance between the subject and the imaging device 5 increases. In other words, by capturing the first training image while changing the distance L between the imaging device 5 and the display device 4, a second training image with a changed amount of parallax is obtained.

[0098] In this first embodiment, the imaging device 5 does not have a lens, but the imaging device 5 may have an optical system, such as a lens, between the imaging device 5 and the display device 4. By having a lens in the imaging device 5, it is possible to make the distance (optical path length) between the imaging device 5 and the display device 4 longer or shorter than the actual distance. This makes it possible to produce a larger or smaller parallax compared to when no optical system is inserted. If the moving device 3 cannot significantly change the distance between the imaging device 5 and the display device 4, the configuration in which the imaging device 5 has an optical system is effective. Also, when an optical system is inserted between the imaging device 5 and the display device 4, the distance between the imaging device 5 and the display device 4 may be changed by changing the focal length of the optical system, inserting or removing the optical system, rather than physically moving the imaging device 5 or the display device 4.

[0099] Note that the distance L may not be the actual distance measured, but rather the optical path length as described above. The display control unit 27 may determine the display size of the first training image displayed on the display device 4 based on the optical path length between the imaging device 5 and the display device 4.

[0100] Returning to Figure 3, the next step is for the imaging control unit 28 to cause the imaging device 5 to capture the first training image displayed on the display device 4, thereby acquiring a second training image (step S106). The imaging device 5 captures the image so that the display device 4 is within its field of view. As mentioned above, the display size of the first training image is changed according to the distance between the display device 4 and the imaging device 5, so even if the distance between the display device 4 and the imaging device 5 is changed, a second training image with the same image size can be acquired.

[0101] Next, the memory control unit 29 acquires the correct answer information associated with the first training image displayed on the display device 4 from the first memory unit 21, and stores a dataset in the second memory unit 22 that includes pairs of the second training image acquired by the imaging device 5 and the correct answer information associated with the first training image displayed on the display device 4 (step S107). As a result, the second memory unit 22 stores a dataset that includes pairs of the second training image acquired by the imaging device 5 and its correct answer information.

[0102] Next, the image acquisition unit 26 determines whether the imaging device 5 has captured all of the first training images stored in the first storage unit 21 (step S108). If it is determined that the imaging device 5 has not captured all of the first training images (NO in step S108), the process returns to step S103, and the image acquisition unit 26 newly acquires the uncaptured first training images from among the multiple first training images stored in the first storage unit 21. Then, the processes from steps S104 to S106 are executed using the newly acquired first training images. After that, in step S107, a dataset containing pairs of second training images acquired by the imaging device 5 and correct answer information associated with the new first training images displayed on the display device 4 is stored in the second storage unit 22, and then step S108 is executed.

[0103] On the other hand, if it is determined that the imaging device 5 has captured all of the first training images (YES in step S108), the distance acquisition unit 24 determines whether or not imaging by the imaging device 5 has been completed at all of the preset imaging positions (step S109). If it is determined that imaging by the imaging device 5 has not been completed at all of the imaging positions (NO in step S109), the process returns to step S101, and the distance acquisition unit 24 acquires the distance between the display device 4 and the imaging device 5 corresponding to the imaging position that has not been captured, from among the multiple imaging positions stored in the third storage unit 23. Then, the movement instruction unit 25 instructs the movement device 3 to move the imaging device 5 to a predetermined imaging position according to the distance acquired by the distance acquisition unit 24. The movement device 3 moves the imaging device 5 to the predetermined imaging position according to the instructions of the movement instruction unit 25.

[0104] On the other hand, if it is determined that imaging by the imaging device 5 has been completed at all imaging positions (YES in step S109), the process ends.

[0105] In the example shown in Figure 3, the memory control unit 29 obtains the correct answer information associated with the first training image displayed on the display device 4 from the first storage unit 21, but is not limited to this.

[0106] For example, the image acquisition unit 26 may acquire a first training image and correct answer information corresponding to the first training image from the first storage unit 21. In this case, the image acquisition unit 26 outputs the correct answer information acquired from the first storage unit 21 to the storage control unit 29, and the storage control unit 29 stores a dataset in the second storage unit 22 that includes a pair of a second training image acquired by the imaging device 5 and correct answer information associated with the first training image displayed on the display device 4, that is, the correct answer information output from the image acquisition unit 26.

[0107] As described above, in the imaging system 1 of this first embodiment, the display size of the first training image displayed on the display device 4 is changed according to the distance between the imaging device 5 and the display device 4. As a result, the position and size of the subject image captured by passing through the pinhole on the optical axis do not change, but images are captured with a changed amount of parallax between pinholes in the multi-pinhole camera. Since the position and size of the subject image captured by passing through the pinhole on the optical axis do not change, the correct answer information attached to the first training image displayed on the display device 4 can be used as is. As a result, it is possible to create a dataset that corresponds to a multi-pinhole camera. By performing training processing using machine learning with such a dataset, it is possible to achieve high-precision recognition that is independent of the distance to the subject.

[0108] In the above description, the movement instruction unit 25 moves the imaging device 5 in the optical axis direction of the imaging device 5, but the imaging device 5 may also be moved in a direction intersecting the optical axis direction of the imaging device 5. When a lensless camera or the like is used as the imaging device 5, a phenomenon called vignetting occurs, where the amount of light decreases from near the center of the image sensor 202 towards the outer edge. phenomenon This occurs. The second training image is affected by vignetting. In response, the movement instruction unit 25 moves the imaging device 5 in a direction intersecting the optical axis, and by changing the positional relationship between the imaging device 5 and the display device 4, multiple second training images can be obtained for a single subject, each affected by different vignetting.

[0109] Furthermore, the movement instruction unit 25 may move the display device 4 instead of the imaging device 5. For example, if the display device 4 is a display and the movement device 3 is a rail with a linear motor, the movement device 3 may move the display device 4 by mounting the display device 4 and driving the linear motor. Also, for example, if the display device 4 is a projector and a screen, the movement device 3 may move at least one of the screen and the projector, or change the distance between the screen and the projector. In addition, the display control unit 27 may change the size of the first training image projected by the projector.

[0110] Here, we will describe a modified version of the first embodiment in which the display device 4 is moved instead of the imaging device 5.

[0111] Figure 11 is a flowchart illustrating the dataset creation process in the imaging control device 2 according to a modified example of the first embodiment of this disclosure.

[0112] In a modified version of the first embodiment, the movement instruction unit 25 instructs the movement device 3 to move the display device 4 according to the distance between the display device 4 and the imaging device 5 acquired by the distance acquisition unit 24. The reference position of the display device 4 is predetermined. For example, when the distance acquisition unit 24 acquires a third distance, the movement instruction unit 25 instructs the movement device 3 to move the display device 4 to the reference position corresponding to the third distance. Then, when the distance acquisition unit 24 acquires a second distance, the movement instruction unit 25 instructs the movement device 3 to move the display device 4 from the reference position by the difference between the third distance and the second distance.

[0113] The moving device 3 moves the display device 4 based on instructions from the moving instruction unit 25.

[0114] The process in step S121 shown in Figure 11 is the same as the process in step S101 shown in Figure 3, so the explanation is omitted.

[0115] Next, the movement instruction unit 25 instructs the movement device 3 to move the display device 4 to a predetermined display position corresponding to the distance acquired by the distance acquisition unit 24 (step S122). Upon receiving the instruction from the movement instruction unit 25, the movement device 3 moves the display device 4 to the predetermined display position.

[0116] The processes in steps S123 to S128 shown in Figure 11 are the same as the processes in steps S103 to S108 shown in Figure 3, so their explanation is omitted.

[0117] If it is determined that the imaging device 5 has captured all of the first training images (YES in step S128), the distance acquisition unit 24 determines whether imaging by the imaging device 5 has been completed at all of the pre-set display positions (step S129). Here, the display position is a position where the display device 4 displays the first training image, which is pre-set to change the distance between the imaging device 5 and the display device 4. If it is determined that imaging by the imaging device 5 has not been completed at all of the display positions (NO in step S129), the process returns to step S121, and the distance acquisition unit 24 acquires the distance between the display device 4 and the imaging device 5 corresponding to the uncaptured display positions from among the multiple imaging positions stored in the third storage unit 23. Then, the movement instruction unit 25 instructs the movement device 3 to move the display device 4 to a predetermined display position according to the distance acquired by the distance acquisition unit 24. The movement device 3 moves the display device 4 to the predetermined display position according to the instructions of the movement instruction unit 25.

[0118] On the other hand, if it is determined that imaging by the imaging device 5 has been completed at all display positions (YES in step S129), the process ends.

[0119] In this first embodiment, the moving device 3 moves either the imaging device 5 or the display device 4, but the disclosure is not limited thereto, and both the imaging device 5 and the display device 4 may be moved.

[0120] (Second Embodiment) In the first embodiment described above, a plurality of distances between the display device 4 and the imaging device 5 are predetermined, and at least one of the display device 4 and the imaging device 5 is moved to each position corresponding to the predetermined plurality of distances, and the display size of the first training image displayed on the display device 4 is changed according to the distance between the imaging device 5 and the display device 4. In contrast, in the second embodiment, at least one of the display device 4 and the imaging device 5 is moved to an arbitrary position, the distance between the imaging device 5 and the display device 4 is measured, and the display size of the first training image displayed on the display device 4 is changed according to the measured distance.

[0121] Figure 12 is a block diagram showing an example of the overall configuration of imaging system 1A according to the second embodiment of this disclosure. In Figure 12, the same reference numerals are used for the same components as in Figure 1, and their descriptions are omitted.

[0122] The imaging system 1A comprises an imaging control device 2A, a moving device 3, a display device 4, an imaging device 5, and a distance measuring device 6.

[0123] The distance measuring device 6 is, for example, a laser distance meter or a ToF(T)meter. i The camera (me of flight) measures the distance between the display device 4 and the imaging device 5 according to the instructions of the imaging control device 2A. For example, the display device 4 and the imaging device 5 are mounted on rails. At least one of the display device 4 and the imaging device 5 is movable on the rails in the direction of the optical axis of the imaging device 5.

[0124] The imaging control device 2A includes a first storage unit 21, a second storage unit 22, a distance acquisition unit 24A, a movement instruction unit 25A, an image acquisition unit 26, a display control unit 27A, an imaging control unit 28, a storage control unit 29, and a fourth storage unit 30.

[0125] The movement instruction unit 25A instructs a change in the distance between the display device 4 and the imaging device 5. AThe system receives input commands from the user to move the imaging device 5 and instructs the moving device 3 to move the imaging device 5 in accordance with the received input commands. The moving instruction unit 25A receives input commands to move the imaging device 5 closer to the display device 4 or to move the imaging device 5 away from the display device 4.

[0126] In this second embodiment, the movement instruction unit 25A moves the imaging device 5 in the optical axis direction of the imaging device 5, but the imaging device 5 may also be moved in a direction intersecting the optical axis direction of the imaging device 5. Furthermore, the movement instruction unit 25A may move the display device 4 instead of the imaging device 5. Moreover, the movement instruction unit 25A may move both the imaging device 5 and the display device 4.

[0127] Furthermore, the imaging system 1A does not necessarily have to be equipped with a moving device 3, and the imaging control device 2A does not necessarily have to be equipped with a moving instruction unit 25A. In this case, the user may manually change the distance between the display device 4 and the imaging device 5. For example, the user may move at least one of the imaging device 5 and the display device 4, which are mounted on rails.

[0128] When the movement of the imaging device 5 by the moving device 3 is complete, the distance acquisition unit 24A instructs the distance measuring device 6 to measure the distance between the display device 4 and the imaging device 5. The distance acquisition unit 24A acquires the distance between the display device 4 and the imaging device 5 measured by the distance measuring device 6. The distance acquisition unit 24A may also accept an input operation from the user to measure the distance between the display device 4 and the imaging device 5, and instruct the distance measuring device 6 to measure the distance according to the received input operation. In particular, if the user manually moves the imaging device 5, the distance acquisition unit 24A may accept an instruction from the user to start measurement by the distance measuring device 6.

[0129] The fourth storage unit 30 stores the reference distance between the display device 4 and the imaging device 5, and the reference display size of the first training image to be displayed on the display device 4, in association with each other. The display size of the first training image is proportional to the distance between the display device 4 and the imaging device 5. Therefore, if the reference distance between the display device 4 and the imaging device 5 and the reference display size of the first training image relative to the reference distance are predetermined, the display size can be determined from the measured distance.

[0130] The display control unit 27A causes the first training image to be displayed on the display device 4 based on the distance measured by the distance measuring device 6. When the distance between the display device 4 and the imaging device 5 is changed, the display control unit 27A changes the display size of the first training image so that the size of the second training image obtained by imaging by the imaging device 5 is maintained, and displays the first training image with the changed display size on the display device 4.

[0131] The display control unit 27A changes the display size of the first training image in proportion to the distance between the display device 4 and the imaging device 5. A reference display size corresponding to the reference distance between the display device 4 and the imaging device 5 is stored in advance in the fourth storage unit 30. The reference distance is, for example, the distance when the display device 4 and the imaging device 5 are positioned at their furthest relative positions. Therefore, the display control unit 27A uses the reference distance and reference display size stored in the fourth storage unit 30 to calculate the display size from the distance measured by the distance measuring device 6, and changes the display size of the first training image to the calculated display size.

[0132] Next, the dataset creation process in the imaging control device 2A according to the second embodiment of this disclosure will be described.

[0133] Figure 13 is a flowchart illustrating the dataset creation process in the imaging control device 2A according to the second embodiment of this disclosure.

[0134] First, the movement instruction unit 25A instructs the movement device 3 to move the imaging device 5 to an arbitrary imaging position (step S141). The movement device 3 moves the imaging device 5 to the arbitrary position in accordance with the instructions of the movement instruction unit 25A.

[0135] Next, the distance acquisition unit 24A instructs the distance measuring device 6 to measure the distance between the display device 4 and the imaging device 5 (step S142). The distance measuring device 6 measures the distance between the display device 4 and the imaging device 5 according to the instructions of the distance acquisition unit 24A.

[0136] Next, the distance acquisition unit 24A acquires the distance between the display device 4 and the imaging device 5, which has been measured by the distance measuring device 6, from the distance measuring device 6 (step S143).

[0137] Next, the display control unit 27 A When the distance between the display device 4 and the imaging device 5 is changed according to the distance measured by the distance measuring device 6, the display control unit 27A determines the display size of the first training image in order to maintain the size of the second training image obtained by imaging by the imaging device 5 (step S144). The display control unit 27A determines the display size of the first training image from the distance measured by the distance measuring device 6 using the reference distance and reference display size stored in the fourth storage unit 30. More specifically, the display control unit 27A calculates the display size of the first training image to be changed by multiplying the value obtained by dividing the distance measured by the distance measuring device 6 by the reference distance by the reference display size.

[0138] The fourth storage unit 30 may store a table that associates the distance between the display device 4 and the imaging device 5 with the display size of the first training image. In this case, the display control unit 27A may read the display size associated with the distance measured by the distance measuring device 6 from the fourth storage unit 30 and determine the read display size as the display size of the first training image.

[0139] Next, the image acquisition unit 26 acquires a first training image from the first storage unit 21 (step S145). The image acquisition unit 26 acquires a first training image that has not been captured from among the multiple first training images stored in the first storage unit 21.

[0140] Next, when the distance between the display device 4 and the imaging device 5 is changed, the display control unit 27A changes the display size of the first training image so that the size of the second training image obtained by imaging by the imaging device 5 is maintained (step S146). At this time, the display control unit 27A changes the display size of the first training image to the determined display size.

[0141] Next, the display control unit 27A causes the display device 4 to display the first training image with the changed display size (step S147).

[0142] The processes in steps S148 to S149 shown in Figure 11 are the same as the processes in steps S106 to S107 shown in Figure 3, so their explanation is omitted.

[0143] Next, the image acquisition unit 26 determines whether the imaging device 5 has captured all of the first training images stored in the first storage unit 21 (step S150). If it is determined that the imaging device 5 has not captured all of the first training images (NO in step S150), the process returns to step S145, and the image acquisition unit 26 newly acquires the uncaptured first training images from among the multiple first training images stored in the first storage unit 21. Then, the processes from steps S146 to S148 are executed using the newly acquired first training images. After that, in step S149, a dataset containing pairs of second training images acquired by the imaging device 5 and correct answer information associated with the new first training images displayed on the display device 4 is stored in the second storage unit 22, and then step S150 is executed.

[0144] On the other hand, if it is determined that the imaging device 5 has captured all of the first training images (YES in step S150), the distance acquisition unit 24A determines whether or not imaging by the imaging device 5 has been completed at all of the planned imaging positions (step S151). Note that multiple imaging positions and the number of imaging cycles are predetermined. If the imaging position is changed and imaging is performed a predetermined number of times, the distance acquisition unit 24A determines that imaging by the imaging device 5 has been completed at all imaging positions. If it is determined that imaging by the imaging device 5 has not been completed at all imaging positions (NO in step S151), the process returns to step S141, and the movement instruction unit 25A instructs the movement device 3 to move the imaging device 5 to an arbitrary imaging position. The distance acquisition unit 24A instructs the measurement of the distance between the display device 4 and the imaging device 5 at the new imaging position and acquires the distance measured by the distance measuring device 6.

[0145] On the other hand, if it is determined that imaging by the imaging device 5 has been completed at all imaging positions (YES in step S151), the process ends.

[0146] In the example shown in Figure 13, the memory control unit 29 obtains the correct answer information associated with the first training image displayed on the display device 4 from the first storage unit 21, but is not limited to this.

[0147] As described in the first embodiment, the image acquisition unit 26 may acquire a first training image and correct answer information corresponding to the first training image from the first storage unit 21. In this case, the image acquisition unit 26 outputs the correct answer information acquired from the first storage unit 21 to the storage control unit 29, and the storage control unit 29 stores a dataset in the second storage unit 22 that includes a pair of a second training image acquired by the imaging device 5 and correct answer information associated with the first training image displayed on the display device 4, that is, the correct answer information output from the image acquisition unit 26.

[0148] As described above, in the imaging system 1A of this second embodiment, the display size of the first training image displayed on the display device 4 is changed according to the distance between the imaging device 5 and the display device 4. As a result, the position and size of the subject image captured by passing through the pinhole on the optical axis do not change, but an image is captured with a changed amount of parallax between pinholes in the multi-pinhole camera. Since the position and size of the subject image captured by passing through the pinhole on the optical axis do not change, the correct answer information attached to the first training image displayed on the display device 4 can be used as is. As a result, it is possible to create a dataset that corresponds to a multi-pinhole camera. By performing training processing using machine learning with such a dataset, it is possible to achieve high-precision recognition that is independent of the distance to the subject.

[0149] Furthermore, the imaging system 1A does not necessarily have to be equipped with a moving device 3, and at least one of the imaging device 5 and the display device 4 may be moved manually by the user. In this case, the distance measuring device 6 measures the distance between the display device 4 and the imaging device 5 according to the instructions of the distance acquisition unit 24A. This simplifies the configuration of the imaging system 1A and reduces the manufacturing cost of the imaging system 1A.

[0150] (Third embodiment) In the first and second embodiments, a dataset containing pairs of second training images and correct information is stored in the second memory unit, while in the third embodiment, a machine learning model is trained using the dataset containing pairs of second training images and correct information stored in the second memory unit.

[0151] Figure 14 is a block diagram showing an example of the overall configuration of imaging system 1B according to the third embodiment of this disclosure. In Figure 14, the same reference numerals are used for the same components as in Figure 1, and their descriptions are omitted.

[0152] The imaging system 1B comprises an imaging control device 2B, a mobile device 3, a display device 4, and an imaging device 5.

[0153] The imaging control device 2B includes a first storage unit 21, a second storage unit 22, a third storage unit 23, a distance acquisition unit 24, a movement instruction unit 25, an image acquisition unit 26, a display control unit 27, an imaging control unit 28, a storage control unit 29, a training unit 31, and a model storage unit 32.

[0154] The training unit 31 trains a machine learning model using a dataset containing pairs of second training images and correct answer information stored in the second memory unit 22. In this third embodiment, the machine learning model applied to the classifier is a machine learning model using a neural network such as deep learning, but other machine learning models may also be used. For example, the machine learning model may be a machine learning model using random forest or genetic programming.

[0155] Machine learning in the training unit 31 is implemented, for example, by backpropagation (BP) in deep learning. Specifically, the training unit 31 inputs a second training image into the machine learning model and obtains the recognition result output by the machine learning model. The training unit 31 then adjusts the machine learning model so that the recognition result becomes the correct information. The training unit 31 improves the recognition accuracy of the machine learning model by repeating the adjustment of the machine learning model for multiple sets (e.g., thousands of sets) of different second training images and correct information.

[0156] The model memory unit 32 stores a trained machine learning model. This machine learning model is also an image recognition model used for image recognition.

[0157] In this third embodiment, the imaging control device 2B includes a training unit 31 and a model storage unit 32. However, the disclosure is not limited thereto, and an external computer connected to the imaging control device 2B via a network may also include a training unit 31 and a model storage unit 32. In this case, the imaging control device 2B may further include a communication unit for transmitting a dataset to the external computer. Alternatively, an external computer connected to the imaging control device 2B via a network may also include a model storage unit 32. In this case, the imaging control device 2B may further include a communication unit for transmitting a trained machine learning model to the external computer.

[0158] In this third embodiment, the imaging system 1B can use the depth information of the subject, which is included in the disparity information, as training data, and is therefore effective in improving the recognition ability of the machine learning model. For example, the machine learning model can recognize that a small object in the image is a subject located at a distance, and can prevent it from being mistaken for dust or ignored. As a result, the machine learning model constructed by machine learning using the second training image can improve its recognition performance.

[0159] In this third embodiment, the imaging system 1B displays a first training image stored in the first storage unit 21, stores a plurality of second training images obtained by changing the distance between the display device 4 and the imaging device 5 in the second storage unit 22, and uses the stored plurality of second training images for training. The training unit 31 may use the plurality of second training images stored in the second storage unit 22 during the training process. Alternatively, the training unit 31 may not use all of the second training images, but randomly select only some of the second training images and use only the selected subset of second training images. Furthermore, the training unit 31 may swap only a portion of the plurality of second training images to create a single image by patching together second training images of various depths, and use the created image for training.

[0160] As described above, the imaging system 1B is effective not only for training and optimizing machine learning parameters, but also for optimizing the device parameters of the imaging device 5. When a multi-pinhole camera is used as the imaging device 5, the recognition performance and privacy protection performance of the imaging device 5 depend on device parameters such as the size of each pinhole, the shape of each pinhole, the arrangement of each pinhole, and the number of pinholes. Therefore, in order to realize an optimal recognition system, it is necessary to optimize not only the machine learning parameters, but also the device parameters of the imaging device 5, such as the size of each pinhole, the shape of each pinhole, the arrangement of each pinhole, and the number of pinholes. In this third embodiment, the imaging system 1B can select device parameters that have a high recognition rate and high privacy protection performance as the optimal device parameters by training and evaluating the second training image obtained when the device parameters of the imaging device 5 are changed.

[0161] In each of the above embodiments, each component may be implemented by dedicated hardware or by executing a software program suitable for each component. Each component may also be implemented by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory. Furthermore, the program may be executed by another independent computer system by recording and transferring the program on a recording medium, or by transferring the program over a network.

[0162] Some or all of the functions of the apparatus according to the embodiments of this disclosure are typically implemented as an integrated circuit, or LSI (Large Scale Integration). These may be individually integrated onto a single chip, or some or all of them may be integrated onto a single chip. Furthermore, the integration is not limited to LSIs, but may also be implemented using dedicated circuits or general-purpose processors. An FPGA (Field Programmable Gate Array) that can be programmed after the LSI is manufactured, or a reconfigurable processor that can reconfigure the connections and settings of the circuit cells inside the LSI may also be used.

[0163] Furthermore, some or all of the functions of the apparatus according to the embodiments of this disclosure may be realized by a processor such as a CPU executing a program.

[0164] Furthermore, all figures used above are illustrative examples provided to illustrate this disclosure, and this disclosure is not limited to these illustrative figures.

[0165] Furthermore, the order in which the steps shown in the flowchart above are performed is illustrative for the purpose of specifically illustrating this disclosure, and other orders are acceptable as long as similar effects are achieved. Also, some of the steps above may be performed simultaneously (in parallel) with other steps. [Industrial applicability]

[0166] The technology disclosed herein is useful as a technique for creating datasets used to train machine learning models, as it can improve the recognition accuracy of machine learning models while protecting the privacy of the subjects.

Claims

1. An image acquisition unit that acquires the first training image for the machine learning model from the first storage unit, A distance acquisition unit that acquires the distance between a display device and an imaging device that acquires a blurred image by imaging, A display control unit that displays the first training image on the display device based on the distance, An imaging control unit that causes the imaging device to capture the first training image displayed on the display device and acquire a second training image, A storage control unit stores a dataset in a second storage unit that includes a pair of the second training image obtained by imaging with the imaging device and the correct answer information corresponding to the first training image. Equipped with, When the distance is changed, the display control unit changes the display size of the first training image so as to maintain the size of the second training image obtained by imaging by the imaging device, and displays the first training image with the changed display size on the display device. The imaging control unit causes the imaging device to capture the first training image, whose display size has been changed. Information processing system.

2. The system further includes a change instruction unit that instructs a change in the distance between the display device and the imaging device. The information processing system according to claim 1.

3. The change instruction unit instructs the change of the distance multiple times, The display control unit changes the display size of the first training image each time the distance is changed multiple times, so as to maintain the size of the second training image obtained by imaging by the imaging device, and displays the first training image with the changed display size on the display device. The imaging control unit causes the imaging device to capture the first training image, whose display size has been changed, and acquires a plurality of second training images. The memory control unit stores in the second memory unit a dataset containing a pair of each of the plurality of second training images obtained by imaging by the imaging device and the correct answer information. The information processing system according to claim 2.

4. The change instruction unit instructs a moving device that moves at least one of the imaging device and the display device to move at least one of the imaging device and the display device. The information processing system according to claim 2 or 3.

5. The change instruction unit moves at least one of the imaging device and the display device in the optical axis direction of the imaging device. The information processing system according to claim 4.

6. The change instruction unit moves at least one of the imaging device and the display device in a direction intersecting the optical axis direction of the imaging device. The information processing system according to claim 4.

7. The distance acquisition unit acquires the distance from a distance measuring device that measures the distance between the display device and the imaging device. The information processing system according to any one of claims 1 to 3.

8. The first training image stored in the first memory unit is a clear image acquired by an imaging device different from the imaging device. The information processing system according to any one of claims 1 to 3.

9. The display control unit changes the display size of the first training image in proportion to the distance between the display device and the imaging device. The information processing system according to any one of claims 1 to 3.

10. The system further comprises a training unit that trains the machine learning model using a dataset containing pairs of the second training images and the correct answer information stored in the second memory unit. The information processing system according to any one of claims 1 to 3.

11. The correct answer information corresponding to the first training image is stored in the first storage unit. The memory control unit acquires the correct answer information from the first memory unit. The information processing system according to any one of claims 1 to 3.

12. The correct answer information corresponding to the first training image is stored in the first storage unit. The image acquisition unit acquires the correct answer information from the first storage unit and outputs the acquired correct answer information to the storage control unit. The information processing system according to any one of claims 1 to 3.

13. Computers The first training image for the machine learning model is obtained from the first storage unit. The distance between the display device and the imaging device that acquires a blurred image through imaging is obtained. Based on the distance, the first training image is displayed on the display device. The first training image displayed on the display device is captured by the imaging device to acquire a second training image. A dataset including a pair of the second training image obtained by imaging with the imaging device and the correct answer information corresponding to the first training image is stored in the second storage unit. In the display of the first training image, when the distance is changed, the display size of the first training image is changed so that the size of the second training image obtained by imaging by the imaging device is maintained, and the first training image with the changed display size is displayed on the display device. In acquiring the second training image, the first training image, whose display size has been changed, is to be captured by the imaging device. Information processing methods.

Citation Information

Patent Citations

  • Mark recognizing device and mark recognizing method

    JP1996122267A

  • Learning device, method for learning, and program

    JP2019200769A

  • Identification system, identification device, method for identification, and program

    JP2019200772A