Object Classification Device and Object Classification Method
The use of a filter array and machine learning classification model for multi-wavelength image capture addresses the limitations of conventional RGB and hyperspectral cameras, achieving high-precision object recognition with reduced computational demands.
Patent Information
- Application Number
- JP2023147062
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-10-15
- Filing Date
- 2023-09-11
- Publication Date
- 2025-07-28
- Estimated Expiration
- 2039-09-24
AI Technical Summary
Conventional RGB images and hyperspectral cameras face limitations in object recognition accuracy due to trade-offs between wavelength resolution and spatial resolution, and high computational demands for processing large multi-wavelength data sets.
An object recognition method using a filter array with translucent filters two-dimensionally arranged to capture multi-wavelength images, combined with a machine learning classification model trained on diverse image data sets, enables high-precision object recognition without reconstructing images in each wavelength band.
This approach enhances object recognition accuracy by maintaining high spatial resolution and reducing computational resources, allowing for efficient and precise identification of objects.
Smart Images

Figure 0007713657000001 
Figure 0007713657000002 
Figure 0007713657000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to an object classification device and an object classification method.
Background Art
[0002] In object recognition using machine learning, generally, a monochrome image or an RGB image is utilized as learning data. On the other hand, attempts have also been made to perform object recognition using a multi-spectral image that contains information on more wavelengths than an RGB image.
[0003] Patent Document 1 discloses a spectral camera in which a plurality of filters that transmit light in different wavelength ranges are spatially arranged in a mosaic pattern as a sensor for acquiring a multi-spectral image. Patent Document 2 discloses a method of learning an image of immune cells for a plurality of image channels by a convolutional neural network in order to improve the recognition accuracy of immune cells in an image. Patent Document 3 discloses a method of machine learning using a multi-spectral image or a hyperspectral image as training data.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Patent Document 2
Patent Document 3
Summary of the Invention
Problems to be Solved by the Invention
[0005] The present disclosure provides a novel object recognition method that enables high-precision object recognition from encoded image data.
Means for Solving the Problems
[0006] An object recognition method according to one aspect of the present disclosure includes acquiring image data of an image including feature information indicating features of an object, and recognizing the object included in the image based on the feature information. The image data includes an image sensor and a filter array disposed in an optical path of light incident on the image sensor, the filter array including a plurality of translucent filters two-dimensionally arranged along a plane intersecting the optical path, the plurality of filters including two or more filters having mutually different wavelength dependencies of light transmittance, and the light transmittance of each of the two or more filters having a maximum value in a plurality of wavelength ranges. The image is acquired by imaging the image with a first imaging device including the filter array.
Advantages of the Invention
[0007] According to the present disclosure, high-precision object recognition becomes possible.
Brief Description of the Drawings
[0008]
Figure 1
Figure 2A
Figure 2B
Figure 2C
Figure 2D
Figure 3A
Figure 3B
Figure 4A
Figure 4B
Figure 4C
Figure 4D
Figure 4E
Figure 5A
Figure 5B
Figure 5C
Figure 6A
Figure 6B
Figure 6C
Figure 7
Figure 8
[0009] Before describing embodiments of the present disclosure, the findings on which the present disclosure is based will be described.
[0010] In object recognition using conventional RGB images, there were limitations in its recognition ability. For example, there were cases where it was impossible to distinguish between a real object and its signboard or poster. This is generally due to the small difference in the amounts of the R, G, and B components of the light reflected from the real object and the amounts of the R, G, and B components of the light reflected from the signboard or poster. To distinguish between a real object and its signboard or poster, for example, it is conceivable to utilize multi-wavelength spectral data. This may enable detection of minute differences in spectral data due to differences in the materials of the objects.
[0011] In conventional hyperspectral cameras, as disclosed in, for example, Patent Document 1, a plurality of wavelength filters with different transmission wavelength ranges are two-dimensionally arranged. When an image of one frame is acquired in a single shot like video shooting, the number of wavelength ranges and the spatial resolution are in a trade-off relationship. That is, in order to acquire a multi-wavelength image, if many filters with different transmission wavelength ranges are spatially dispersed and arranged, the spatial resolution of the image acquired for each wavelength range becomes low. Therefore, even if a hyperspectral image is used for object recognition in the expectation of improving the recognition accuracy of an object, in reality, due to the low spatial resolution, the recognition accuracy may decrease.
[0012] It is also conceivable to improve both the wavelength resolution and the resolution by increasing the number of pixels of the image sensor. In this case, a large-capacity three-dimensional data in which multi-wavelength data is added to two-dimensional data of space is handled. When applying machine learning to such large-sized data, a lot of time or resources are consumed for preprocessing, learning, communication, and storage of the data.
[0013] Based on the above considerations, the inventor came up with the object recognition method described in the following items.
[0014] [Item 1] The object recognition method according to the first item includes acquiring image data of an image including feature information indicating features of an object, and recognizing the object included in the image based on the feature information. The image data is an image sensor and a filter array disposed in an optical path of light incident on the image sensor, and includes a plurality of light-transmissive filters two-dimensionally arranged along a plane intersecting the optical path. The plurality of filters includes two or more filters having different wavelength dependencies of light transmittance, and the light transmittance of each of the two or more filters has a maximum value in a plurality of wavelength ranges. The image is acquired by imaging the image with a first imaging device including a filter array.
[0015] [Item 2] In the object recognition method according to the first item, recognizing the object is performed by applying a classification model learned by a machine learning algorithm to the image data. The classification model may be pre-learned by a plurality of first training data sets each including learning image data and label data for identifying the object included in the learning image indicated by the learning image data.
[0016] [Item 3] In the object recognition method according to the second item, the plurality of learning image data included in the plurality of first training data sets may include learning image data generated by a second imaging device different from the first imaging device.
[0017] [Item 4] In the object recognition method according to the third item, the second imaging device may include a filter array having characteristics equivalent to those of the filter array in the first imaging device.
[0018] [Item 5] The object recognition method according to any one of the second to fourth items may further include further learning the classification model by a second training data set including the image data and second label data for identifying the object after the object is recognized.
[0019] [Item 6] In the object recognition method according to any one of Items 2 to 5, the position of the object in the learning image data included in the plurality of first training data sets may be different from each other in the plurality of learning image data.
[0020] [Item 7] In the object recognition method according to any one of Items 2 to 6, the learning image data may be obtained by imaging the object in a state where the object occupies a predetermined range or more in the learning image.
[0021] [Item 8] In the object recognition method according to any one of Items 1 to 7, obtaining the image data is performed using an imaging device including a display, and the object recognition method may further include displaying, on the display, an auxiliary display for notifying a user of an area where the object should be located or a range that the object should occupy in the image before the image data is obtained.
[0022] [Item 9] In the object recognition method according to any one of Items 1 to 8, the plurality of filters have different wavelength dependencies of light transmittance, and the light transmittance of each of the plurality of filters may have a maximum value in a plurality of wavelength ranges.
[0023] [Item 10] The vehicle control method according to Item 10 is a vehicle control method using the object recognition method according to any one of Items 1 to 9, wherein the first imaging device is attached to a vehicle, and includes controlling the operation of the vehicle based on the result of recognizing the object.
[0024] [Item 11] The information display method according to item 11 is an information display method using an object recognition method according to any one of items 1 to 9, and based on the result of recognizing the object, at least one selected from the group consisting of the name of the object and the description of the object is obtained from a database, and the at least one selected from the group consisting of the name of the object and the description of the object is displayed on a display.
[0025] [Item 12] The object recognition method according to item 12 includes obtaining image data of an image including feature information indicating features of an object, and recognizing the object included in the image based on the feature information. The image data is obtained by repeatedly performing, a plurality of times, an operation of imaging the image in a state where a part of the plurality of light sources is emitting light, while changing a combination of light sources included in the part of the plurality of light sources, by a first imaging device including an image sensor and a light source array including a plurality of light sources that emit light in different wavelength ranges.
[0026] [Item 13] In the object recognition method according to item 12, the recognizing of the object is performed by applying a classification model learned by a machine learning algorithm to the image data, and the classification model may be pre-learned by a plurality of first training data sets each including learning image data and label data for identifying the object included in the learning image indicated by the learning image data.
[0027] [Item 14] In the object recognition method according to item 13, the plurality of learning image data included in the plurality of first training data sets may include learning image data generated by a second imaging device different from the first imaging device.
[0028] [Item 15] In the object recognition method according to item 14, the second imaging device may include a light source array having characteristics equivalent to those of the light source array in the first imaging device.
[0029] [Item 16] The object recognition method according to any one of items 13 to 15 may further include further learning the classification model by a second training data set including the image data and second label data for identifying the object after the object is recognized.
[0030] [Item 17] In the object recognition method according to any one of items 13 to 16, the positions of the objects in the learning image data included in the plurality of first training data sets within the learning images of the objects may be different from each other in the plurality of learning image data.
[0031] [Item 18] In the object recognition method according to any one of items 13 to 17, the learning image data may be obtained by imaging the object in a state where the object occupies a predetermined range or more within the learning image.
[0032] [Item 19] In the object recognition method according to any one of items 12 to 18, obtaining the image data is performed using an imaging device including a display, and the object recognition method may further include displaying, on the display, an auxiliary display for notifying a user of an area where the object should be located or a range that the object should occupy in the image before the image data is obtained.
[0033] [Item 20] The vehicle control method according to item 20 is a vehicle control method using the object recognition method according to any one of items 12 to 19, wherein the first imaging device is attached to a vehicle and controls an operation of the vehicle based on a result of recognizing the object.
[0034] [Item 21] The information display method according to Item 21 is an information display method using an object recognition method according to any one of Items 12 to 19, and based on the result of recognizing the object, at least one selected from the group consisting of the name of the object and the description of the object is obtained from a database, and the at least one selected from the group consisting of the name of the object and the description of the object is displayed on a display.
[0035] [Item 22] The object recognition device according to Item 22 includes an image sensor that generates image data of an image including feature information indicating features of an object, and a filter array disposed in an optical path of light incident on the image sensor, the filter array including a plurality of light-transmissive filters two-dimensionally arranged along a plane intersecting the optical path, the plurality of filters including two or more filters having different wavelength dependencies of light transmittance, and each of the two or more filters having a maximum value in a plurality of wavelength ranges, and a signal processing circuit that recognizes the object included in the image based on the feature information.
[0036] [Item 23] The object recognition device according to Item 23 includes an image sensor that generates an image signal of an image including an object, a light source array including a plurality of light sources that emit light in different wavelength ranges, and a control circuit that controls the image sensor and the plurality of light sources, the control circuit repeatedly performing, a plurality of times, an operation of causing the image sensor to capture an image in a state where a part of the plurality of light sources emits light while changing a combination of light sources included in the part of the plurality of light sources, and a signal processing circuit that recognizes the object included in the image based on feature information indicating features of the object included in image data constituted by the image signals generated for each of the plurality of times of imaging by the image sensor.
[0037] [Item 24] The object recognition device according to item 24 includes a memory and a signal processing circuit. The signal processing circuit receives two-dimensional image data of an image including a plurality of pixels, where information in a plurality of wavelength bands is multiplexed in the data of each of the plurality of pixels, and the luminance distribution of each of the plurality of pixels is encoded, and is multi / hyperspectral image data. Based on the feature information included in the two-dimensional image data, the signal processing circuit recognizes an object included in the scene indicated by the two-dimensional image data.
[0038] [Item 25] In the object recognition device according to item 24, the feature information may be extracted from the two-dimensional image data without reconstructing each image in the plurality of wavelength bands based on the two-dimensional image data.
[0039] [Item 26] The object recognition device according to item 24 may further include an imaging device that acquires the two-dimensional image data.
[0040] [Item 27] In the object recognition device according to item 26, the two-dimensional image data may be acquired by imaging the object in a state where the object occupies a predetermined range or more in the imaging area of the imaging device.
[0041] [Item 28] The object recognition device according to item 27 may further include a display that displays an auxiliary display for notifying the user of an area where the object should be located or a range that the object should occupy in the image captured by the imaging device before the two-dimensional image data is acquired by the imaging device.
[0042] [Item 29] In the object recognition device according to item 26, the imaging device may include an image sensor and a filter array disposed in an optical path of light incident on the image sensor, the filter array including a plurality of translucent filters two-dimensionally arranged along a plane intersecting the optical path, the plurality of filters including two or more filters having mutually different wavelength dependencies of light transmittance, and each of the two or more filters having a maximum value in a plurality of wavelength ranges.
[0043] [Item 30] In the object recognition device according to item 29, the plurality of filters may include a plurality of subsets arranged periodically.
[0044] The embodiments described below all show comprehensive or specific examples. The numerical values, shapes, materials, components, arrangement positions of the components, etc. shown in the following embodiments are merely examples and are not intended to limit the present disclosure. Also, among the components in the following embodiments, components not described in the independent claims indicating the most general concept are described as optional components.
[0045] In the present disclosure, all or part of a circuit, unit, device, member, or part, or all or part of the functional blocks in a block diagram may be executed by one or more electronic circuits including a semiconductor device, a semiconductor integrated circuit (IC), or a large scale integration (LSI). The LSI or IC may be integrated on one chip or may be configured by combining a plurality of chips. For example, functional blocks other than memory elements may be integrated on one chip. Here, although it is called an LSI or IC, the name may change depending on the degree of integration, and it may be called a system LSI, a very large scale integration (VLSI), or an ultra large scale integration (ULSI). A Field Programmable Gate Array (FPGA) programmed after the manufacture of the LSI, or a reconfigurable logic device capable of reconfiguring the bonding relationship inside the LSI or setting up the circuit partitions inside the LSI can also be used for the same purpose.
[0046] Furthermore, the functions or operations of all or part of a circuit, unit, device, member, or part can be executed by software processing. In this case, the software is recorded on a non-transitory recording medium such as one or more ROMs, optical disks, hard disk drives, etc., and when the software is executed by a processor, the functions specified by the software are executed by the processor and peripheral devices. The system or device may include one or more non-transitory recording media on which the software is recorded, a processor, and the required hardware devices, such as an interface.
[0047] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings.
[0048] (Embodiment 1) FIG. 1 is a diagram schematically showing an example of an object recognition apparatus 300 in exemplary embodiment 1 of the present disclosure. FIG. 1 shows, as an example, a situation in which a mushroom is photographed. The object 70 to be photographed may be any object. The object recognition apparatus 300 in embodiment 1 includes an imaging device 150, a signal processing circuit 200, a display 400, and a memory 500. The imaging device 150 includes an optical system 40, a filter array 100C, and an image sensor 60. The object recognition apparatus 300 may be a computer such as a smartphone or a tablet computer, for example. A camera mounted on these computers may function as the imaging device 150.
[0049] The filter array 100C is disposed in the optical path of the light incident on the image sensor 60. In the present embodiment, the filter array 100C is disposed at a position facing the image sensor 60. The filter array 100C may be disposed at other positions. The image of the light from the object 70 is encoded by the filter array 100C. Here, "encoding" means modulating the image by attenuating the light incident on the filter array 100C at an attenuation rate depending on the wavelength and position of the light. The image data generated based on the image modulated in this way is referred to as "encoded image data". Details of the configuration of the filter array 100C and encoding will be described later.
[0050] The image sensor 60 can be a monochrome type image pickup device having a plurality of light detection cells which are a plurality of pixels two-dimensionally arranged on an imaging surface. The image sensor 60 can be, for example, a CCD (Charge-Coupled Device) sensor, a CMOS (Complementary Metal Oxide Semiconductor) sensor, an infrared array sensor, a terahertz array sensor, or a millimeter wave array sensor. The light detection cells include, for example, photodiodes. The image sensor 60 does not necessarily have to be a monochrome type image pickup device. For example, a color type image pickup device having R / G / B, R / G / B / IR, or R / G / B / W filters may be used. The image sensor 60 may have detection sensitivity not only in the visible wavelength range but also in the wavelength ranges of X-rays, ultraviolet rays, near infrared rays, mid infrared rays, far infrared rays, microwaves, and radio waves.
[0051] The image sensor 60 is disposed in the optical path of the light that has passed through the filter array 100C. The image sensor 60 receives the light that has passed through the filter array 100C and generates an image signal. Each light detection cell in the image sensor 60 outputs a photoelectric conversion signal corresponding to the amount of received light. An image signal is generated by the plurality of photoelectric conversion signals output from the plurality of light detection cells. FIG. 1 schematically shows an example of the captured image 120 constituted by the image signal, that is, the encoded image data.
[0052] The optical system 40 includes at least one lens. In the example shown in FIG. 1, the optical system 40 is depicted as one lens, but it may be constituted by a combination of a plurality of lenses. The optical system 40 may have a zoom function as will be described later. The optical system 40 forms an image of the light from the object 70 on the filter array 100C.
[0053] The signal processing circuit 200 is a circuit that processes the image signal output from the image sensor 60. The signal processing circuit 200 can be realized, for example, by a combination of a central processing unit (CPU), an arithmetic processor for image processing (GPU), and a computer program. Such a computer program is stored in a recording medium such as a memory, and a processor such as a CPU or GPU can execute the recognition process described later by executing the program. The signal processing circuit 200 may be a digital signal processor (DSP) or a programmable logic device (PLD) such as a field programmable gate array (FPGA). The signal processing circuit 200 may be included in a server computer connected to a device such as the imaging device 150 or a smartphone via a network such as the Internet.
[0054] The signal processing circuit 200 recognizes the object 70 from the encoded image data. For the recognition of the object 70, for example, a model learned by a known machine learning algorithm can be used. Details of the object recognition method will be described later.
[0055] The display 400 displays information associated with the recognized object 70. The display 400 can be, for example, a display of a smartphone or a tablet computer. The display 400 may be a display connected to a personal computer or the like, or a display built into a laptop computer.
[0056] Next, the configuration and details of the encoding of the filter array 100C will be described.
[0057] FIG. 2A is a diagram schematically showing an example of a filter array 100C. The filter array 100C has a plurality of regions arranged two-dimensionally. In this specification, the region may be referred to as a "cell". A filter having a spectrally transmissivity set individually is disposed in each region. Here, the "spectrally transmissivity" means a light transmissivity having wavelength dependence. The spectrally transmissivity is represented by a function T(λ) with the wavelength of incident light being λ. The spectrally transmissivity T(λ) can take a value between 0 and 1. Thus, the filter array 100C includes a plurality of filters arranged two-dimensionally along a plane intersecting an optical path.
[0058] In the example shown in FIG. 2A, the filter array 100C has 48 rectangular regions arranged in 6 rows and 8 columns. In actual applications, more regions may be provided. The number may be on the same order as the number of pixels of a general image sensor such as an image sensor. The number of pixels is, for example, from several hundred thousand to tens of millions. In one example, the filter array 100C may be disposed directly above the image sensor and arranged such that each region corresponds to one pixel of the image sensor. Each region faces, for example, one or a plurality of pixels of the image sensor.
[0059] FIG. 2B is a diagram showing an example of the spatial distribution of the light transmissivity of each of a plurality of wavelength regions W1, W2, ···, Wi included in a target wavelength region. In the example shown in FIG. 2B, the difference in shading of each region represents the difference in transmissivity. The lighter the region, the higher the transmissivity, and the darker the region, the lower the transmissivity. As shown in FIG. 2B, the spatial distribution of the light transmissivity differs depending on the wavelength region.
[0060] Figures 2C and 2D are diagrams showing examples of the spectral transmittance of region A1 and region A2 included in a plurality of regions of the filter array 100C shown in Figure 2A, respectively. The spectral transmittance of region A1 and the spectral transmittance of region A2 are different from each other. Thus, the spectral transmittance of the filter array 100C varies depending on the region. However, it is not necessarily the case that the spectral transmittance of all regions is different. The spectral transmittance of at least two regions among the plurality of regions in the filter array 100C is different from each other. That is, the filter array 100C includes two or more filters having different spectral transmittances from each other. The spectral transmittance of each of the two or more filters has a maximum value in a plurality of wavelength ranges and a minimum value in other plurality of wavelength ranges.
[0061] Here, the meanings of "maximum value" and "minimum value" in the present disclosure will be described. When the maximum value of the spectral transmittance of the filter of interest is normalized to 1 and the minimum value is normalized to 0, a value that exceeds 0.5 and the difference from the adjacent minimum value is 0.2 or more is defined as the "maximum value" in the present disclosure. Similarly, when the above normalization is performed, a value that is less than 0.5 and the difference from the adjacent maximum value is 0.2 or more is defined as the "minimum value" in the present disclosure. The spectral transmittances of all of the plurality of filters in the filter array 100C may be different from each other. In this case, the spectral transmittance of each filter may have a maximum value in a plurality of wavelength ranges and a minimum value in other plurality of wavelength ranges. In one example, the number of patterns of the spectral transmittance of the plurality of filters included in the filter array 100C may be the same as or more than the number i of wavelength ranges included in the target wavelength range. Typically, the filter array 100C can be designed such that the spectral transmittances of more than half of the filters are different.
[0062] The filter array 100C modulates incident light into light having a plurality of intensity peaks that are discrete with respect to wavelength for each region, and superimposes and outputs these multi-wavelength lights. Thereby, the image of the light that has passed through the filter array 100C is encoded.
[0063] The wavelength resolution of the spectral transmittance in each region can be set to about the bandwidth of a desired wavelength range. In other words, among the wavelength ranges including one maximum value in the spectral transmittance curve, the width of the range taking a value equal to or greater than the average value of the minimum value closest to the maximum value and the maximum value can be set to about the bandwidth of the desired wavelength range. In this case, if the spectral transmittance is decomposed into frequency components by, for example, Fourier transform, the value of the frequency component corresponding to that wavelength range becomes relatively large.
[0064] The filter array 100C is typically divided into a plurality of cells corresponding to a plurality of regions partitioned in a grid pattern, as shown in FIG. 2A. These cells have different spectral transmittances from each other. The wavelength distribution and spatial distribution of the light transmittance in each region of the filter array 100C can be, for example, a random distribution or a quasi-random distribution.
[0065] The concepts of random distribution and quasi-random distribution are as follows. First, each region in the filter array 100C can be considered as a vector element having a value from 0 to 1, for example, according to the light transmittance. Here, when the transmittance is 0, the value of the vector element is 0, and when the transmittance is 1, the value of the vector element is 1. In other words, a set of regions arranged in a row in the row direction or column direction can be considered as a multi-dimensional vector having a value from 0 to 1. Therefore, it can be said that the filter array 100C has a plurality of multi-dimensional vectors in the column direction or row direction. At this time, the random distribution means that any two multi-dimensional vectors are independent, that is, not parallel. Also, the quasi-random distribution means that a configuration in which some of the multi-dimensional vectors are not independent is included. Therefore, in the random distribution and quasi-random distribution, a vector having as elements the transmittance values of the light in the first wavelength range in each region belonging to a set of regions arranged in one row or column included in a plurality of regions, and a vector having as elements the transmittance values of the light in the first wavelength range in each region belonging to a set of regions arranged in another row or column are independent of each other. Similarly, for a second wavelength range different from the first wavelength range, a vector having as elements the transmittance values of the light in the second wavelength range in each region belonging to a set of regions arranged in one row or column included in a plurality of regions, and a vector having as elements the transmittance values of the light in the second wavelength range in each region belonging to a set of regions arranged in another row or column are independent of each other.
[0066] When the filter array 100C is arranged near or directly above the image sensor 60, the cell pitch, which is the mutual interval between a plurality of regions in the filter array 100C, may be made substantially the same as the pixel pitch of the image sensor 60. In this way, the resolution of the image of the encoded light emitted from the filter array 100C substantially matches the resolution of the pixels. When the filter array 100C is arranged away from the image sensor 60, the cell pitch may be made finer according to the distance.
[0067] In the example shown in FIGS. 2A to 2D, a grayscale transmittance distribution is assumed in which the transmittance of each region can take any value between 0 and 1. However, it is not necessarily a grayscale transmittance distribution. For example, a binary-scale transmittance distribution in which the transmittance of each region can take either a value of approximately 0 or approximately 1 may be adopted. In the binary-scale transmittance distribution, each region transmits most of the light in at least two of the plurality of wavelength ranges included in the target wavelength range and does not transmit most of the light in the remaining wavelength ranges. Here, "most" refers to generally 80% or more.
[0068] A part of all the cells, for example, half of the cells, may be replaced with a transparent region. Such a transparent region transmits light in all the wavelength ranges W1 to Wi included in the target wavelength range with a similarly high transmittance. The high transmittance is, for example, 0.8 or more. In such a configuration, the plurality of transparent regions may be arranged, for example, in a checkered pattern. That is, in the two arrangement directions of the plurality of regions in the filter array 100C, regions with different light transmittances depending on the wavelength and the transparent regions may be alternately arranged. In the example shown in FIG. 2A, the two arrangement directions are the horizontal direction and the vertical direction. By extracting the components transmitted through the checkered transparent regions, a monochrome image can be simultaneously acquired with one camera.
[0069] The filter array 100C can be composed of at least one selected from the group consisting of a multilayer film, an organic material, a diffraction grating structure, and a microstructure containing metal. In the case of a multilayer film, for example, a multilayer film including a dielectric multilayer film or a metal film is used. At this time, in each cell, at least one of the thickness, material, and stacking order of the multilayer film can be designed to be different. Thereby, different spectral characteristics can be realized in each cell. Further, with the multilayer film, spectral characteristics having a sharp rise or fall can be realized. When an organic material is used, different spectral characteristics can be realized in each cell by different pigments or dyes, or by stacking different materials. In the case of a diffraction grating structure, different spectral characteristics can be realized in each cell by providing diffraction structures with different diffraction pitches or depths. In the case of a microstructure containing metal, different spectral characteristics can be realized by spectroscopy using the plasmon effect.
[0070] The filter array 100C is arranged near or directly above the image sensor 60. Here, "near" means being close enough so that the image of light from the optical system 40 is formed on the surface of the filter array 100C in a state where it is somewhat clear. "Directly above" means that the two are close enough that almost no gap is generated. The filter array 100C and the image sensor 60 may be integrated. The filter array 100C is a mask having a spatial distribution of light transmittance. The filter array 100C modulates and passes the intensity of the incident light.
[0071] FIG. 3A and FIG. 3B are diagrams schematically showing an example of the two-dimensional distribution of the filter array 100C.
[0072] As shown in FIG. 3A, the filter array 100C may be composed of a binary mask. The black part represents light shielding, and the white part represents light transmission. The light passing through the white part is transmitted 100%, and the light passing through the black part is shielded 100%. The two-dimensional distribution of the transmittance of the mask can be a random distribution or a quasi-random distribution. The two-dimensional distribution of the transmittance of the mask does not necessarily have to be completely random. This is because the encoding by the filter array 100C is performed to distinguish each image of each wavelength. Also, the ratio of the black part to the white part does not have to be 1:1. For example, the white part:black part = 1:9 may be acceptable. As shown in FIG. 3B, the filter array 100C may be a mask having a grayscale transmittance distribution.
[0073] As shown in FIGS. 3A and 3B, the filter array 100C has a spatial distribution of different transmittances for each of the wavelength ranges W1, W2, ···, Wi. The spatial distribution of the transmittance of each wavelength range does not match even if it is translated.
[0074] The image sensor 60 can be a monochrome type imaging device having two-dimensional pixels. However, the image sensor 60 does not necessarily have to be composed of a monochrome type imaging device. For the image sensor 60, for example, a color type imaging device having filters of R / G / B, R / G / B / IR, or R / G / B / W may be used. By using a color type imaging device, the amount of information regarding the wavelength can be increased. Thereby, it is possible to complement the characteristics of the filter array 100C, and the filter design becomes easier.
[0075] Next, the process of acquiring the image data indicating the captured image 120 by the object recognition device 300 of the present embodiment will be described. The image of the light from the object 70 is formed by the optical system 40 and encoded by the filter array 100C installed immediately before the image sensor 60. As a result, images having encoded information different for each wavelength range overlap with each other and are formed on the image sensor 60 as a multiplexed image. Thereby, the captured image 120 is obtained. At this time, since a spectroscopic element such as a prism is not used, no spatial shift of the image occurs. As a result, even a multiplexed image can maintain a high spatial resolution. As a result, it becomes possible to improve the accuracy of object recognition.
[0076] A band-pass filter may be installed in a part of the object recognition device 300 to limit the wavelength range. When the wavelength range of the object 70 is known to some extent, the discrimination range can also be limited by limiting the wavelength range. As a result, high recognition accuracy of the object can be realized.
[0077] Next, an object recognition method using the object recognition device 300 in the present embodiment will be described.
[0078] FIG. 4A is a flowchart showing an example of an object recognition method using the object recognition device 300 in the present embodiment. This object recognition method is executed by the signal processing circuit 200. The signal processing circuit 200 executes the processes of steps S101 to S104 shown in FIG. 4A by executing a computer program stored in the memory 500.
[0079] First, the user captures the object 70 with the imaging device 150 provided in the object recognition device 300. Thereby, the encoded captured image 120 is obtained.
[0080] In step S101, the signal processing circuit 200 acquires the image data generated by the imaging device 150. The image data indicates the encoded captured image 120.
[0081] In step S102, the signal processing circuit 200 performs preprocessing on the acquired image data. The preprocessing is performed to improve the recognition accuracy. The preprocessing may include, for example, region extraction, smoothing processing for noise removal, and feature extraction. The preprocessing may be omitted if it is unnecessary.
[0082] In step S103, the signal processing circuit 200 applies the learned classification model to the image data to identify the object 70 included in the scene indicated by the preprocessed image data. The classification model is pre-learned by, for example, a known machine learning algorithm. Details of the classification model will be described later.
[0083] In step S104, the signal processing circuit 200 outputs information associated with the object 70. The signal processing circuit 200 outputs information such as the name and / or detailed information of the object 70 to the display 400, for example. The display 400 displays an image indicating the information. The information is not limited to an image and may be presented by, for example, voice.
[0084] Next, the classification model used in the object recognition method will be described.
[0085] FIG. 4B is a flowchart showing an example of the generation process of the classification model.
[0086] In step S201, the signal processing circuit 200 collects a plurality of training data sets. Each of the plurality of training data sets includes learning image data and label data. The label data is information for identifying the object 70 included in the scene indicated by the learning image data. The learning image data is image data encoded in the same manner as the aforementioned image data. The plurality of learning image data included in the plurality of training data sets may include learning image data generated by the imaging device 150 in the present embodiment or other imaging devices. Details of the plurality of training data sets will be described later.
[0087] In step S202, the signal processing circuit 200 performs preprocessing on the learning image data included in each training data. The preprocessing is as described above.
[0088] In step S203, the signal processing circuit 200 generates a classification model by machine learning from a plurality of training data sets. For machine learning, for example, algorithms such as deep learning, support vector machine, decision tree, genetic programming, or Bayesian network can be used. When deep learning is used, for example, algorithms such as convolutional neural network (CNN) or recurrent neural network (RNN) can be used.
[0089] In this embodiment, by using a model trained by machine learning, information regarding an object in a scene can be directly obtained from the encoded image data. To perform the same in the prior art, many operations were required. For example, it was necessary to reconstruct the image data in each wavelength band from the encoded image data by a method such as compressive sensing, and to identify an object from those image data. In contrast, in this embodiment, it is not necessary to reconstruct the image data in each wavelength band from the encoded image data. Therefore, the time or computational resources consumed for the reconstruction process can be saved.
[0090] FIG. 4C is a diagram schematically showing an example of a plurality of training data sets in the present embodiment. In the example shown in FIG. 4C, each training data set includes encoded image data indicating one or more mushrooms and label data indicating whether the mushroom is an edible mushroom or a poisonous mushroom. In this way, for each training data set, the encoded image data and the label data indicating the correct label correspond one-to-one. The correct label can be information indicating, for example, the name, characteristics of the object 70, a sensory evaluation such as "tasty" or "tasteless", or a determination such as "good" or "bad". Generally, the larger the number of training data sets, the higher the learning accuracy can be. Here, the position of the object 70 in the images of the plurality of learning image data included in the plurality of training data sets may vary depending on the learning image data. The encoded information is different for each pixel. Therefore, the more learning image data with different positions of the object 70 in the image, the higher the accuracy of object recognition by the classification model can be improved.
[0091] In the object recognition device 300 according to the present embodiment, the classification model is incorporated in the signal processing circuit 200 before being used by the user. As another method, the encoded image data indicating the captured image 120 may be transmitted to a classification system separately prepared outside via a network or the cloud. In the classification system, for example, high-speed processing by a supercomputer is possible. Thereby, even if the processing speed of the user-side terminal is vulnerable, as long as it can be connected to the network, the recognition result of the object 70 can be provided to the user at high speed.
[0092] The image data acquired in step S101 in FIG. 4A and the learning image data acquired in step S201 in FIG. 4B can be encoded, for example, by a filter array having equivalent characteristics. In that case, the recognition accuracy of the object 70 can be increased. Here, the filter array having equivalent characteristics does not necessarily have exactly the same characteristics, and the spectral transmission characteristics may be different in some of the filters. For example, the characteristics of several percent to several tens of percent of the filters in total may be different. When the learning image data is generated by another imaging device, the other imaging device may be provided with a filter array having characteristics equivalent to those of the filter array 100C included in the imaging device 150.
[0093] The recognition result of the object 70 may be fed back to the classification model. Thereby, the classification model can be further trained.
[0094] FIG. 4D is a diagram schematically showing an example of feeding back the recognition result of the object 70 to the classification model. In the example shown in FIG. 4D, the learned classification model is applied to the pre-processed encoded image data, and a classification result is output. Then, the result is added to the data set, and further machine learning is performed using the data set. Thereby, the model can be further trained and the prediction accuracy can be improved.
[0095] FIG. 4E is a flowchart showing in more detail the operation when the recognition result is fed back to the classification model.
[0096] Steps S301 to S304 shown in FIG. 4E are the same as steps S101 to S104 shown in FIG. 4A, respectively. Thereafter, steps S305 to S307 are executed.
[0097] In step S305, the signal processing circuit 200 generates a new training data set including the image data acquired in step S301 and the label data indicating the object 70 recognized in step S303.
[0098] In step S306, the signal processing circuit 200 further trains the classification model with a new plurality of training data sets. This training process is the same as the training processes shown in steps S202 and S203 shown in FIG. 4B.
[0099] In step S307, the signal processing circuit 200 determines whether to continue recognizing the object 70. If the determination is Yes, the signal processing circuit 200 executes the process of step S301 again. If the determination is No, the signal processing circuit 200 ends the recognition of the object 70.
[0100] In this way, by feeding back the recognition result of the object 70 to the classification model, the recognition accuracy of the classification model can be improved. Furthermore, it becomes possible to create a classification model suitable for the user.
[0101] When a classification system is separately provided, the user may transmit a data set including the recognition result of the object 70 to the classification system via a network for feedback. The data set may include data indicating the captured image 120 generated by imaging, or pre-processed data thereof, and label data indicating the recognition result by the classification model or the correct label based on the user's knowledge. The user who transmits the data set for feedback may be given an incentive such as a reward or points from the provider of the classification system. Authentication of the access permission of the captured image 120 captured by the user or the possibility of automatic transmission may be displayed on the display 400, for example, by a screen pop-up, before transmission.
[0102] The filter array 100C can multiplex a plurality of wavelength information for one pixel instead of one wavelength information for one pixel. The captured image 120 includes the multiplexed two-dimensional information. The two-dimensional information is spectral information encoded, for example, randomly with respect to space and wavelength. When a fixed pattern is used as the filter array 100C, the encoding pattern is learned by machine learning. Thereby, although it is two-dimensional input data, information of substantially three dimensions (that is, two dimensions of position and one dimension of wavelength) is utilized for object recognition.
[0103] Since the image data in the present embodiment is data in which wavelength information is multiplexed, it is possible to increase the spatial resolution per wavelength as compared with a hyperspectral image that sacrifices the conventional spatial resolution. Further, the object recognition device 300 in the present embodiment can acquire image data of one frame with a single shot. Thereby, as compared with the conventional high-resolution scanning type hyperspectral imaging method, it is possible to recognize a moving object or an object that is resistant to camera shake.
[0104] In the imaging of a conventional hyperspectral image, there has been a problem that the detection sensitivity per wavelength is low. For example, when decomposing into 40 wavelengths, the amount of light decreases to 1 / 40 per pixel as compared with the case of not decomposing. On the other hand, in the method according to the present embodiment, as illustrated in FIGS. 3A and 3B, for example, about 50% of the incident light amount is detected by the image sensor 60. As a result, the detected light amount per pixel becomes higher than that of a conventional hyperspectral image. As a result, the signal-to-noise ratio of the image increases.
[0105] Next, an example of another function of the imaging device that implements the object recognition method in the present embodiment will be described.
[0106] FIG. 5A is a diagram schematically showing a function of displaying a recommended area for object recognition to assist imaging by a camera. When the object 70 is imaged extremely small or extremely large on the image sensor 60, a difference occurs between the image of the imaged object 70 and the image of the training data set recognized during learning, and the recognition accuracy decreases. The filter array 100C has different wavelength information included for each pixel, for example. For this reason, if the object 70 is detected only in a part of the imaging area of the image sensor 60, the wavelength information is biased. To prevent the bias of the wavelength information, the object 70 can be photographed as widely as possible in the imaging area of the image sensor 60. Further, when the image of the object 70 is photographed in a state where it protrudes from the imaging area of the image sensor 60, the information on the spatial resolution of the object 70 is missing. Therefore, the recommended area for object recognition is slightly inside the imaging area of the image sensor 60. In the example shown in FIG. 5A, an auxiliary display 400a indicating the recommended area for object recognition is displayed on the display 400. In FIG. 5A, the entire area of the display 400 corresponds to the imaging area of the image sensor 60. For example, an area of 60% to 98% of the horizontal or vertical width of the imaging area can be displayed on the display 400 as the recommended area for object recognition. The recommended area for object recognition may be an area of 70% to 95% or 80% to 90% of the horizontal or vertical width of the imaging area. Thus, the auxiliary display 400a may be displayed on the display 400 before the image data is acquired by the imaging device 150. The auxiliary display 400a notifies the user of the area where the object 70 should be located or the range that the object 70 should occupy in the imaged scene. Similarly, each of the plurality of learning image data included in the plurality of training data sets can be acquired by imaging the object 70 in a state where it occupies a predetermined range or more in the image.
[0107] FIG. 5B is a diagram schematically showing how an object 70 is enlarged by an optical system having a zoom function. In the example shown in the left part of FIG. 5B, the object 70 before enlargement is displayed on the display 400, and in the example shown in the right part of FIG. 5B, the object 70 after enlargement is displayed on the display 400. Thus, with the optical system 40 having a zoom function, the object 70 can be widely imaged on the image sensor 60.
[0108] FIG. 5C is a diagram schematically showing a modified example of the filter array 100C. In the example shown in FIG. 5C, a group of regions AA composed of a collection of a plurality of regions (A1, A2,...) is periodically arranged. The plurality of regions have different spectral characteristics. By "periodic" it means that the group of regions AA is repeated two or more times in the vertical direction and / or the horizontal direction while maintaining the spectral characteristics. With the filter array 100C shown in FIG. 5C, the spatial bias of wavelength information can be prevented. Further, in the learning of object recognition, instead of the entire filter array 100C shown in FIG. 5C, learning may be performed only by the group of regions AA which is a subset of the periodic structure. Thereby, the learning time can be shortened. By periodically arranging filters having the same spectral characteristics in space, object recognition becomes possible even when an object is imaged on a part rather than the entire imaging region.
[0109] The image encoded by the filter array 100C may include, for example, wavelength information multiplexed randomly. Therefore, this image is difficult for the user to view. Thus, the object recognition device 300 may separately include a normal camera for display to the user. That is, the object recognition device 300 may have a binocular configuration including the imaging device 150 and a normal camera. Thereby, a visible monochrome image that is not encoded can be displayed on the display 400 for the user. As a result, it becomes easier for the user to grasp the positional relationship between the object 70 and the imaging region of the image sensor 60.
[0110] The object recognition device 300 may have a function of extracting the contour of the object 70 in the image. By extracting the contour, unnecessary background around the object 70 can be removed. The image data with the unnecessary background removed may be used as learning image data. In that case, it is possible to further improve the recognition accuracy. The object recognition device 300 may have a function of displaying the recognition result of the contour on the display 400 so that the user can finely adjust the contour.
[0111] Figs. 6A to 6C are diagrams schematically showing application examples of the object recognition device 300 in the present embodiment.
[0112] Part (a) of Fig. 6A shows an application example for discriminating the type of a plant. Part (b) of Fig. 6A shows an application example for displaying the name of a food. Part (c) of Fig. 6A shows an application example for analyzing mineral resources. Part (d) of Fig. 6A shows an application example for specifying the type of an insect. In addition, the object recognition device 300 in the present embodiment is effective for uses such as security authentication / lock release such as face authentication, or person detection. In the case of a normal monochrome image or RGB image, there is a possibility of misrecognizing an object at a glance by the human eye. On the other hand, by adding multi-wavelength information as in the present embodiment, it becomes possible to improve the recognition accuracy of the object.
[0113] FIG. 6B shows an example in which detailed information about an object 70 is displayed on a smartphone implementing the object recognition method according to the present embodiment. In this example, the object recognition device 300 is mounted on the smartphone. By simply holding the smartphone over the object 70, it is possible to identify what the object 70 is, and based on the result, collect and display the name and description information of the object 70 from a database via a network. In this way, it is possible to utilize a portable information device such as a smartphone as an "image search encyclopedia". When complete identification is difficult, a plurality of candidates may be presented in descending order of likelihood. In this way, based on the recognition result of the object 70, data indicating the name and description information of the object 70 may be acquired from a database, and the name and / or description information may be displayed on the display 400.
[0114] FIG. 6C shows an example in which a plurality of objects existing in the street are recognized by a smartphone. The smartphone is equipped with the object recognition device 300. When the object 70 is identified as an inspection object on a production line, the inspection device acquires only information of a specific wavelength corresponding to the object 70. On the other hand, in a situation where the target of the object 70 is not specified, such as in street use, it is effective to acquire multi-wavelength information like the object recognition device 300 in the present embodiment. The object recognition device 300 may be arranged on the display 400 side of the smartphone according to the usage example, or may be arranged on the surface opposite to the display 400.
[0115] In addition, the object recognition method according to the present embodiment can be applied to a wide range of fields where recognition by artificial intelligence (AI) can be performed, such as a map application, autonomous driving, or car navigation. As described above, the object recognition device can also be mounted on a portable device such as a smartphone, a tablet, or a head-mounted display device. As long as it can be photographed by a camera, a living body such as a person, a face, or an animal can also be the object 70.
[0116] The captured image 120 represented by the image data input to the signal processing circuit 200 is a multiplexed encoded image. Therefore, it is difficult to determine at a glance what is shown in the captured image 120. However, the captured image 120 contains feature information that is information indicating the features of the object 70. Therefore, the AI can directly recognize the object 70 from the captured image 120. As a result, it is not necessary to perform the arithmetic processing of image reconstruction that takes a relatively long time.
[0117] (Embodiment 2) The object recognition device 300 according to Embodiment 2 is applied to a sensing device for autonomous driving. Hereinafter, detailed descriptions of the same contents as those in Embodiment 1 will be omitted, and the description will focus on the points different from Embodiment 1.
[0118] FIG. 7 is a diagram schematically showing an example of vehicle control using the object recognition device 300 in the present embodiment. The object recognition device 300 mounted on the vehicle can sense the environment outside the vehicle and recognize one or more objects 70 around the vehicle that enter the field of view of the object recognition device 300. The objects 70 around the vehicle may include, for example, oncoming vehicles, parallel vehicles, parked vehicles, pedestrians, bicycles, roads, lanes, white lines, sidewalks, curbs, ditches, signs, signals, utility poles, stores, trees, obstacles, or falling objects.
[0119] The object recognition device 300 includes an imaging device similar to that in Embodiment 1. The imaging device generates image data of a moving image at a predetermined frame rate. The image data shows a captured image 120 in which light from an object 70 around the vehicle passes through the filter array 100C and is multiplexed. The signal processing circuit 200 acquires the image data, extracts one or more objects 70 within the field of view from the image data, estimates what each of the extracted objects 70 is, and labels each object 70. Based on the recognition result of the object 70, the signal processing circuit 200 can, for example, understand the surrounding environment, judge danger, or display the target driving trajectory 420. Data such as the surrounding environment, danger information, and the target driving trajectory 420 can be used for controlling in-vehicle devices such as the steering or transmission of the vehicle body. Thereby, automatic driving can be enabled. The recognition result such as the object recognition label or the travel route may be displayed on a display 400 installed in the vehicle as shown in FIG. 7 so that the driver can grasp it. Thus, the vehicle control method in the present embodiment includes controlling the operation of the vehicle to which the imaging device 150 is attached based on the recognition result of the object 70.
[0120] In conventional object recognition using RGB or monochrome images, it is difficult to distinguish between a photograph and the actual object. For this reason, for example, there were cases where a photograph of a signboard or poster and the actual object were misrecognized. However, in the object recognition device 300, by using multi-wavelength information, the difference in spectral distribution between the paint of the signboard and the actual vehicle can be considered. Thereby, it is possible to improve the recognition accuracy. Further, in the object recognition device 300, two-dimensional data with multi-wavelength information superimposed is acquired. As a result, the data amount is smaller than that of conventional three-dimensional hyperspectral data. As a result, the time required for data reading and transfer, and the processing time of machine learning can be shortened.
[0121] In addition to misrecognition between a photograph and an actual object, there are cases where an object may accidentally appear as something else in a camera image. In the example shown in FIG. 7, depending on the degree of growth or viewing angle of a street tree, it may appear as a human figure. Therefore, in conventional object recognition based on shape, the street tree shown in FIG. 7 may be misrecognized as a human. In this case, in an autonomous driving environment, by misrecognizing that a person has jumped out, deceleration or sudden braking of the vehicle body may be instructed. As a result, an accident may be induced. For example, on a highway, it is unacceptable for the vehicle body to suddenly stop due to misrecognition. Even in such an environment, the object recognition device 300 can improve the recognition accuracy compared to conventional object recognition by utilizing multi-wavelength information.
[0122] The object recognition device 300 can be used in combination with various sensors such as a millimeter-wave radar, a laser range finder (Lidar), or GPS. Thereby, the recognition accuracy can be further improved. For example, by linking with information on a pre-recorded road map, the generation accuracy of the trajectory of the target driving can be improved.
[0123] (Embodiment 3) In Embodiment 3, different from Embodiment 1, by using a plurality of light sources with different emission wavelength ranges instead of the filter array 100C, encoded image data is acquired. Hereinafter, detailed descriptions of the same contents as in Embodiment 1 will be omitted, and the description will focus on the points different from Embodiment 1.
[0124] FIG. 8 is a diagram schematically showing an example of the object recognition device 300 in this embodiment. The object recognition device 300 in this embodiment includes an imaging device 150, a signal processing circuit 200, a display 400, and a memory 500. The imaging device 150 includes an optical system 40, an image sensor 60, a light source array 100L, and a control circuit 250.
[0125] The light source array 100L includes a plurality of light sources that each emit light in a different wavelength range. The control circuit 250 controls the image sensor 60 and the plurality of light sources included in the light source array 100L. The control circuit 250 repeatedly performs, a plurality of times, the operation of causing the image sensor 60 to capture an image while changing the combination of the light sources to be emitted, with some or all of the plurality of light sources in a light-emitting state. As a result, light having mutually different spectral characteristics is emitted from the light source array 100L for each imaging. The combination of the light sources to be emitted does not include exactly the same combination. However, among the plurality of combinations, some light sources may overlap in two or more combinations. Therefore, the captured images 120G1, 120G2, 120G3, ···, 120Gm obtained in each of the photographing times T1, T2, T3, ···, Tm have different intensity distributions. In the present embodiment, the image data input to the signal processing circuit 200 is a set of image signals generated by the image sensor 60 in the imaging device 150 for each of a plurality of times of imaging.
[0126] The control circuit 250 may not only change each light source to a binary state of on or off, but also adjust the light amount of each light source. Even when such adjustment is performed, a plurality of image signals having different wavelength information can be obtained. Each light source can be, for example, an LED, an LD, a laser, a fluorescent lamp, a mercury lamp, a halogen lamp, a metal halide lamp, or a xenon lamp, but is not limited thereto. Also, when emitting light in a wavelength range of the terahertz order, a light source such as an ultrafast fiber laser such as a femtosecond laser can be used.
[0127] The signal processing circuit 200 performs learning and classification of the object 70 using all or any of the captured images 120G1, 120G2, 120G3, ···, 120Gm included in the image data.
[0128] The control circuit 250 may cause the light source array 100L to emit light having not only a spatially uniform illuminance distribution but also, for example, light having a spatially random intensity distribution. The light emitted from the plurality of light sources may have different two-dimensional illuminance distributions for each wavelength. As shown in FIG. 8, the image of the light emitted from the light source array 100L toward the object 70 and passing through the optical system 40 is formed on the image sensor 60. In this case, the light incident on each pixel of the image sensor 60 or on a plurality of pixels has spectral characteristics including a plurality of different spectral peaks, similar to the example shown in FIG. 2. Thereby, as in the first embodiment, object recognition in a single shot becomes possible.
[0129] Similar to the first embodiment, the plurality of learning image data included in the plurality of training data sets includes the learning image data generated by the imaging device 150 or another imaging device. When generating the learning image data by another imaging device, the other imaging device may include a light source array having characteristics equivalent to those of the light source array 100L included in the imaging device 150. When the image data of the recognition target and each learning image data are encoded by a light source array having equivalent characteristics, high recognition accuracy of the object 70 can be obtained.
[0130] The object recognition method in the present disclosure includes acquiring image data in which a plurality of wavelength informations are multiplexed for each pixel, and applying a classification model learned by a machine learning algorithm to the image data in which the plurality of wavelength informations are multiplexed, thereby recognizing an object included in the scene indicated by the image data. Further, the object recognition method in the present disclosure includes enhancing the classification model learning using the image data in which the plurality of wavelength informations are multiplexed. The means for obtaining the image data in which the plurality of wavelength informations are multiplexed for each pixel is not limited to the imaging device described in the above-described embodiment.
[0131] The present disclosure also includes a program and a method for defining the operations executed by the signal processing circuit 200.
Industrial Applicability
[0132] The object recognition device in the present disclosure can be used in measuring instruments that accurately identify objects during measurement. The object recognition device can be applied to, for example, plant / food / biological species identification, route guidance / navigation, mineral exploration, biosensing / medical / cosmetic sensing, food foreign matter / residual pesticide inspection systems, remote sensing systems, and in-vehicle sensing systems such as autonomous driving.
Explanation of Signs
[0133] 40 Optical system 60 Image sensor 70 Object 100C Filter array 100L Light source array 120 Captured image 200 Signal processing circuit 250 Control circuit 300 Object recognition device 400 Display 400a Auxiliary display 420 Trajectory of target travel 500 Memory
Claims
1. A memory, a signal processing circuit, and comprising: The signal processing circuit acquires image data output from an image sensor, the image data being generated by light modulated to have maxima in a plurality of wavelength ranges incident on each of a plurality of pixels of the image sensor, inputs the image data into a machine learning model, the machine learning model being trained to classify an object included in a scene indicated by the image data upon receiving the input of the image data, and outputs a result of classification by the machine learning model, an object classification device.
2. The classification by the machine learning model is performed without reconstructing each image in the plurality of wavelength ranges based on the image data. The object classification device according to claim 1.
3. Further comprising an imaging device that acquires the image data. The object classification device according to claim 1.
4. The image data is acquired by imaging the object in a state where the object occupies a predetermined range or more in an imaging area of the imaging device. The object classification device according to claim 3.
5. Further comprising a display that displays an auxiliary display for notifying a user of an area where the object should be located or a range that the object should occupy in an image captured by the imaging device before the image data is acquired by the imaging device. The object classification device according to claim 4.
6. The imaging device includes the image sensor and a filter array disposed in an optical path of light incident on the image sensor, the filter array including a plurality of light-transmissive filters two-dimensionally arranged along a plane intersecting the optical path, the plurality of filters including two or more filters having mutually different wavelength dependencies of light transmittance, and each light transmittance of the two or more filters having a maximum value in a plurality of wavelength ranges, and including: The image data is generated by the light that has passed through the filter array being imaged by the image sensor. The object classification device according to claim 3.
7. The plurality of filters include a plurality of subsets arranged periodically. The object classification device according to claim 6.
8. A method executed by a computer, Obtain image data output from an image sensor, the image data being generated by light modulated to have maxima in a plurality of wavelength bands incident on each of a plurality of pixels of the image sensor. Input the image data into a machine learning model, the machine learning model being trained to receive the input of the image data and classify an object included in a scene indicated by the image data. Output a result of classification by the machine learning model. Object classification method. **Claim 9** A step of obtaining image data output from an image sensor, the image data being generated by light modulated to have maxima in a plurality of wavelength bands incident on each of a plurality of pixels of the image sensor. A step of inputting the image data into a machine learning model, the machine learning model being trained to receive the input of the image data and classify an object included in a scene indicated by the image data. A step of outputting a result of classification by the machine learning model. Causing a computer to execute the steps. Program.
Citation Information
Patent Citations
A spectral camera equipped with a mosaic filter for each pixel.
JP2015501432A
Imaging device and spectroscopy system
JP2016156801A
Data processor, data processing method, program, and electronic equipment
JP2018096834A
Systems and methods for analyzing remote sensing imagery
US20170076438A1
Multi-point spectral system
US20170163901A1