Imaging device, imaging method, and recording medium

WO2026204642A1PCT designated stage Publication Date: 2026-10-01SONY SEMICON SOLUTIONS CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2026/010636
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-28
Filing Date
2026-03-18
Publication Date
2026-10-01

Smart Images

  • Figure JP2026010636_01102026_PF_FP_ABST
    Figure JP2026010636_01102026_PF_FP_ABST
Patent Text Reader

Abstract

Provided is an imaging device that outputs information relating to a feature region required by a post-stage processing unit to the post-stage processing unit as information relating to a bit depth required by the post-stage processing unit. A pre-processing unit comprises: a feature region detection unit that detects a feature region of a target object for image analysis by the post-stage processing unit; a data format determination unit that determines a data format on the basis of a data amount of an image of the feature region and a post-stage bit depth required by the post-stage processing unit; and a virtual channel number determination unit that dynamically determines the number of virtual channels on the basis of the data amount of the image of the feature region and a post-stage frame rate of the post-stage processing unit. An output interface unit converts the format of the image of the feature region according to the data format determined by the data format determination unit, dynamically notifies the post-stage processing unit of the number of virtual channels determined by the virtual channel number determination unit, and outputs the image of the feature region to the post-stage processing unit.
Need to check novelty before this filing date? Find Prior Art

Description

Image Pickup Apparatus, Image Pickup Method, and Recording Medium

[0001] The present disclosure relates to an image pickup apparatus, an image pickup method, and a recording medium.

[0002] For example, as disclosed in Patent Document 1 etc., sensing is performed outside the image sensor, a necessary image region is notified from outside the image sensor to the image sensor, and by utilizing the feature region cropping function of the image sensor and MIPI (Mobile Industry Processor Interface), data with a narrowed range and a reduced number of pixels is output. Accordingly, a signal acquisition apparatus that reduces the burden of post-processing outside the image sensor and realizes high-speed processing is known.

[0003] Japanese Patent Laid-Open No. 2021-140032

[0004] Incidentally, although the cropped image of the feature region can reduce the amount of data transmitted to the post-processing unit, it is generally compressed data, and part of bit information is missing, which may make it impossible to perform high-precision image analysis in the post-processing unit in some cases.

[0005] According to one aspect of the present disclosure, the post-processing unit outputs information of a necessary feature region to the post-processing unit as bit depth information required by the post-processing unit.

[0006] An imaging device relating to one aspect of the present disclosure includes an image sensor, a pre-processing unit that performs pre-processing necessary for image analysis of an image captured by the image sensor, and an output interface unit that outputs pre-processing information to the subsequent processing unit, wherein the pre-processing unit includes a feature region detection unit that detects a feature region of an object to be analyzed by the subsequent processing unit, a data format determination unit that determines a data format based on the amount of data of the image of the feature region and the subsequent bit depth required by the subsequent processing unit, and a virtual channel number determination unit that dynamically determines the number of virtual channels based on the amount of data of the image of the feature region and the subsequent frame rate of the subsequent processing unit, wherein the output interface unit formats the image of the feature region according to the data format determined by the data format determination unit, dynamically notifies the subsequent processing unit of the number of virtual channels determined by the virtual channel number determination unit, and outputs the image of the feature region to the subsequent processing unit.

[0007] This is a block diagram illustrating the overview configuration of an imaging system including an imaging device according to an embodiment. This is a flowchart illustrating the preprocessing procedure by the preprocessing unit. This is a diagram illustrating the overview configuration of a robot picking device using the imaging device according to an embodiment of this disclosure. This is a diagram illustrating the generation of a distance image by the metadata generation unit. This is a diagram illustrating the relationship between a feature image generated from an image and metadata. This is a diagram illustrating the virtual channel of MIPI. This is a block diagram illustrating the overview configuration of an imaging system including an imaging device according to a first modified example. This is a flowchart illustrating the preprocessing procedure by the preprocessing unit according to the first modified example. This is a block diagram illustrating the overview configuration of an imaging system including an imaging device according to a second modified example. This is a diagram illustrating the first application example. This is a diagram illustrating the generation of metadata in the first application example. This is a diagram illustrating the second application example. This is a diagram illustrating the third application example. This is a diagram illustrating the fourth application example. This is a diagram illustrating the fifth application example. This is a hardware configuration diagram showing an example of a computer that realizes the functions of the preprocessing unit.

[0008] Embodiments of this disclosure will be described in detail below with reference to the drawings. In each of the following embodiments, the same parts will be denoted by the same reference numerals to avoid redundant descriptions.

[0009] This disclosure will be described in the following order of items: 1. Embodiments 1-1. Overview of the imaging device according to the embodiment 1-2. Preprocessing procedure 1-3. Application to a robot picking device 1-4. Example of metadata transmission 1-5. Number of virtual channels 2. First modification 2-1. Configuration 2-2. Preprocessing procedure 3. Second modification 4. First application example 5. Second application example 6. Third application example 7. Fourth application example 8. Fifth application example 9. Hardware configuration 10. Summary

[0010] (1. Embodiments) (1-1. Overview of the Imaging Device According to an Embodiment) Figure 1 is a block diagram showing the overview configuration of an imaging system 1 including an imaging device 10 according to an embodiment. As shown in Figure 1, the imaging system 1 includes an imaging device 10 and a post-processing unit 30. The imaging device 10 is also called a solid-state imaging device or image sensor, and includes an image sensor 11, a high dynamic range image generation unit 12, an image processing unit 13, and an output interface unit 27. The image processing unit 13 also includes a pre-processing unit 20.

[0011] The downstream processing unit 30 is connected to the downstream of the imaging device 10 and performs image analysis. The image processing unit 13 performs image processing on the image captured by the image sensor 11, and in particular the preprocessing unit 20 performs the preprocessing necessary for the image analysis of the downstream processing unit 30, and outputs the preprocessing information that has been processed to the downstream processing unit 30 via the output interface unit 27.

[0012] The high dynamic range image generation unit 12 synthesizes multiple images captured by the image sensor 11 while changing the exposure (exposure period) to generate a dynamic range image (HDR (High Dynamic Range) image) with a wide dynamic range, and outputs an image with less overexposure and underexposure to the preprocessing unit 20. Alternatively, the high dynamic range image generation unit 12 may be omitted, and the standard dynamic range image (SDR (Standard Dynamic Range) image) captured by the image sensor 11 may be output to the preprocessing unit 20.

[0013] The preprocessing unit 20 includes a feature region detection unit 21, a metadata generation unit 22, a data format determination unit 23, and a virtual channel number determination unit 24. The data handled by the preprocessing unit 20, specifically the subsequent bit depth 25 and the subsequent frame rate 26, are shown with reference numerals in the figure.

[0014] The feature region detection unit 21 detects the feature region (ROI: Region of Interest) of the object targeted for image analysis by the subsequent processing unit 30 within the image (HDR image) captured by the image sensor 11. This feature region is detected, for example, using a trained model. The trained model is machine-trained using training data that includes the target object. The trained model may be a multilayer neural network, for example, a deep neural network (DNN), or more specifically, a convolutional neural network (CNN).

[0015] The metadata generation unit 22 generates metadata using the image of the feature region detected by the feature region detection unit 21. The metadata is, for example, distance information of the detected feature region.

[0016] The data format determination unit 23 determines the data format based on the amount of data for the feature region image and metadata, and the subsequent bit depth 25 required by the subsequent processing unit 30. The data format is a combination of a data type such as RAW, RGB, or YUV and a bit depth, such as RAW24 (bit depth: 24 bits), RGB888 (bit depth: 24 bits), or YUV422 (bit depth: 12 bits). By using a data format with a high bit depth, the subsequent processing unit 30 can perform high-precision image analysis. In particular, converting the data to a data type that the subsequent processing unit 30 can handle reduces the processing load on the subsequent processing unit 30.

[0017] The virtual channel number determination unit 24 dynamically determines the number of virtual channels in the output interface unit 27 based on the amount of image and metadata data in the feature region and the subsequent frame rate 26 of the subsequent processing unit 30. The clocks of unused virtual channels are stopped to reduce power consumption. This power consumption reduction allows the imaging device 10 to satisfy power consumption limitations, especially when the imaging device 10 has a stacked die structure.

[0018] The output interface unit 27 is, for example, MIPI as a high-speed parallel interface. MIPI can also use multiple virtual channels. The output interface unit 27 converts the image and metadata of the feature region according to the data format determined by the data format determination unit 23, dynamically notifies the subsequent processing unit 30 of the number of virtual channels determined by the number of virtual channels 24, and outputs the image and metadata of the feature region to the subsequent processing unit.

[0019] The bit depth is the number of bits assigned to one pixel, and the subsequent bit depth 25 is the bit depth used by the subsequent processing unit 30 for image analysis. By the preprocessing unit 20 matching the data format to the bit depth of the subsequent processing unit 30, the load on the subsequent processing unit 30 is reduced. The subsequent frame rate 26 is the frame rate for data processing in the subsequent processing unit 30.

[0020] When the subsequent frame rate 26 is 60 fps and the subsequent processing unit 30 prioritizes accuracy, the subsequent bit depth 25 increases and the number of virtual channels decreases. On the other hand, when the subsequent frame rate 26 is 120 fps and the subsequent processing unit 30 prioritizes processing speed, the subsequent bit depth 25 decreases and the number of virtual channels increases. The virtual channel number determination unit 24 dynamically changes the number of virtual channels to match the subsequent frame rate 26 of the subsequent processing unit 30 in response to changes in the data volume of the image and metadata of the feature region, changes in parameters, and changes in the data format (change in processing mode). Here, a change in processing mode refers to a change in data format such as a data volume reduction mode (RAW6), a subsequent load reduction mode (RGB444), or a high-precision mode (RAW24).

[0021] The preprocessor unit 20 functions as a control unit and stores programs corresponding to each of the functional units of the feature region detection unit 21, metadata generation unit 22, data format determination unit 23, and virtual channel number determination unit 24 in a storage device such as a non-volatile memory or magnetic disk drive. These programs are loaded into memory and executed by the CPU to run the corresponding processes. The subsequent bit depth 25 and subsequent frame rate 26 are stored in memory as parameters. The feature region detection unit 21, metadata generation unit 22, data format determination unit 23, and virtual channel number determination unit 24 may be stored as firmware or as signal processing hardware such as an FPGA.

[0022] The subsequent processing unit 30 has a control unit 31, which has an image analysis unit 32 and a result output unit 33. The image analysis unit 32 performs predetermined image analysis processing using preprocessing information output from the preprocessing unit 20. The result output unit 33 then outputs the analysis results. The predetermined image analysis processing is a three-dimensional shape analysis of the object to be picked used in robot picking.

[0023] (1-2. Preprocessing Procedure) Figure 2 is a flowchart showing the preprocessing procedure performed by the preprocessing unit 20. As shown in Figure 2, the preprocessing unit 20 first receives the captured image output from the high dynamic range image generation unit 12 (step S101). Then, the feature region detection unit 21 detects the feature region of the target object from the captured image (step S102). Furthermore, the metadata generation unit 22 generates metadata for the feature region (step S103). After that, the data format determination unit 23 determines the data format based on the amount of data for the feature region image and metadata and the subsequent bit depth 25 required by the subsequent processing unit 30 (step S104). Furthermore, the virtual channel number determination unit 24 determines the number of virtual channels based on the amount of data for the feature region image and metadata and the subsequent frame rate 26 of the subsequent processing unit 30 (step S105). Subsequently, the output interface unit 27 converts the image and metadata of the feature region according to the data format determined by the data format determination unit 23, dynamically notifies the subsequent processing unit 30 of the number of virtual channels determined by the virtual channel number determination unit 24, and outputs the image and metadata of the feature region to the subsequent processing unit 30 (step S106), thus ending this process. This process is performed each time an captured image is input.

[0024] (1-3. Application to a robot picking device) Figure 3 is a diagram showing the schematic configuration of a robot picking device using an imaging device according to an embodiment of the present disclosure. As shown in Figure 3, the robot arm 41 is a device that performs tasks such as grasping and carrying objects in place of a human arm. The robot arm 41 is equipped with a plurality of links, each of which is rotatably mounted. A gripping part 42 for grasping the object 43 to be picked is positioned at the tip of the robot arm 41. The robot arm 41 extends and retracts and rotates based on the control of a control device (not shown), and can carry the grasped object 43 to any position within its range of motion.

[0025] Figure 3 shows an example in which a robot arm 41 grasps an object 43 being transported by a belt conveyor 44 and transports it to a predetermined location such as a workpiece. The robot arm 41 keeps its tip end in a predetermined position near the belt conveyor 44, and when the object 43 moves within its range of motion, it moves its tip end and grasps the object 43 with its gripping part 42.

[0026] An imaging device 10 and a downstream processing unit 30 are positioned at the tip of the robot arm 41. The imaging device 10 captures a moving image of the area near the belt conveyor 44 within the movable range of the robot arm 41. The downstream processing unit 30 recognizes the moving target object 43 from the moving image captured by the imaging device 10, detects whether the target object 43 has entered an area where it can be grasped by the gripping unit 42 of the robot arm 41, and transmits this information to the control unit of the robot arm 41. As a result, the control unit of the robot arm 41 moves the robot arm 41 and drives the gripping unit 42 to grasp the target object 43.

[0027] Here, the feature region detection unit 21 of the imaging device 10 detects the feature region of the target object 43 within the image of the measurement range. The metadata generation unit 22 generates a distance image as metadata based on the image within the feature region. This distance image is output to the downstream processing unit 30. The downstream processing unit 30 generates 3D point cloud data based on the distance image. Based on the 3D point cloud data, the downstream processing unit 30 detects the position and orientation of the target object 43 and notifies the control unit. As a result, the gripping unit 42 can pick the target object 43 with high accuracy.

[0028] Figure 4 illustrates the generation of a distance image by the metadata generation unit 22. The metadata generation unit 22 calculates the distance of the target object 43. The distance of the target object 43 can be calculated based on the principle of triangulation by detecting feature points from the pattern light projected onto the target object 43, and associating the detected feature points with feature points on the captured image.

[0029] Figure 4(A) shows an example in which a pattern light 51 from a projector 50 is projected onto a target object 43, and an image is generated by an imaging device 10. The pattern light 51 represents an example of a stripe code composed of multiple vertical lines. The projector 50 projects such a pattern light 51. The imaging device 10, located at a predetermined distance from the projector 50, images the target object 43 onto which the pattern light 51 is projected, and obtains an image corresponding to the distance to the target object 43. The image captured when there is no target object 43 is called the projected image. This projected image is an image based on the pattern light 51.

[0030] Figure 4(B) shows an example of an captured image 60. In the captured image 60, an image 62 is formed in the region of the target object 43 that changes according to the shape of the target object 43, and an image based on pattern light 51 (projected image 61) is formed in the other regions. The metadata generation unit 22 sets feature points in the image 62. These feature points can be set on the edges of the projection pattern, etc. The black dots in Figure 4(B) represent examples of feature points. Next, the metadata generation unit 22 calculates feature quantities from the pixel values ​​around the set feature points. Next, the metadata generation unit 22 associates the set feature points with pixels of the projected image 61 that have the same feature quantities. Figure 4(C) shows an example of when feature points (x, y) in the image 62 and points (xp, yp) in the projected image 61 are associated.

[0031] Next, the metadata generation unit 22 calculates the distance between the feature point (x, y) and the point (xp, yp) on the image. Then, the metadata generation unit 22 can calculate the distance (3D shape) to the target object 43 from the distance between the feature point (x, y) and the point (xp, yp) on the image by applying the principle of triangulation. By repeating this process, a distance image can be generated. In other words, the subsequent processing unit performs image analysis of the 3D measurement of the target object, the image of the feature region is a high-bit-depth image of the target object, and the metadata is a distance image calculated based on the coordinate position data of the pattern light irradiated onto the target object.

[0032] (1-4. Example of Metadata Transmission) Figure 5 shows the relationship between the feature image D2 (ROI image) and metadata D3 (distance image) generated from the captured image D1. The feature image D2 is an image of the feature region E within the captured image D1 that contains the target object 43. The metadata D3 is a distance image generated based on the feature image D2 onto which pattern light is projected.

[0033] As shown in Figure 5, if the captured image D1 is, for example, 4000 x 3000 pixels, the feature image D2 is an image of only the feature region E, and therefore will be, for example, 512 x 512 pixels. In other words, the image transmitted is not the captured image D1, but only the feature image D2, which is narrowed down to the feature region E, thus reducing the amount of data transmitted and enabling high-speed output to the subsequent processing unit 30. Also, the metadata D3 will be, for example, 512 x 512 pixels. Therefore, even if metadata D3 is included in the transmission of the feature image D2, the amount of data transmitted is reduced. Moreover, since the subsequent processing unit 30 does not need to generate a distance image, the processing load is reduced.

[0034] Furthermore, feature image D2 is masked except for region EA, which corresponds to feature region E within the frame region FD of captured image D1, and the pixel data of feature region EA is transmitted as feature image D2 during the horizontal period. Therefore, region EB other than region EA becomes blank space, and pixel data is not transmitted. Consequently, metadata D3 can be sent during the horizontal period (horizontal padding period) corresponding to this blank region EA. In other words, metadata D3 can be transmitted by effectively utilizing the horizontal padding period of the frame.

[0035] (1-5. Number of Virtual Channels) Figure 6 is a diagram illustrating MIPI virtual channels. As shown in Figure 6, the output interface unit 27 can transmit data through multiple virtual channels VC1 and VC2. Although Figure 6 shows two virtual channels VC1 and VC2, the number of virtual channels can be increased or decreased up to, for example, 16 channels.

[0036] MIPI's virtual channels VC1 and VC2 are called lane distribution functions, and each virtual channel VC1 and VC2 transfers data in a predetermined data sequence. All data handled is in 8-bit units. The virtual channel of the sent data is notified to the downstream processing unit 30 by the data type ID (virtual channel number), allowing the downstream processing unit 30 to retrieve the data.

[0037] Virtual channel VC1 is the master channel, and virtual channel VC2 is the slave channel. Each virtual channel VC1 and VC2 is clock-controlled by the timing generator 70. The parallel-to-serial conversion unit 71 converts the parallel data to serial data and outputs it. When changing from two virtual channels to one, virtual channel VC1 must always be active, but the other slave virtual channel, for example, virtual channel VC2, can be turned off by the timing generator 70. In this case, the clock of virtual channel VC2 is turned off. Therefore, the power consumption of the unused virtual channel VC2 can be reduced.

[0038] (2. First Modification) (2-1. Configuration) Figure 7 is a block diagram showing the outline configuration of an imaging system 1 including an imaging device 10 according to the first modification. In the above embodiment, the preprocessing unit 20 incorporated the programs for the feature region detection unit 21 and the metadata generation unit 22, as well as the data format determination unit 23 and the virtual channel number determination unit 24, and parameters including the subsequent bit depth 25 and the subsequent frame rate 26. In contrast, the preprocessing unit 20 of the first modification is equipped with a general-purpose CPU and, through a connection with the subsequent processing unit 30, writes the programs for the feature region detection unit 21 and the metadata generation unit 22, as well as the data format determination unit 23 and the virtual channel number determination unit 24, and parameters including the subsequent bit depth 25 and the subsequent frame rate 26 from the subsequent processing unit 30 to the preprocessing unit 20.

[0039] Specifically, the pre-processing unit 20 is provided with an input / output interface unit 28, and the post-processing unit 30 is provided with an input / output interface unit 34. When the post-processing unit 30 is connected to the imaging device 10, the connection is made via the input / output interface units 28 and 34, and the programs for the feature region detection unit 21 and metadata generation unit 22 corresponding to the post-processing unit 30 are written to the pre-processing unit 20 as firmware. As a result, an imaging device 10 is created that has a feature region detection unit 21 and metadata generation unit 22 corresponding to the processing content of the post-processing unit 30 connected to the imaging device 10.

[0040] (2-2. Preprocessing Procedure) Figure 8 is a flowchart showing the preprocessing procedure by the preprocessing unit 20 of the first modified example. As shown in Figure 8, first, the preprocessing unit 20 writes the programs of the feature region detection unit 21 and metadata generation unit 22 corresponding to the subsequent processing unit 30 as firmware to the preprocessing unit 20, and also writes parameters including the subsequent bit depth 25 and subsequent frame rate 26 to the preprocessing unit 20 (step S100). After that, the processing of steps S101 to S106 is performed, which is the same as the processing of steps S101 to S106 shown in Figure 2. After the processing of step S106, the preprocessing unit 20 determines whether or not imaging by the imaging device 10 has finished (step S107). If imaging by the imaging device 10 has not finished (step S107: No), it proceeds to step S101 and repeats the processing described above. On the other hand, if imaging by the imaging device 10 has finished (step S107: Yes), the preprocessing unit 20 terminates this process.

[0041] (3. Second Modification) Figure 9 is a block diagram showing the outline configuration of the imaging system 1 including the imaging device 10 according to the second modification. In the second modification, the metadata generation unit 22 in the preprocessing unit 20 of the imaging device 10 of the embodiment is removed. In the second modification, only the image of the feature region detected by the feature region detection unit 21 is output to the subsequent processing unit 30 as an uncompressed high-bit-depth image.

[0042] (4. First Application Example) FIG. 10 is a diagram illustrating the first application example. FIG. 11 is a diagram illustrating generation of metadata in the first application example. In the first application example, the configuration of the image capturing apparatus 10 according to the embodiment is used. The post-processing unit 30 performs image analysis for three-dimensional measurement of a target object, and the feature region detection unit 21 of the embodiment detects a feature region of the target object. The image of the feature region is a high bit-depth image for the target object, and the metadata is width-direction center position coordinate data of slit light with which the target object is irradiated.

[0043] In the first application example, a target object 80 is irradiated with slit light 81. In the ROI image of the feature image D2 shown in FIG. 10, the target object 80 is irradiated with the slit light 81. The target object 80 is conveyed in a direction orthogonal to the slit light 81. Accordingly, as the target object 80 moves, the slit light 81 is substantially the same as irradiating patterned light. Meanwhile, as shown in FIG. 10, metadata D4 is the width-direction center position coordinate data of the slit light 81 with which the target object 80 is irradiated.

[0044] The post-processing unit 30 generates distance information of the target object 80 by using a plurality of pieces of metadata D4 associated with the relative movement of the target object 80 with respect to the slit light 81, and further generates three-dimensional data.

[0045] Here, the reason why the ROI image is set to have a high bit-depth is that, as shown in FIG. 11, a luminance distribution L of the slit light 81 in a moving direction of the target object 80 is a Gaussian distribution. Therefore, the center position in the width direction of the luminance distribution L is the maximum value in the width direction. To obtain the coordinates (xc, yc) of this maximum value with high accuracy, the ROI image is provided with a high bit-depth. Note that the ROI image does not need to be color pixels, and monochrome pixels are sufficient.

[0046] Since the preprocessing unit 20 performs acquisition processing of the width-direction center position coordinate data of the slit light 81 as metadata, the load on the post-processing unit 30 can be reduced. In addition, since metadata is generated using a high bit-depth image in the image capturing apparatus 10, highly accurate three-dimensional data can be obtained.

[0047] (5. Second Application Example) Figure 12 is a diagram illustrating the second application example. In the second application example, the configuration of the imaging device 10 according to the second modification is used. The post-processing unit 30 performs image analysis for graphic code reading, and the characteristic region detection unit 21 detects a characteristic region of the graphic code. The image of the characteristic region is an image with a high bit depth for the graphic code, and the output interface unit 27 outputs the image of the characteristic region to the post-processing unit 30 in an uncompressed format.

[0048] As shown in Figure 12, the characteristic region detection unit 21 detects a characteristic region E of the graphic code 90. A characteristic image D2 of the characteristic region E is an image having a high bit depth for the graphic code 90, for example, RAW24 (bit depth: 24 bits). The output interface unit 27 outputs the characteristic image D2 to the post-processing unit 30 in an uncompressed format. This allows the post-processing unit 30 to identify the graphic code 90 with high accuracy.

[0049] (6. Third Application Example) Figure 13 is a diagram illustrating the third application example. In the third application example, the configuration of the imaging device 10 according to the embodiment is used. The preprocessing unit 20 alternately outputs images of the characteristic region in different data formats temporally.

[0050] In the third application example, the post-processing unit 30 performs tablet inspection, and this inspection includes performing scratch inspection based on luminance and material abnormality inspection based on color almost simultaneously. For this reason, the characteristic region detection unit 21 detects an ROI image in which the tablet 95 is present. Then, for this ROI image, a characteristic image D21 format-converted to YUV420 (bit depth: 10 bits) for scratch inspection and a characteristic image D22 format-converted to RGB444 (bit depth: 12 bits) for color inspection are generated. Then, the characteristic image D21 and the characteristic image D22 are output alternately temporally. Note that when the data amount of the characteristic image D21 and the characteristic image D22 is not large, the characteristic image D21 and the characteristic image D22 may be transmitted as one frame. In this way, by alternately transmitting a plurality of characteristic images that have undergone format conversion suitable for a plurality of inspections, the load on the post-processing unit 30 is reduced, and a plurality of inspections can be performed simultaneously.

[0051] (7. Fourth Application Example) Figure 14 is a diagram illustrating the fourth application example. In the fourth application example, the configuration of the imaging device 10 of the embodiment is used. The preprocessing unit 20 outputs metadata of the feature region image and the depth image alternately in time, and the postprocessing unit 30 combines the feature region image and the depth image and displays them in a segmented view.

[0052] As shown in Figure 14, the preprocessing unit 20 transmits the feature image and the distance image to the subsequent processing unit 30. As shown in Figure 14, the subsequent processing unit 30 divides the feature region E of the target object 100 into divided region E1 and divided region E2, displays a portion of the feature image D23 (ROI image) in divided region E1, and displays a portion of the distance image D33 in divided region E2. The proportion and position of divided regions E1 and E2 relative to the feature region E can be set arbitrarily.

[0053] (8. Fifth Application Example) Figure 15 is a diagram illustrating the fifth application example. In the fifth application example, the configuration of the imaging device 10 of the second modification is used. The downstream processing unit 30 performs image analysis for tilt correction of the target object. The feature region detection unit 21 detects the feature region of the target object. The image of the feature region is a high-bit-depth image of the target object. The output interface unit 27 outputs the image of the feature region to the downstream processing unit 30 uncompressed.

[0054] As shown in Figure 15, the preprocessing unit 20 transmits a feature image D24 in which the target object 110 is positioned diagonally to the subsequent processing unit 30 at a high bit depth. As shown in Figure 15, the subsequent processing unit 30 performs tilt correction on the feature image D24 and displays the tilt-corrected feature image D100. Note that the correction angle increases, so the correction distance increases. In this case, it is preferable to use the imaging device 10 of the embodiment and transmit distance information at a high bit depth to perform tilt correction. Note that the ROI area decreases as the correction angle increases, so high speed can be ensured even when transmitting feature information at a high bit depth.

[0055] The feature region detection unit 21 detects the feature region E of the target object 110 not by using a pre-trained model, but by image processing including edge detection and centroid detection.

[0056] (9. Hardware Configuration) The preprocessing unit 20 according to the embodiments and modifications of the present disclosure described above is implemented by a computer 1000 having a configuration such as that shown in Figure 16. Figure 16 is a hardware configuration diagram showing an example of a computer 1000 that implements the functions of the preprocessing unit 20. The computer 1000 has a CPU 1100, RAM 1200, ROM 1300, secondary storage device 1400, communication interface 1500, and input / output interface 1600. The parts of the computer 1000 are connected by a bus 1050.

[0057] The CPU 1100 operates based on programs stored in the ROM 1300 or secondary storage device 1400, and controls each part. For example, the CPU 1100 loads the programs stored in the ROM 1300 or secondary storage device 1400 into the RAM 1200 and executes processing corresponding to various programs.

[0058] ROM 1300 stores boot programs such as the BIOS (Basic Input Output System) that are executed by the CPU 1100 when the computer 1000 starts up, as well as programs that depend on the computer 1000's hardware.

[0059] The secondary storage device 1400 is a computer-readable recording medium that non-temporarily stores programs executed by the CPU 1100 and data used by such programs. Specifically, the secondary storage device 1400 is a recording medium that stores programs according to this embodiment and its modifications.

[0060] The communication interface 1500 is an interface for the computer 1000 to connect to the external network 1550. For example, the CPU 1100 can receive data from other devices or transmit data it has generated to other devices via the communication interface 1500.

[0061] The input / output interface 1600 is an interface for connecting the input / output device 1650 and the computer 1000. For example, the CPU 1100 receives data from input devices such as a keyboard or mouse via the input / output interface 1600. The CPU 1100 also transmits data to output devices such as a display, speaker, or printer via the input / output interface 1600. The input / output interface 1600 may also function as a media interface for reading programs recorded on a predetermined recording medium (media). Examples of media include optical recording media such as DVDs (Digital Versatile Discs) and PDs (Phase Change Rewritable Disks), magneto-optical recording media such as MOs (Magneto-Optical Disks), tape media, magnetic recording media, or semiconductor memory.

[0062] For example, when computer 1000 functions as a preprocessing unit 20, the CPU 1100 of computer 1000 executes a program loaded onto RAM 1200 to realize functions such as a feature area detection unit 21, a metadata generation unit 22, a data format determination unit 23, and a virtual channel number determination unit 24. The secondary storage device 1400 stores the program according to this disclosure, as well as data such as the subsequent bit depth 25 and the subsequent frame rate 26. The CPU 1100 reads and executes the program data 1450 from the secondary storage device 1400, but as another example, these programs may be obtained from other devices via an external network 1550.

[0063] (10. Summary) As described above, the imaging device according to this disclosure performs feature region detection within the imaging device, enabling uncompressed calculations even with high bit depth images such as high dynamic range images. Furthermore, by performing feature region detection and metadata generation within the imaging device, the processing load on the subsequent processing unit is reduced, and power consumption in the subsequent stage can be reduced. Moreover, by coordinating the imaging device with feature region detection functionality with the subsequent processing unit, or by setting parameters, and utilizing output interfaces such as MIPI, it is possible to convert high bit depth feature information and high-precision metadata of only the necessary regions into an optimal format and output with the minimum number of virtual channels. This enables high-speed output with optimized power consumption. At this time, since the imaging data has a high bit depth, it is possible to output without reducing the frame rate even when more advanced image analysis is required.

[0064] The effects described in this disclosure are merely illustrative and not limited to those disclosed. Other effects may also occur.

[0065] While embodiments of this disclosure have been described above, the technical scope of this disclosure is not limited to the embodiments described above, and various modifications are possible without departing from the spirit of this disclosure. Furthermore, components from different embodiments and modifications may be combined as appropriate.

[0066] Furthermore, this technology can also take the following configuration: (1) An imaging device comprising: an image sensor; a pre-processing unit that performs pre-processing necessary for image analysis of an image captured by the image sensor and is connected to a subsequent processing unit that performs image analysis; and an output interface unit that outputs the pre-processing information to the subsequent processing unit, wherein the pre-processing unit comprises: a feature region detection unit that detects the feature region of an object to be analyzed by the subsequent processing unit; a data format determination unit that determines the data format based on the amount of data of the image of the feature region and the subsequent bit depth required by the subsequent processing unit; and a virtual channel number determination unit that dynamically determines the number of virtual channels based on the amount of data of the image of the feature region and the subsequent frame rate of the subsequent processing unit, wherein the output interface unit converts the image of the feature region in the data format determined by the data format determination unit, dynamically notifies the subsequent processing unit of the number of virtual channels determined by the virtual channel number determination unit, and outputs the image of the feature region to the subsequent processing unit. (2) The imaging apparatus according to (1), wherein the preprocessing unit comprises a metadata generation unit that generates metadata for the feature region, the data format determination unit determines the data format based on the amount of data for the image of the feature region and the metadata and the subsequent bit depth required by the subsequent processing unit, the virtual channel count determination unit dynamically determines the number of virtual channels based on the amount of data for the image of the feature region and the metadata and the subsequent frame rate of the subsequent processing unit, and the output interface unit converts the image of the feature region and the metadata in the format determined by the data format determination unit, dynamically notifies the subsequent processing unit of the number of virtual channels determined by the virtual channel count determination unit, and outputs the image of the feature region and the metadata to the subsequent processing unit.(3) The imaging apparatus according to (1) or (2), comprising a high dynamic range image generation unit that generates a high dynamic range image based on an image captured by the image sensor, wherein the preprocessing unit performs preprocessing on the high bit depth image output from the high dynamic range image generation unit. (4) The imaging apparatus according to (2) or (3), wherein the preprocessing unit is equipped with a general-purpose CPU and, through connection with the downstream processing unit, has been programmed by the downstream processing unit to the preprocessing unit to function as the feature region detection unit and the metadata generation unit, and parameters including the downstream bit depth and the downstream frame rate. (5) The imaging apparatus according to any one of (1) to (4), wherein the feature region detection unit detects the feature region of a target object using a trained model. (6) The imaging apparatus according to any one of (1) to (4), wherein the feature region detection unit detects the feature region of a target object by image processing including edge detection processing. (7) The imaging device according to any one of (1) to (6), wherein the pre-processing unit outputs images of feature regions alternately in time in different data formats. (8) The imaging device according to any one of (2) to (6), wherein the pre-processing unit outputs metadata of feature regions and depth images alternately in time, and the post-processing unit displays a combination of the feature region images and depth images in a segmented view. (9) The imaging device according to any one of (2) to (6), wherein the post-processing unit performs image analysis of an object to be picked by a robot, the feature region detection unit detects the feature region of the object, the image of the feature region is an uncompressed image, and the metadata is a depth image of the object. (10) The imaging device according to any one of (2) to (6), wherein the downstream processing unit performs image analysis of a three-dimensional measurement of the target object, the feature region detection unit detects the feature region of the target object, the image of the feature region is a high-bit depth image of the target object, and the metadata is a distance image calculated based on the coordinate position data of the pattern light irradiated onto the target object.(11) The imaging device according to any one of (2) to (6), wherein the subsequent processing unit performs image analysis of a three-dimensional measurement of the target object, the feature region detection unit detects the feature region of the target object, the image of the feature region is a high-bit-depth image of the target object, and the metadata is the coordinate data of the center position in the width direction of the slit light irradiated onto the target object. (12) The imaging device according to any one of (1) to (6), wherein the subsequent processing unit performs image analysis of a graphic code reading, the feature region detection unit detects the feature region of the graphic code, the image of the feature region is a high-bit-depth image of the graphic code, and the output interface unit outputs the image of the feature region to the subsequent processing unit uncompressed. (13) The imaging apparatus according to any one of (1) to (6), wherein the subsequent processing unit performs image analysis for tilt correction of the target object, the feature region detection unit detects the feature region of the target object, the image of the feature region is a high-bit-depth image of the target object, and the output interface unit outputs the image of the feature region to the subsequent processing unit uncompressed. (14) The imaging apparatus according to any one of (2) to (6), wherein the metadata is output to the subsequent processing unit during the horizontal padding period outside the feature region in the frame of the image captured by the image sensor. (15) An imaging method that performs the following processes: a computer detects a feature region of an object to be analyzed by a subsequent processing unit from an image captured by an image sensor; determines a data format based on the amount of data of the feature region image and the subsequent bit depth required by the subsequent processing unit; dynamically determines the number of virtual channels based on the amount of data of the feature region image and the subsequent frame rate of the subsequent processing unit; converts the image of the feature region using the determined data format; dynamically notifies the subsequent processing unit of the determined number of virtual channels and outputs the image of the feature region to the subsequent processing unit.(16) A computer-readable recording medium on which a program is recorded that causes a computer to perform the following processes: detect a feature region of an object to be analyzed by a subsequent processing unit from an image captured by an image sensor; determine a data format based on the amount of data of the feature region image and the subsequent bit depth required by the subsequent processing unit; dynamically determine the number of virtual channels based on the amount of data of the feature region image and the subsequent frame rate of the subsequent processing unit; format the image of the feature region according to the determined data format; dynamically notify the subsequent processing unit of the determined number of virtual channels and output the image of the feature region to the subsequent processing unit.

[0067] 1. Imaging System 10. Imaging Device 11. Image Sensor 12. High Dynamic Range Image Generation Unit 13. Image Processing Unit 20. Preprocessing Unit 21. Feature Region Detection Unit 22. Metadata Generation Unit 23. Data Format Determination Unit 24. Virtual Channel Count Determination Unit 25. Post-stage Bit Depth 26. Post-stage Frame Rate 27. Output Interface Unit 28, 34. Input / Output Interface Unit 30. Post-stage Processing Unit 31. Control Unit 32. Image Analysis Unit 33. Result Output Unit

Claims

1. An imaging device comprising: an image sensor; a pre-processing unit that performs pre-processing necessary for image analysis of an image captured by the image sensor and is connected to a subsequent processing unit that performs image analysis; and an output interface unit that outputs pre-processing information to the subsequent processing unit, wherein the pre-processing unit comprises: a feature region detection unit that detects a feature region of an object to be analyzed by the subsequent processing unit; a data format determination unit that determines a data format based on the amount of data of the image of the feature region and the subsequent bit depth required by the subsequent processing unit; and a virtual channel number determination unit that dynamically determines the number of virtual channels based on the amount of data of the image of the feature region and the subsequent frame rate of the subsequent processing unit, wherein the output interface unit converts the image of the feature region to the data format determined by the data format determination unit, dynamically notifies the subsequent processing unit of the number of virtual channels determined by the virtual channel number determination unit, and outputs the image of the feature region to the subsequent processing unit.

2. The imaging apparatus according to claim 1, wherein the preprocessing unit comprises a metadata generation unit that generates metadata for the feature region, the data format determination unit determines the data format based on the amount of data for the image of the feature region and the metadata and the subsequent bit depth required by the subsequent processing unit, the virtual channel count determination unit dynamically determines the number of virtual channels based on the amount of data for the image of the feature region and the metadata and the subsequent frame rate of the subsequent processing unit, and the output interface unit converts the format of the image of the feature region and the metadata according to the data format determined by the data format determination unit, dynamically notifies the subsequent processing unit of the number of virtual channels determined by the virtual channel count determination unit, and outputs the image of the feature region and the metadata to the subsequent processing unit.

3. The imaging apparatus according to claim 1, further comprising a high dynamic range image generation unit that generates a high dynamic range image based on an image captured by the image sensor, wherein the preprocessing unit performs preprocessing on the high bit depth image output from the high dynamic range image generation unit.

4. The imaging apparatus according to claim 2, wherein the pre-processing unit is equipped with a general-purpose CPU, and is connected to the post-processing unit, and the post-processing unit has written a program to the pre-processing unit from the post-processing unit that causes the general-purpose CPU to function as the feature region detection unit and the metadata generation unit, and parameters including the post-processing bit depth and the post-processing frame rate.

5. The imaging apparatus according to claim 1, wherein the feature region detection unit detects the feature region of a target object using a trained model.

6. The imaging apparatus according to claim 1, wherein the feature region detection unit detects the feature region of a target object by image processing including edge detection processing.

7. The imaging apparatus according to claim 1, wherein the preprocessing unit outputs images of feature regions alternately in time in different data formats.

8. The imaging apparatus according to claim 2, wherein the pre-processing unit outputs metadata of the feature region image and the depth image alternately in time, and the post-processing unit combines the feature region image and the depth image and displays them separately.

9. The imaging apparatus according to claim 2, wherein the downstream processing unit performs image analysis of an object to be picked by a robot, the feature region detection unit detects a feature region of the object, the image of the feature region is an uncompressed image, and the metadata is a distance image of the object.

10. The imaging apparatus according to claim 2, wherein the downstream processing unit performs image analysis of a three-dimensional measurement of the target object, the feature region detection unit detects the feature region of the target object, the image of the feature region is a high-bit-depth image of the target object, and the metadata is a distance image calculated based on the coordinate position data of the pattern light irradiated onto the target object.

11. The imaging apparatus according to claim 2, wherein the subsequent processing unit performs image analysis of a three-dimensional measurement of the target object, the feature region detection unit detects the feature region of the target object, the image of the feature region is a high-bit-depth image of the target object, and the metadata is the coordinate data of the center position in the width direction of the slit light irradiated onto the target object.

12. The imaging apparatus according to claim 1, wherein the subsequent processing unit performs image analysis of the graphic code reading, the feature region detection unit detects the feature region of the graphic code, the image of the feature region is a high-bit-depth image of the graphic code, and the output interface unit outputs the image of the feature region to the subsequent processing unit uncompressed.

13. The imaging apparatus according to claim 1, wherein the subsequent processing unit performs image analysis for tilt correction of the target object, the feature region detection unit detects the feature region of the target object, the image of the feature region is a high-bit-depth image of the target object, and the output interface unit outputs the image of the feature region to the subsequent processing unit uncompressed.

14. The imaging apparatus according to claim 2, wherein the metadata is output to the subsequent processing unit during the horizontal padding period outside the feature area in the frame of the image captured by the image sensor.

15. An imaging method that performs the following processes: a computer detects a feature region of an object to be analyzed by a subsequent processing unit from an image captured by an image sensor; determines a data format based on the amount of data of the feature region image and the subsequent bit depth required by the subsequent processing unit; dynamically determines the number of virtual channels based on the amount of data of the feature region image and the subsequent frame rate of the subsequent processing unit; converts the format of the feature region image according to the determined data format; dynamically notifies the subsequent processing unit of the determined number of virtual channels and outputs the feature region image to the subsequent processing unit.

16. A computer-readable recording medium on which a program is recorded that causes a computer to execute the following processes: detecting a feature region of an object to be analyzed by a subsequent processing unit from an image captured by an image sensor; determining a data format based on the amount of data of the feature region image and the subsequent bit depth required by the subsequent processing unit; dynamically determining the number of virtual channels based on the amount of data of the feature region image and the subsequent frame rate of the subsequent processing unit; formatting the image of the feature region according to the determined data format; dynamically notifying the subsequent processing unit of the determined number of virtual channels and outputting the image of the feature region to the subsequent processing unit.