Information processing device, information processing method, and program

JP2026141576APending Publication Date: 2026-09-04CANON KK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025028237
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2026-09-04

AI Technical Summary

Benefits of technology

【0009】 本発明によれば、入力データのビット深度を低くしても、出力の精度の低下を抑制できる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026141576000001_ABST
    Figure 2026141576000001_ABST
Patent Text Reader

Abstract

This technology provides a way to suppress the decrease in output accuracy even when the bit depth of the input data is reduced. [Solution] The information processing device comprises: acquisition means for acquiring input data; bit depth conversion means for converting the bit depth of the input data to a lower bit depth; feature extraction means for extracting features from the input data whose bit depth has been converted; and bit depth expansion means for expanding the bit depth of the features to a higher bit depth.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technology for converting the bit depth of data. Background Art

[0002] In recent years, the accuracy of image recognition technologies such as image classification, object detection, and object tracking has improved dramatically with the advent of Deep Neural Networks (hereinafter abbreviated as DNN). Generally, the computation of DNN involves a large amount of calculation and also consumes a large amount of memory. Therefore, DNN computation is often performed at a low bit depth such as 8 bits. On the other hand, a high bit depth such as 10 bits or 16 bits is sometimes required for the input and output of DNN.

[0003] Patent Document 1 discloses a technology related to a neural network for use when the bit depth of data input to the neural network is greater than the bit depth handled by the neural network. Said technology divides an input having a high bit depth into upper bits and lower bits to obtain a low-bit input. Each of the divided low-bit data items is input to a neural network, and outputs from the neural network are connected in the bit direction, thereby generating a high-bit output. Prior Art Documents Patent Documents

[0004] Patent Document 1 Japanese Unexamined Patent Publication No. 2023-81714 Patent Document 2 Japanese Patent No. 6556033 Non-Patent Documents

[0005] Non-Patent Document 1 Liu, 6 others, “SSD:Single Shot Multibox Detector”, In: ECCV2016 [Retrieved January 31, 2020], Internet <URL:https: / / www.cs.utoronto.ca / ~bonner / courses / 2020s / csc2547 / papers / discriminative / object-detection / ssd,-liu,-eccv-2016.pdf> [Non-Patent Document 2] Jiankang Deng, et al., “RetinaFace: Single-stage Dense Face Localisation in the Wild”, In: arXiv2019 [Retrieved January 31, 2015], Internet<URL:https: / / arxiv.org / pdf / 1905.00641> [Non-Patent Document 3] Liangyu Chen, et al., “Simple Baseline for Image Restoration”, In: arXiv 2022 [Retrieved January 31, 2025], Internet<URL:https: / / arxiv.org / pdf / 2204.04676> [Overview of the project] [Problems that the invention aims to solve]

[0006] However, with the method described above, information from one bit (e.g., the lower bit) is lost when calculating the other bit (e.g., the higher bit), resulting in a significant decrease in output accuracy. In other words, the method described above significantly reduces accuracy by lowering the bit depth.

[0007] Therefore, the present invention provides a technology that can suppress a decrease in output accuracy even when the bit depth of the input data is reduced. [Means for solving the problem]

[0008] To solve this problem, for example, the information processing apparatus of the present invention has the following configuration. That is, A means for acquiring input data, Bit depth conversion means for converting the bit depth of the input data to a lower bit depth, A feature extraction means for extracting features from the input data whose bit depth has been converted, Bit depth extension means for extending the bit depth of the aforementioned feature to a higher bit depth, It is equipped with. [Effects of the Invention]

[0009] According to the present invention, even if the bit depth of the input data is reduced, the decrease in output accuracy can be suppressed. [Brief explanation of the drawing]

[0010] [Figure 1] A block diagram showing the overall configuration, including the main parts of the imaging device of the embodiment. [Figure 2] A diagram of the imaging optical system used to explain the amount of defocus. [Figure 3] A block diagram illustrating the functional configuration of the imaging device according to the first embodiment. [Figure 4] A block diagram showing the functional configuration of the defocus range estimation unit. [Figure 5] A block diagram showing the functional configuration of the input data acquisition unit. [Figure 6] A flowchart of the defocus range estimation process in the first embodiment. [Figure 7] A flowchart illustrating the input data acquisition process for S601. [Figure 8] A flowchart explaining the bit depth conversion process of S602. [Figure 9] A flowchart explaining the input integration process of S603. [Figure 10] A diagram illustrating the integration of inputs in the channel direction. [Figure 11] A diagram illustrating the integration of spatial inputs. [Figure 12] A diagram illustrating the integration of inputs using elemental summation. [Figure 13]A flowchart showing CNN processing for extracting features of S604. [Figure 14] A diagram illustrating the process from input data to bounding box map generation. [Figure 15] A diagram explaining the defocus range of each part of a subject. [Figure 16] A diagram explaining details of bit depth conversion for input data. [Figure 17] A flowchart explaining the bit depth expansion processing of S605. [Figure 18] A flowchart explaining another example of bit depth expansion processing for S605. [Figure 19] A diagram showing transition states in bit depth expansion processing. [Figure 20] A block diagram showing the functional configuration of a learning device that learns a defocus range. [Figure 21] A flowchart of learning processing for learning a defocus range. [Figure 22] A block diagram showing the functional configuration of an imaging apparatus according to a second embodiment. [Figure 23] A flowchart of noise reduction processing according to the second embodiment. [Figure 24] A flowchart explaining the bit depth conversion of S2401. [Figure 25] A diagram explaining integration of images that are input data. [Figure 26] A flowchart showing CNN processing for extracting features of S2403. [Figure 27] A diagram showing an example of expanding bit depth by integrating features of S2404. [Figure 28] A block diagram showing the hardware configuration of a system control unit included in the imaging apparatus 10. DESCRIPTION OF EMBODIMENTS

[0011] The embodiments will be described in detail below with reference to the attached drawings. Although several features are described in the embodiments, not all of these features are necessarily essential to the present invention, and the features may be combined in any way. Furthermore, in the attached drawings, the same or similar configurations are given the same reference numeral, and redundant descriptions are omitted.

[0012] (First Embodiment) In this embodiment, we will describe a case in which an interchangeable-lens imaging device estimates the defocus range by considering the depth of the subject and captures an image that is in focus on the subject.

[0013] This embodiment will now be described with reference to the drawings. Figure 1 is a block diagram showing the overall configuration of the imaging device 10, including its main components. Referring to Figure 1, the overall configuration of the imaging device according to the first embodiment will be described.

[0014] As shown in Figure 1, the imaging device 10 is, for example, a lens-interchangeable digital camera. The imaging device 10 comprises a camera body 100, a lens unit 200, and a lens mount mechanism 113. The camera body 100 is mechanically and electrically connected to the lens unit 200 via the lens mount mechanism 113, so as to be detachable.

[0015] The camera body 100 includes an image sensor 101, a system control unit 102, a shutter 103, a memory 104, a power switch 105, a mode switching unit 106, a rear monitor 107, a touch panel 108, a viewfinder display unit 109, an eyepiece lens 110, and an eyepiece detection unit 111.

[0016] The image sensor 101 converts the optical signal, which is an optical image formed by light from a subject, into an electrical signal and outputs it. The image sensor 101 may be an electronic device such as a CMOS (Complementary Metal Oxide Semiconductor) type imaging sensor or a CCD (Charge Coupled Device) type image sensor.

[0017] The system control unit 102 controls the camera body 100 and incorporates a processor such as a well-known CPU (Central Processing Unit). The system control unit 102 further includes an image processing unit for the video signal obtained by the image sensor 101. The system control unit 102 further includes a phase-difference AF unit that performs focus detection processing using a phase-difference detection method based on focus detection image data (signals for phase-difference AF) obtained from the image sensor 101 and the image processing unit. More specifically, the image processing unit generates a pair of image data formed by light beams passing through a pair of pupil regions of the imaging optical system as focus detection image data. The phase-difference AF unit detects the amount of focus shift based on the amount of shift of the pair of image data. In this way, the phase-difference AF unit of this embodiment performs phase-difference AF (image plane phase-difference AF) based on the output of the image sensor 101 without using a dedicated AF sensor.

[0018] Memory 104 stores programs, variables, constants, etc., for the operation of the system control unit 102. Memory 104 includes, for example, electrically erasable and restorable non-volatile memory. Memory 104 stores various parameters, setting values ​​such as ISO sensitivity, shooting modes, and various correction data.

[0019] The power switch 105 receives an operation from the user to switch the power of the camera body 100 on and off. The power switch 105 outputs the received operation to the system control unit 102.

[0020] The mode switching unit 106 receives user input to switch and set various shooting modes, including live view shooting mode and video shooting mode. The mode switching unit 106 outputs the received input to the system control unit 102.

[0021] The rear monitor 107 includes a display device and LEDs, etc. The rear monitor 107 displays operating status and shooting information such as characters, images, and sounds, as well as messages, in accordance with the execution of a program by the system control unit 102. The display device may be a liquid crystal display device or an organic EL (Electro-Luminescence) display device, etc.

[0022] The touch panel 108 is located on the display surface of the rear monitor 107. The touch panel 108 detects contact with a finger or pen and notifies the system control unit 102 of the contact position on the rear monitor 107. The system control unit 102 then executes the operation or function associated with the contact position.

[0023] The viewfinder display unit 109 displays shooting information in accordance with the execution of a program in the system control unit 102, and together with the eyepiece lens 110, constitutes an electronic viewfinder (EVF). The viewfinder display unit 109 is, for example, a small liquid crystal display device.

[0024] The eyepiece detection unit 111 detects the photographer's eyepiece state and outputs the eyepiece state to the system control unit 102. The system control unit 102 displays the aforementioned shooting information on either the rear monitor 107 or the viewfinder display unit 109 according to the acquired eyepiece state.

[0025] Next, the configuration of the lens unit 200 will be described. The lens unit 200 guides incident light to the image sensor 101. The lens unit 200 includes a photographic lens 201, an aperture 202, a lens drive circuit 203, an aperture control circuit 204, and a lens control unit 205.

[0026] Although only one lens is shown in the illustration for simplification, the photographic lens 201 may be a group of multiple lenses. When light rays from a subject enter the photographic lens 201, the aperture 202 and shutter 103 cause the light rays to form an optical image on the image sensor 101.

[0027] The lens control unit 205 controls the entire lens unit 200 based on instructions from the system control unit 102. Specifically, the lens control unit 205 moves the photographic lens of the lens unit 200 in the optical axis direction via the lens drive circuit 203 to focus on any subject. The lens control unit 205 controls the aperture 202 via the aperture control circuit 204 to adjust the depth of field and the amount of light. The lens control unit 205 has a memory that stores various constants, variables, and programs for lens operation. The lens control unit 205 has a non-volatile memory that holds information for controlling the lens unit 200, such as the maximum aperture value, minimum aperture value, and focal length.

[0028] The system control unit 102 of the camera body 100 calculates the amount of defocus using the output information of the image sensor 101. Based on the calculated amount of defocus, the system control unit 102 communicates via the lens control unit 205 of the lens unit 200 and controls the lens drive circuit 203 to adjust the focus.

[0029] Figure 28 is a block diagram showing the hardware configuration of the system control unit 102 of the imaging device 10. The system control unit 102 is an example of a computer (also called an information processing device). The system control unit 102 includes a processor 2901, a memory 2902, a storage 2903, a communication IF 2904, an input IF 2905, an output IF 2906, and a bus 2907. The processor 2901, memory 2902, storage 2903, communication IF 2904, input IF 2905, and output IF 2906 are connected to each other via the bus 2907 so that they can send and receive information.

[0030] The processor 2901 is an arithmetic processing unit, such as a CPU (Central Processing Unit). The system control unit 102 may have other processors, such as an MPU (Micro Processing Unit), GPU (Graphics Processing Unit), NPU (Neural Processing Unit), and QPU (Quantum Processing Unit), in place of or in addition to the CPU. The processor 2901 realizes various functions of the system control unit 102 by reading computer programs (hereinafter also referred to as programs) stored in the storage 2903 and loading them into the memory 2902. Some or all of the functions of the system control unit 102 may be realized by one or more circuits, such as an ASIC (Application Specific Integrated Circuit) and a PLD (Programmable Logic Device) including an FPGA (Field Programmable Gate Array).

[0031] Memory 2902 corresponds to memory 104 and is a high-speed read / write storage device such as RAM (Random Access Memory). Memory 2902 functions as a work area when processor 2901 executes a program. Memory 2902 temporarily stores the program and parameters necessary for program execution.

[0032] Storage 2903 is a non-volatile storage device such as an HDD (Hard Disk Drive) or an SSD (Solid State Drive). Storage 2903 retains programs, parameters necessary for program execution, and program execution results even when power is not supplied. Storage 2903 stores, for example, trained models for extracting features (also called feature quantities) and model parameters.

[0033] The communication IF2904 is an interface for enabling communication with external devices via a wired or wireless network.

[0034] Input IF2905 is an interface for receiving information from an input device such as a touch panel 108. The input device may be, for example, a mouse or a keyboard.

[0035] Output IF2906 is an interface for outputting images and other information to the rear monitor 107, etc.

[0036] (Explanation of how to calculate the amount of defocus) Here, we will explain the amount of defocus used as depth information for the image in this embodiment. Figure 2 is a diagram of the imaging optical system for explaining the amount of defocus of the imaging optical system. Specifically, Figure 2 shows the relationship between the amount of defocus and the phase difference (image shift amount) between the first focus detection signal and the second focus detection signal acquired from the image sensor.

[0037] The imaging plane 2000 is the plane on which the image sensor 101 is positioned. The exit pupil of the imaging optical system is divided into two parts: a first pupil region 2011 and a second pupil region 2012. The amount of defocus d may be the distance from the imaging position C of the light beams from subjects 2021 and 2022 to the imaging plane 2000, or the magnitude of the distance (i.e., the absolute value of the distance). The amount of defocus in a front-focus state, where the imaging position C is on the subject side of the imaging plane 2000, is negative (d<0). The amount of defocus in a back-focus state, where the imaging position C is on the opposite side of the subject from the imaging plane 2000, is positive (d>0). In a focused state, where the imaging position C is on the imaging plane 2000, d=0. The imaging optical system is in focus with respect to subject 2021 (d=0) and in a front-focus state with respect to subject 2022 (d<0). The front-pinned state (d<0) and back-pinned state (d>0) are examples of the defocused state (|d|>0).

[0038] In the front-focused state (d<0), the light beam from subject 2022 that passes through the first pupil region 2011 (or second pupil region 2012) is focused and then spreads out with a width of Γ1 or Γ2 centered on the centroid position G1 (or centroid position G2) of the light beam. As a result, a blurred image is formed on the image sensor 2000. This blurred image is received by each first focus detection pixel (or each second focus detection pixel) on the image sensor, and a first focus detection signal (or second focus detection signal) is generated. In other words, the first focus detection signal (or second focus detection signal) is a signal that represents a subject image on the image sensor 2000 in which subject 2022 is blurred by a blur width Γ1 (or blur width Γ2) at the centroid position G1 (or centroid position G2) of the light beam.

[0039] The blur width Γ1 (or blur width Γ2) of the subject image increases roughly in proportion to the increase in the magnitude of the defocus amount d |d|. Similarly, the magnitude of the image shift amount p (=difference in the centroid position of the light beam G1-G2) between the first focus detection signal and the second focus detection signal |p| also increases roughly in proportion to the increase in the magnitude of the defocus amount d |d|. The same applies even in a back-focused state (d>0), although the direction of the image shift between the first focus detection signal and the second focus detection signal is opposite to that in a front-focused state.

[0040] Thus, as the amount of defocus in the imaging signal increases, the amount of image shift between the first focus detection signal and the second focus detection signal also increases. Based on this relationship, the phase-difference AF unit of the system control unit 102 performs focus detection using an imaging plane phase-difference detection method, which calculates the amount of defocus from the amount of image shift between the first focus detection signal and the second focus detection signal obtained using the image sensor 101. Specifically, the phase-difference AF unit of the system control unit 102 converts the amount of image shift into a detected defocus amount using a conversion coefficient calculated based on the baseline length. In this embodiment, the unit of defocus is the product of the aperture F value and the allowable circle of confusion diameter δ in the optical system of the imaging device during image capture [Fδ].

[0041] (Functional configuration of the first embodiment) The functional configuration and operation of the first embodiment will be described below. Figure 3 is a block diagram showing the functional configuration of the imaging device 10 of the first embodiment. The imaging device 10 includes a defocus range estimation unit 301 and a control unit 302. The processor 2901 of the system control unit 102 may implement some or all of the functions of the defocus range estimation unit 301 and the control unit 302 by reading and executing a program. Some or all of the functions of the defocus range estimation unit 301 and the control unit 302 may be implemented by circuits such as ASICs.

[0042] The defocus range estimation unit 301 estimates the defocus range of the subject. The defocus range may be the range of the amount of defocus the subject has.

[0043] Figure 14 illustrates the process from input data to the generation of a BB map. The input data is a diagram illustrating the defocus of image 1400 obtained when the imaging device 10 images a person 1401. Figure 14(a) shows image 1400 obtained when the imaging device 10 images a person 1401. Figure 14(c) shows a defocus map 1407, which is a map showing the amount of defocus in each region of image 1400. The defocus map 1407 may show the magnitude of the defocus amount using grayscale.

[0044] Figure 15 is a diagram illustrating the defocus range of each part of the subject. Figure 15(a) shows the defocus range for each part of the person 1401. Defocus range 1500 shows the defocus range of the person's left eye. Defocus range 1501 shows the defocus range of the person's face. Defocus range 1502 shows the defocus range of the person's entire body. Defocus ranges 1500, 1501, and 1502 each visualize the extent of the object in the depth direction as viewed from the imaging device 10. The focus position 1503, shown by the thick line, indicates the focus position of the imaging device 10.

[0045] Figure 15(b) is a schematic diagram showing the estimated defocus range of a person's left eye, right eye, face, and entire body. The horizontal axis represents the magnitude of the defocus amount. The nearest point in the horizontal axis indicates the side closer to the imaging device 10. Conversely, the farthest point in the horizontal axis indicates the side further from the imaging device 10. The length of the arrows extending horizontally represents the defocus range, which is the range of values ​​for the defocus amount of each part.

[0046] As shown in Figure 15(a), for example, in the depth direction of the entire body of person 1401 as viewed from the imaging device 10, the closest point is, for example, the tip of person 1401's nose. On the other hand, the furthest point is, for example, person 1401's heels. Therefore, the maximum value of the defocus amount of the entire body of person 1401 (closest value) is the defocus amount of the tip of person 1401's nose, and the minimum value of the defocus amount (farthest value) is the defocus amount of person 1401's heels. The range defined by these values ​​is the defocus range of the entire body of person 1401.

[0047] In this way, the defocus range estimation unit 301 estimates the defocus range, which is the range of values ​​for the amount of defocus, by taking into account the perspective relationship in the depth direction of the estimated object, such as the eyes, face, and whole body of the subject.

[0048] The control unit 302 calculates the amount of drive for the photographic lens 201 based on the defocus range estimated by the defocus range estimation unit 301 and controls the focus position. In addition, the control unit 302 adjusts the depth of field (DoF) when controlling the aperture 202.

[0049] Figure 4 is a block diagram showing the functional configuration of the defocus range estimation unit 301. The defocus range estimation unit 301 includes an input data acquisition unit 401, a bit depth conversion unit 402, an input integration unit 403, a feature extraction unit 404, and a bit depth expansion unit 405.

[0050] The input data acquisition unit 401 acquires one or more input data necessary for estimating the defocus range. The input data includes, for example, image data of the subject (hereinafter also referred to as the image) and a defocus map. The multiple input data may have different bit depths.

[0051] The bit depth conversion unit 402 converts the bit depth of one or more input data acquired by the input data acquisition unit 401. The bit depth conversion unit 402 may convert the bit depth of any of the input data among the multiple input data. For example, the bit depth conversion unit 402 converts the bit depth of the defocus map included in the input data to a lower bit depth. If the bit depths of the multiple input data are different, the bit depth conversion unit 402 may unify the bit depths by converting the higher bit depth.

[0052] The input integration unit 403 integrates multiple input data. For example, the input integration unit 403 integrates the input data acquired by the input data acquisition unit 401 and the input data whose bit depth has been converted by the bit depth conversion unit 402.

[0053] The feature extraction unit 404 extracts one or more features (also called feature quantities) from the input data that has been converted in bit depth and integrated by the input integration unit 403. The feature extraction unit 404 may extract features using a trained model such as machine learning. The feature extraction unit 404 may extract features related to the defocus range.

[0054] The bit depth extension unit 405 extends and increases the bit depth of the feature data extracted by the feature extraction unit 404. The bit depth extension unit 405 extends the bit depth by, for example, integrating multiple features. The bit depth extension unit 405 extends the bit depth of features related to the defocus range, for example. The bit depth extension unit 405 may extend the bit depth of features that have been reduced by the bit depth conversion unit 402 to the same bit depth as the original input data.

[0055] Figure 5 is a block diagram showing the functional configuration of the input data acquisition unit 401. The input data acquisition unit 401 includes an image acquisition unit 501, a subject detection unit 502, a subject identification unit 503, a defocus map acquisition unit 504, an image cropping unit 505, and a defocus map cropping unit 506.

[0056] The image acquisition unit 501 acquires images captured by the imaging device 10. The images to be acquired are, for example, images 1400 containing people.

[0057] The subject detection unit 502 detects a subject from the image acquired by the image acquisition unit 501. The subject is, for example, a person. The subject detection unit 502 may use techniques such as those described in Non-Patent Document 1 to detect objects such as people as subjects from within the image. The subject detection unit 502 acquires a bounding box (hereinafter abbreviated as BB) indicating the area of ​​the subject, for example, by object detection. The subject detection unit 502 may also use techniques such as those described in Non-Patent Document 2 to detect the face and eyes of a person. Figure 14(b) shows image 1402 with the BBs of various parts and the whole body of person 1401 superimposed from image 1400. BB1403 shows the BB of the left eye. BB1404 shows the BB of the right eye. BB1405 shows the BB of the face. BB1406 shows the BB of the whole body.

[0058] The subject identification unit 503 identifies the subject to be focused on from among the detected subjects. The subject identification unit 503 may identify the subject by receiving a touch from the user via the touch panel 108 while the subject displayed on the rear monitor 107 is shown. In addition, the subject identification unit 503 may identify the subject by automatically detecting the main subject in the image, etc., even without a touch from the touch panel 108. For example, the subject identification unit 503 may automatically detect the main subject in the image using the technology described in Patent Document 2.

[0059] The defocus map acquisition unit 504 acquires a defocus map corresponding to the image captured by the imaging device 10.

[0060] The image cropping unit 505 crops the area containing the subject from the image and resizes it to a predetermined size.

[0061] The defocus map extraction unit 506, similar to the image extraction unit 505, extracts the area containing the subject from the defocus map and resizes the defocus map to a predetermined size.

[0062] The BB map generation unit 507 generates a BB map based on the BB information of the subject obtained by the subject identification unit 503. Figure 14(d) shows an example of a BB map 1408. The BB map has predetermined values ​​input into the areas indicated by the BB of each part of the subject and the whole body (for example, BB 1403 of the right eye, BB 1404 of the left eye, BB 1405 of the face, and BB 1406 of the whole body). Note that there does not need to be only one BB map; there may be a separate map for each part of the subject.

[0063] Figure 6 is a flowchart showing the flow of the defocus range estimation process in the first embodiment. In the following description, each step will be preceded by an S, and the notation of the steps will be omitted. However, the system control unit 102 of the imaging device 10 does not necessarily have to perform all the steps described in the flowchart of Figure 6. Figure 6 shows the processes performed by the system control unit 102 as each step.

[0064] In S601, the input data acquisition unit 401 acquires input data (also simply called input). The input data includes, for example, images captured by the imaging device 10 and a defocus map. Figure 7 shows a flowchart illustrating the input data acquisition process in S601. The input data acquisition process will be explained with reference to Figure 7.

[0065] In step S6011, the image acquisition unit 501 acquires the captured image.

[0066] In S6012, the subject detection unit 502 detects a subject in the image and acquires the subject's BB.

[0067] In S6013, the subject identification unit 503 identifies the subject to be photographed.

[0068] In S6014, the defocus map acquisition unit 504 acquires a defocus map corresponding to the captured image.

[0069] In S6015, the image cropping unit 505 crops a portion of the image containing the subject based on the subject's BB obtained in S6013, and resizes the image to a predetermined size.

[0070] In S6016, the defocus map extraction unit 506 extracts a portion of the defocus map containing the subject based on the subject's BB obtained in S6013, and resizes the defocus map to a predetermined size.

[0071] In S6017, the BB map generation unit 507 generates a BB map as shown in Figure 14(d) based on the BB of the subject obtained in S6013.

[0072] Returning to Figure 6, in S602, the bit depth conversion unit 402 converts the bit depth of the defocus map obtained in S6016. Here, the bit depth of the defocus map is L bits (for example, 16 bits). The bit depth of the image obtained in S6011 is M bits (for example, 8 bits). The feature extraction unit 404 is capable of handling a bit depth of M bits. Also, L bits ≠ M bits. Specifically, L bits > M bits is acceptable. Thus, the bit depth of the image is M bits, which is the same as the bit depth that the feature extraction unit 404 can handle. On the other hand, the bit depth of the defocus map is different from the bit depth of the image and is higher than the bit depth that the feature extraction unit 404 can handle. Therefore, in S602, the bit depth conversion unit 402 converts the bit depth of the defocus map to a bit depth that the feature extraction unit 404 can handle. In other words, the bit depth conversion unit 402 converts the bit depth of the defocus map to a lower bit depth. It can also be said that the bit depth conversion unit 402 unifies the bit depths of multiple input data with different bit depths.

[0073] Figure 8 shows a flowchart illustrating the bit depth conversion process in S602. If the bit depth conversion unit 402 simply or uniformly converts an L-bit defocus map to M bits, information from quantization will be lost. This may reduce the accuracy of defocus range estimation. Therefore, in S602, the bit depth conversion unit 402 applies at least one of several types of conversions to the bit depth conversion of the defocus map. The conversion referred to here is, for example, a nonlinear conversion. For example, in S6021, the bit depth conversion unit 402 applies a first LUT to convert the bit depth of the defocus map. In S6022, the bit depth conversion unit 402 applies a second LUT, which is different from the first LUT, to convert the bit depth of the defocus map. LUT stands for Look Up Table.

[0074] Figure 16 illustrates the details of the bit depth conversion of the input data. Figure 16(a) shows the transformation 1601 using a nonlinear function corresponding to the first LUT applied in S6021. Figure 16(b) shows the transformation 1602 using a nonlinear function corresponding to the second LUT applied in S6022.

[0075] The bit depth conversion unit 402 reduces information loss during quantization by applying multiple different types of bit depth conversions. Furthermore, accuracy near the focal plane (near a defocus value of 0) is particularly important for estimating the defocus range. Therefore, the bit depth conversion unit 402 converts the bit depth in a way that minimizes the loss of information about defocus values ​​with small absolute values. Conversion 1601 increases the slope in the region where the absolute value of the defocus value is small and decreases the slope in the region where the absolute value of the defocus value is large. As a result, conversion 1601 converts the bit depth in a way that minimizes the loss of information about defocus values ​​near the focal plane. On the other hand, conversion 1602 has a smaller change in slope compared to conversion 1601.

[0076] The bit depth conversion unit 402 may convert the bit depth using a LUT based on conversions 1601 and 1602. The number of bins in the LUT does not have to be L bits. For example, if a LUT with fewer bins than L bits is provided, the bit depth conversion unit 402 may interpolate between the bins using linear interpolation or the like. Also, if the defocus value is negative, the bit depth conversion unit 402 may calculate the absolute value of the defocus value, convert it to M bits using the LUT, and then add a negative sign. If the bit depth conversion unit 402 employs multiple conversions, the number of defocus maps increases by the number of conversions. For example, if the bit depth conversion unit 402 applies two types of conversions, conversion 1601 and conversion 1602, the number of defocus maps increases from one to two. In the following description, the M-bit defocus map obtained by conversion 1601 will be referred to as the first defocus map. The M-bit defocus map obtained by conversion 1602 will be referred to as the second defocus map. Note that transformations 1601, 1602, the transformation using the first LUT, and the transformation using the second LUT are examples of the first nonlinear transformation.

[0077] Returning to Figure 6, in S603, the input integration unit 403 integrates the input data. For example, the input integration unit 403 integrates the image and BB map acquired by the input data acquisition unit 401 with the first defocus map and second defocus map whose bit depth has been converted by the bit depth conversion unit 402. In other words, the input integration unit 403 unifies multiple input data whose bit depth has been unified by the bit depth conversion unit 402.

[0078] Figure 9 is a flowchart illustrating the input integration process in S603. The following explanation will refer to Figure 9 to describe several integration processes.

[0079] Figure 9(a) is a flowchart of the input integration process by channel-direction coupling. Figure 10 is a diagram showing the integration of inputs in the channel direction. In S6031, the input integration unit 403 generates an integrated input 1005 by coupling the first defocus map 1001, the second defocus map 1002, the image 1003, and the BB map 1004 in the channel direction. The integrated input 1005 can also be considered the output of the input integration unit 403. The number of channels in input 1005 is the sum of the number of channels in the first defocus map 1001, the second defocus map 1002, the image 1003, and the BB map 1004.

[0080] Figure 9(b) is a flowchart of the input integration process by spatial coupling. Figure 11 is a diagram showing the integration of inputs in the spatial direction. In S6032, the input integration unit 403 equalizes the number of data channels by broadcasting the first defocus map 1101, the second defocus map 1102, the image 1103, and the BB map 1104 in the channel direction. In S6033, the input integration unit 403 combines the first defocus map 1101, the second defocus map 1102, the image 1103, and the BB map 1104 in the spatial direction to generate an integrated input 1105.

[0081] Figure 9(c) is a flowchart of the input integration process by element sum joining. Figure 12 is a diagram showing the integration of inputs by element sum joining. In S6034, the input integration unit 403 equalizes the number of data channels by broadcasting the first defocus map 1101, the second defocus map 1102, the image 1103, and the BB map 1104 in the channel direction. In S6035, the input integration unit 403 combines the first defocus map 1204, the second defocus map 1203, the image 1201, and the BB map 1202 by calculating the element sum, thereby generating an integrated input 1205.

[0082] Returning to Figure 6, in S604, the feature extraction unit 404 extracts features from the integrated input. The feature extraction unit 404 may use a Convolutional Neural Network (hereinafter abbreviated as CNN) or the like to extract features.

[0083] Figure 13 is a flowchart showing the CNN processing for extracting features in S604. The CNN includes processing using nonlinear operations such as Convolution, ReLU, and Max Pooling. The CNN may have multiple of these elements. The CNN may also have Global Average Pooling, Fully Connected, etc. The feature extraction unit 404 performs Convolution in S6041, ReLU in S6042, and Max Pooling in S4043. After that, the feature extraction unit 404 performs Convolution in S6044 and ReLU in S6045. The feature extraction unit 404 performs Convolution in S6046 and ReLU in S6047 again. The feature extraction unit 404 may repeat Convolution and ReLU three or more times. The feature extraction unit 404 extracts features by performing Global Average Pooling in S6048 and Fully Connected in S6049.

[0084] The feature extraction unit 404 may be a Multilayer Perceptron or a Multi-head Self Attention. The above example does not limit the feature extraction unit 404, and its configuration may be changed as appropriate. The feature extraction unit 404 can be computed with a bit depth of, for example, M bits.

[0085] The number of feature elements output by the feature extraction unit 404 is the product of the number of parts for which the defocus range is estimated, the locations for which the defocus range value is obtained, and the ensemble number N for integrating the outputs described later. For example, if four parts—left eye, right eye, face, and whole body—are output at two locations, near and far, and with an ensemble number of 2 for integration, the feature extraction unit 404 will output 16 elements (=4*2*2=(number of parts)*(number of locations (=near, far))*(ensemble number)).

[0086] Returning to Figure 6, in S605, the bit depth extension unit 405 increases the bit depth of the features by extending the bit depth of the features acquired by the feature extraction unit 404, and outputs the estimated defocus range.

[0087] Figure 17 is a flowchart illustrating the bit depth expansion process of S605. Figure 19 is a diagram showing the transition states of the bit depth expansion process. Figure 19(a) shows one or more M-bit features. Figure 19(b) shows the state in which the M-bit features have been decomposed into N groups of elements. Figure 19(c) shows the output in which the decomposed elements have been combined using element-wise product. Figure 19(d) shows the values ​​of each element of the output feature quantity.

[0088] In the bit depth expansion process, in S6051, the bit depth expansion unit 405 decomposes the M-bit features extracted by the feature extraction unit 404, shown in Figure 19(a), into multiple elements of N groups, shown in Figure 19(b), by decomposing them in the channel direction. N is the number of ensembles at the time of integration and is predetermined.

[0089] Next, in S6052, the bit depth extension unit 405 calculates and integrates the element-wise product of the decomposed features to generate the output shown in Figure 19(c). When the bit depth extension unit 405 integrates features using element-wise product, the bit depth after feature integration becomes M bits * N bits. For example, if M is 8 and N is 2, the bit depth extension unit 405 obtains a 16-bit output. Each element of the 16-bit output corresponds to the near and far sides of the defocus range of each area to be output. As shown in Figure 19(d), the 0th to 7th elements of the feature quantity indicate the defocus amounts on the near and far sides of each area, respectively. The bit depth extension unit 405 may use the M * N bit output as is, or it may convert the bit depth of the output by bit shifting or other means as needed.

[0090] Figure 18 is a flowchart illustrating another example of the bit depth extension process in S605. The bit depth extension unit 405 may perform the bit depth extension process as shown in Figure 18.

[0091] Specifically, in S6051, the bit depth extension unit 405 decomposes the M bits of features extracted by the feature extraction unit 404 in the channel direction.

[0092] In S6053 and S6054, the bit depth extension unit 405 applies a third LUT and a fourth LUT to the decomposed features. Here, the third LUT corresponds to the first LUT, and the fourth LUT corresponds to the second LUT. For example, the transformation by the third LUT is the inverse transformation of the transformation by the first LUT. For example, the transformation by the fourth LUT is the inverse transformation of the transformation by the second LUT. The transformations by the third and fourth LUTs are nonlinear transformations and are examples of the second nonlinear transformation.

[0093] In S6055, the bit depth extension unit 405 integrates the features to which the LUT has been applied by calculating the element sum, thereby generating an output. When M-bit features are integrated by summing them with an ensemble number N, the number of bits in the integrated output will be M+N bits. The bit depth extension unit 405 may also integrate the features by averaging the elements of the features instead of summing them.

[0094] As a result, in S605, the bit depth extension unit 405 can obtain a defocus range for each part with an expanded and increased bit depth. The bit depth extension unit 405 outputs the defocus range with the expanded bit depth to, for example, the control unit 302.

[0095] The control unit 302 can calculate the amount of drive of the photographic lens 201 based on the defocus range of any obtained subject area and control the focus position. In addition, the control (adjustment) of the aperture 202 is performed to adjust the depth of field (DoF). Since the defocus range is output with a bit depth larger than the M bits that can be calculated by the feature extraction unit, finer adjustment of the lens drive amount becomes possible. As a result, the imaging device 10 becomes easier to focus on any subject area.

[0096] (Learning the defocused range) Next, a learning device and learning method for learning the defocus range of a subject will be described. Figure 20 is a block diagram showing the functional configuration of a learning device for learning the defocus range. The learning device 21 may be provided on the imaging device 10, or it may be separate from the imaging device 10. The learning device 21 includes a defocus range estimation unit 2100, a true value acquisition unit 2101, a loss calculation unit 2102, a parameter update unit 2103, a parameter storage unit 2104, and a storage unit 2105.

[0097] The defocus range estimation unit 2100 has the same function as the defocus range estimation unit 301 during inference. The defocus range estimation unit 2100 takes the image and the defocus map as input, estimates the defocus range of each part of the subject, and outputs the estimation result. This estimation result can also be considered the output of the bit depth extension unit 405.

[0098] The true value acquisition unit 2101 acquires the correct label, which includes the preset true values ​​of the defocus range for each part of the subject.

[0099] The images prepared as training data, and the defocus ranges that serve as the correct labels, may be values ​​calculated from the focus detection signal acquired at the same time as the image acquisition, determining the amount of defocus for each focus detection region. For example, the imaging device 10 may calculate the amount of defocus. Alternatively, an external computing device such as a computer may hold the focus detection signal and the image signal and calculate the amount of defocus.

[0100] When associating the correct label for the defocus range with an image, the correct amount of defocus for a part of the subject may be determined by referencing the amount of defocus for that area of ​​the subject in the image and the amount of defocus for at least one of the background and foreground obstacles, and adopting an appropriate range. Users may also associate the correct label for the defocus range with the image while visually confirming it. Alternatively, the defocus range of the correct label may be calculated by segmenting the image and using the region of each part of the subject that does not include background and foreground obstacles, and the region where the focus detection area overlaps.

[0101] The loss calculation unit 2102 calculates the loss by comparing the estimated result output from the defocus range estimation unit 2100 with the correct label obtained by the true value acquisition unit 2101.

[0102] The parameter update unit 2103 updates the parameters of the defocus range estimation unit 2100 based on the loss calculated by the loss calculation unit 2102.

[0103] The parameter storage unit 2104 stores the parameters of the defocus range estimation unit 2100 in the storage unit 2105.

[0104] If the imaging device 10 and the learning device 21 are separate units, the imaging device 10 may update the parameters of the defocus range estimation unit 301 based on the parameters of the defocus range estimation unit 2100.

[0105] Figure 21 is a flowchart of the learning process for learning the defocus range. The flow of processing S601 to S605 is the same as during inference. In the learning process, S601 is executed first.

[0106] In S2201, the true value acquisition unit 2101 acquires the correct labels for the defocus range of each part of the subject. The correct labels include the true values ​​of the near and far defocus ranges of each part of the subject, as shown in Figure 15(b). After this, S602 to S605 are executed.

[0107] In S2202, the loss calculation unit 2102 calculates the loss. The loss is calculated based on the estimated value of the defocus range obtained from S605 and the true value of the correct label of the defocus range obtained in S2201. The loss is obtained, for example, by the L1 norm in equation (1) below.

[0108]

number

[0109] Here, the definitions of each variable are as follows:

number

[0110] In S2203, the parameter update unit 2103 updates the parameters of the defocus range estimation unit 2100 based on the loss calculated in S2201. The parameters may be the weights of elements such as convolutions in a neural network. The parameter update unit 2103 may update the parameters based on backpropagation, such as using Momentum SGD.

[0111] In S2204, the parameter update unit 2103 terminates learning if it determines that the loss obtained by the loss calculation unit 2102 has converged. If the parameter update unit 2103 determines that the loss has not converged, it returns to S601. The parameter update unit 2103 may determine that the loss has converged if the loss falls below a certain value. The parameter update unit 2103 may also determine that the loss has converged if the number of iterations from S601 to S2203 exceeds a predetermined number.

[0112] In S2205, the parameter storage unit 2104 stores the parameters of the defocus range estimation unit 2100. If the imaging device 10 and the learning device 21 are integrated, the storage of the parameters of the defocus range estimation unit 2100 becomes the storage of the parameters of the defocus range estimation unit 301. If the imaging device 10 and the learning device 21 are separate, the parameters of the defocus range estimation unit 301 are then updated with the parameters of the defocus range estimation unit 2100.

[0113] Furthermore, the subject does not have to be of only one type. Also, the subject is not limited to people, but may include multiple types of subjects, such as animals like dogs and cats, and vehicles like cars and trains. In other words, this embodiment does not limit the category or number of types of subjects.

[0114] (Effects of the first embodiment) This embodiment converts the bit depth of the input data to a lower bit depth, extracts features, and then expands the bit depth of the features. As a result, this embodiment can suppress the loss of information in the input data during the bit depth conversion and suppress the decrease in the accuracy of the output features.

[0115] In this embodiment, the feature extraction unit 404 integrates multiple features extracted using element-wise multiplication or the like. This makes it possible for this embodiment to convert the bit depth without splitting the upper bits and lower bits. Therefore, even when the bit depth that the feature extraction unit 404 can handle is lower than the bit depth of the input defocus map, this embodiment makes it possible to estimate the defocus range while suppressing a decrease in accuracy.

[0116] This embodiment converts the bit depth of the input data to a lower bit depth using a nonlinear transformation. Furthermore, this embodiment expands the bit depth of the extracted features using a nonlinear transformation. As a result, this embodiment can suppress the loss of information in important areas (e.g., near the focal plane), thereby further suppressing the decrease in output accuracy.

[0117] In this embodiment, the nonlinear transformation applied to a bit depth conversion that reduces the bit depth may be the inverse transformation of the nonlinear transformation applied to a bit depth extension that increases the bit depth. In this case, this embodiment can further suppress the loss of information due to bit depth conversion.

[0118] This embodiment unifies the bit depth of multiple input data by converting the bit depth to reduce it. As a result, this embodiment can easily integrate multiple input data.

[0119] (Second embodiment) The second embodiment performs various noise reduction processing tasks to reduce image noise. Figure 22 is a block diagram showing the functional configuration of the imaging device of the second embodiment. The imaging device 23 of the second embodiment includes an image acquisition unit 501, a bit depth conversion unit 402, an input integration unit 403, a feature extraction unit 404, and a bit depth expansion unit 405. Note that each component of the second embodiment has the same function as each component of the first embodiment, so the explanation will be simplified.

[0120] The image acquisition unit 501 acquires the image captured by the imaging device 23 as input data.

[0121] The bit depth conversion unit 402 converts the bit depth of the image acquired by the image acquisition unit 501 and generates multiple images.

[0122] The input integration unit 403 integrates the images whose bit depth has been converted in the bit depth conversion unit 402.

[0123] The feature extraction unit 404 extracts features from the image obtained by the input integration unit 403.

[0124] The bit depth extension unit 405 integrates the features extracted by the feature extraction unit 404 and converts the bit depth.

[0125] In this embodiment, the bit depth of the image is N bits. The bit depth that the feature extraction unit 404 can handle is M bits. For example, N > M. The bit depth of the output obtained from the bit depth extension unit 405 is K bits.

[0126] Figure 23 is a flowchart of the noise reduction process in the second embodiment.

[0127] In S601, the image acquisition unit 501 acquires input data including an image of the subject.

[0128] In S2401, the bit depth conversion unit 402 applies several different functions to the input data, which includes an N-bit image, and then quantizes it to convert the bit depth of the image to M bits, which is lower than N bits.

[0129] Figure 24 is a flowchart illustrating the bit depth conversion process in S2401. In S2501, the bit depth conversion unit 402 applies a first function, which is a nonlinear function, to the image. In S2502, a second function, which is different from the first function and is also a nonlinear function, is applied to the image. The nonlinear function may be, for example, the function shown in equation (2).

[0130]

number

[0131] Here, x is the input pixel value. α and β are predetermined parameters. sign(·) is a function that returns the sign of the input data. The first and second functions have different α and β values. Therefore, the bit depth conversion unit 402 can achieve different bit depth conversions on an image by applying the first and second functions, which have different parameters α and β, to the image.

[0132] In S2503, the bit depth conversion unit 402 quantizes the pixel values ​​of the image converted in S2501. In S2504, the bit depth conversion unit 402 quantizes the pixel values ​​of the image converted in S2502. As a result, the bit depth conversion unit 402 converts the bit depth of the image to a lower bit depth and generates two different types of data.

[0133] In S2502 and S2502, the bit depth conversion unit 402 applied two functions as multiple different conversions, but the type and number of functions to be applied are not limited. Furthermore, the bit depth conversion unit 402 may apply different LUTs instead of functions. The LUTs are generated, for example, based on the values ​​of the first and second functions. The function parameters and LUT values ​​may be determined by noise reduction learning.

[0134] Next, in S2402, the input integration unit 403 integrates the M-bit images acquired by the image acquisition unit 501 in S2501. Integration may be performed in the same manner as in the first embodiment. Integration may be performed, for example, by combining the M-bit images in the channel direction.

[0135] Figure 25 illustrates the integration of input images. Images 2601 and 2602 in Figure 25 are multiple M-bit images obtained in S2401. The integrated input 2603 shows the state in which images 2601 and 2602 are integrated in the channel direction. In other words, the integrated input 2603 can be said to be the output of the input integration unit 403.

[0136] Next, in S2403, the feature extraction unit 404 extracts features from the integrated input image. The feature extraction unit 404 may use a Convolutional Neural Network (hereinafter abbreviated as CNN) or the like to extract features.

[0137] Figure 26 is a flowchart showing the CNN processing for extracting features in S2403. The CNN may be a CNN for noise reduction. The feature extraction unit 404 extracts features by performing Convolution in S2701, ReLU in S2702, and MaxPooling in S2703, while reducing the spatial resolution. The feature extraction unit 404 further reduces the spatial resolution while extracting features by performing Convolution in S2704, ReLU in S2705, and MaxPooling in S2706. Subsequently, the feature extraction unit 404 increases the resolution by UpSampling in S2707. After that, the feature extraction unit 404 extracts features through Convolution in S2708 and UpSampling in S2709.

[0138] Next, the feature extraction unit 404 extracts features by performing the Convolution in S2710 and the ReLU in S2711. The feature extraction unit 404 extracts features by performing the Convolution in S2712, which is different from the Convolution in S2710, and the ReLU in S2713, which is different from the ReLU in S2711. As a result, the feature extraction unit 404 outputs two types of M-bit features. These two types of features are referred to as the first feature and the second feature, respectively. The feature extraction unit 404 may output three or more types of features, and the output of the feature extraction unit 404 may be changed as appropriate.

[0139] Next, in S2404, the bit depth extension unit 405 integrates the multiple features obtained in S2403 and performs a conversion by extending the bit depth.

[0140] Figure 27 shows an example of extending the bit depth by integrating the features of S2404. The bit depth extension unit 405 can obtain an M*N bit output by calculating the element-wise product of the first feature 2801 and the second feature 2802, which are each obtained in M ​​bits. Here, N is the number of features to be element-wise multiplied. In this embodiment, N is 2. The bit depth extension unit 405 may output the M*N bit output as is, or it may convert it to K bits by bit shifting or the like before outputting it.

[0141] The method for learning noise reduction on images is described in detail in Non-Patent Document 3.

[0142] <Effects of the second embodiment> In the second embodiment, multiple features output from the feature extraction unit 404 are integrated using element-wise AND or the like. This makes it possible for the second embodiment to convert the bit depth without splitting the upper and lower bits. Therefore, even when the bit depth that the feature extraction unit 404 can handle is lower than the bit depth of the input image, it is possible to perform noise reduction while suppressing a decrease in accuracy.

[0143] The embodiments described above may be combined as appropriate. When combining embodiments, the system may be configured so that users can select functions and other elements.

[0144] (Other embodiments) The present invention can also be realized by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. Furthermore, the present invention can also be realized by a circuit (e.g., an ASIC) that implements one or more functions.

[0145] The inventions described herein include the following information processing apparatus, information processing method, and program. (Item 1) A means for acquiring input data, Bit depth conversion means for converting the bit depth of the input data to a lower bit depth, A feature extraction means for extracting features from the input data whose bit depth has been converted, Bit depth extension means for extending the bit depth of the aforementioned feature to a higher bit depth, An information processing device characterized by comprising: (Item 2) The acquisition means acquires multiple input data, The bit depth conversion means converts the bit depth of at least one of the input data of a plurality of input data, The system includes an input integration means for integrating multiple input data, including the input data whose bit depth has been converted, The feature extraction means extracts features from the integrated plurality of input data. The information processing device described in item 1, characterized by the features described herein. (Item 3) The bit depth conversion means converts the bit depth by a first nonlinear conversion. The information processing device described in item 2, characterized in that it is an information processing device. (Item 4) The feature extraction means learns based on the output of the bit depth expansion means and a preset true value. An information processing device according to any one of items 1 to 3, characterized by the features described in item 1 to 3. (Item 5) The bit depth extension means decomposes and integrates the features in the channel direction. An information processing device according to any one of items 1 to 4, characterized in that it is the same as described in item 1 to 4. (Item 6) The bit depth extension means integrates the features by calculating the element product of the decomposed features. The information processing device described in item 5, characterized by the features described herein. (Item 7) The bit depth extension means applies a second nonlinear transformation to the decomposed features. The information processing device described in item 5, characterized by the features described herein. (Item 8) The bit depth extension means extends the bit depth by calculating the element sum of the features to which the second nonlinear transformation has been applied. The information processing device described in item 7, characterized by the features described herein. (Item 9) The bit depth conversion means converts the bit depth by a first nonlinear transformation, which is the inverse transformation of the second nonlinear transformation. The information processing device described in item 8, characterized by the features described herein. (Item 10) The bit depth extension means has parameters generated by the learning of the feature extraction means. An information processing device according to any one of items 1 to 9, characterized in that it is an information processing device. (Item 11) The input integration means combines the multiple input data by one of the following methods: combining the multiple input data in the channel direction, combining the multiple input data in the spatial direction, or combining the multiple input data by elemental sum. The information processing device described in item 2, characterized in that it is an information processing device. (Item 12) The acquisition means acquires multiple input data with different bit depths. An information processing device according to any one of items 1 to 11, characterized by the features described in item 1 to 11. (Item 13) The bit depth conversion means unifies the bit depth of the multiple input data by converting the bit depth. The information processing device described in item 12, characterized by the features described herein. (Item 14) The acquisition means acquires a plurality of first input data, including an image of the subject, The bit depth extension means outputs information regarding the defocus range of the subject. An information processing device according to any one of items 1 to 13, characterized by the features described in item 1 to 13. (Item 15) The bit depth extension means extends and outputs the bit depth of the features relating to the defocus range for each part of the subject. The information processing device described in item 14, characterized by the features described herein. (Item 16) The acquisition means acquires a plurality of input data, including a first input data and a second input data having a higher bit depth than the first input data. The bit depth conversion means converts the bit depth of the second input data. An information processing device according to any one of items 1 to 15, characterized by the features described herein. (Item 17) The bit depth conversion means converts the bit depth of one or more input data from among multiple input data by a plurality of different conversions. The information processing device described in item 1, characterized by the features described herein. (Item 18) The acquisition process involves obtaining input data, A bit depth conversion step that converts the bit depth of the input data to a lower bit depth, A feature extraction step is performed to extract features from the input data whose bit depth has been converted. A bit depth extension process for extending the bit depth of the aforementioned feature to a higher bit depth, An information processing method characterized by having the following features. (Item 19) A program to cause a computer to function as one of the information processing devices described in any one of items 1 through 17.

[0146] The invention is not limited to the embodiments described above, and various modifications and variations are possible without departing from the spirit and scope of the invention. Accordingly, claims are attached to disclose the scope of the invention. [Explanation of Symbols]

[0147] 102...System control unit, 401...Input data acquisition unit, 402...Bit depth conversion unit, 403...Input integration unit, 404...Feature extraction unit, 405...Bit depth expansion unit, 501...Image acquisition unit.

Claims

1. A means for acquiring input data, Bit depth conversion means for converting the bit depth of the input data to a lower bit depth, A feature extraction means for extracting features from the input data whose bit depth has been converted, Bit depth extension means for extending the bit depth of the aforementioned feature to a higher bit depth, An information processing device characterized by comprising:

2. The acquisition means acquires multiple input data, The bit depth conversion means converts the bit depth of at least one of the input data of a plurality of input data, The system includes an input integration means for integrating multiple input data, including the input data whose bit depth has been converted, The feature extraction means extracts features from the integrated plurality of input data. The information processing apparatus according to feature 1.

3. The bit depth conversion means converts the bit depth by a first nonlinear conversion. The information processing apparatus according to feature 2.

4. The feature extraction means learns based on the output of the bit depth expansion means and a preset true value. The information processing apparatus according to feature 1.

5. The bit depth extension means decomposes and integrates the features in the channel direction. The information processing apparatus according to feature 1.

6. The bit depth extension means integrates the features by calculating the element product of the decomposed features. The information processing apparatus according to feature 5.

7. The bit depth extension means applies a second nonlinear transformation to the decomposed features. The information processing apparatus according to feature 5.

8. The bit depth extension means extends the bit depth by calculating the element sum of the features to which the second nonlinear transformation has been applied. The information processing apparatus according to feature 7.

9. The bit depth conversion means converts the bit depth by a first nonlinear transformation, which is the inverse transformation of the second nonlinear transformation. The information processing apparatus according to feature 8.

10. The bit depth extension means has parameters generated by the learning of the feature extraction means. The information processing apparatus according to feature 1.

11. The input integration means combines the multiple input data by one of the following methods: combining the multiple input data in the channel direction, combining the multiple input data in the spatial direction, or combining the multiple input data by elemental sum. The information processing apparatus according to feature 2.

12. The acquisition means acquires multiple input data with different bit depths. The information processing apparatus according to feature 1.

13. The bit depth conversion means unifies the bit depth of the multiple input data by converting the bit depth. The information processing apparatus according to feature 12.

14. The acquisition means acquires a plurality of first input data, including images of the subject, The bit depth extension means outputs information regarding the defocus range of the subject. The information processing apparatus according to feature 1.

15. The bit depth extension means extends and outputs the bit depth of the features relating to the defocus range for each part of the subject. The information processing apparatus according to feature 14.

16. The acquisition means acquires a plurality of input data, including a first input data and a second input data having a higher bit depth than the first input data. The bit depth conversion means converts the bit depth of the second input data. The information processing apparatus according to feature 1.

17. The bit depth conversion means converts the bit depth of one or more input data from among a plurality of input data by a plurality of different conversions. The information processing apparatus according to feature 1.

18. The acquisition process involves obtaining input data, A bit depth conversion step that converts the bit depth of the input data to a lower bit depth, A feature extraction step is performed to extract features from the input data whose bit depth has been converted. A bit depth extension process for extending the bit depth of the aforementioned feature to a higher bit depth, An information processing method characterized by having the following features.

19. A program for causing a computer to function as one of the means of an information processing device according to any one of claims 1 to 17.

Citation Information

Patent Citations

  • Image processing apparatus, control method thereof, and program

    JP2023081714A

  • Image processing device, image processing method, and program

    JP6556033B2