Image processing apparatus and image processing method
Patent Information
- Application Number
- JP2022072675
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-04-26
- Publication Date
- 2025-05-07
AI Technical Summary
The use of machine learning for subject tracking in battery-powered devices increases power consumption, leading to a shortened operating time.
Implement a dual-tracking system with a first tracking process using machine learning and a second process that does not use machine learning, controlled by a frequency management mechanism to reduce power consumption.
Enables accurate subject tracking using machine learning while minimizing power consumption by alternating the operation frequencies of the tracking processes.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image processing apparatus and an image processing method, and particularly to a subject tracking technique.
Background Art
[0002] For images captured in a time series such as a video, a subject tracking function that tracks a specific region over time is known. Conventionally, the subject tracking function has been realized by template matching that uses the region of the tracking target as a template and searches for the region with the highest correlation with the template. On the other hand, in recent years, a subject tracking function using machine learning such as deep learning (DL) has been proposed (Patent Document 1).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] By using machine learning, tracking with higher accuracy than template matching can be realized. However, the use of machine learning requires a very large number of operations, and in order to execute at high speed, it is necessary to use a high-performance and high-power-consuming circuit. However, in a battery-powered device such as an imaging apparatus, an increase in power consumption leads to a shortening of the operable time.
[0005] In one aspect of the present invention, there is provided an image processing apparatus and an image processing method that enable the use of a subject tracking function using machine learning while suppressing an increase in power consumption.
Means for Solving the Problems
[0006] The above objective is achieved by an image processing device comprising: a first tracking means for applying a first tracking process using machine learning to an image; a second tracking means for applying a second tracking process without machine learning to an image; and a control means for controlling the operation of the first tracking means and the second tracking means, wherein the control means controls the operation frequency of the first tracking means to be lower than the operation frequency of the second tracking means. [Effects of the Invention]
[0007] According to the present invention, it is possible to provide an image processing device and an image processing method that enable the use of a subject tracking function utilizing machine learning while suppressing an increase in power consumption. [Brief explanation of the drawing]
[0008] [Figure 1] Block diagram showing an example of the functional configuration of a digital camera as an example of an image processing apparatus according to the embodiment. [Figure 2] This diagram shows an example of live view display on the display unit. [Figure 3] Timing chart for an example of subject tracking function in an embodiment [Figure 4] Timing chart for another example of subject tracking function in the embodiment [Figure 5] Timing chart for yet another example of subject tracking function in the embodiment [Figure 6] Flowchart relating to the subject tracking function in the embodiment [Figure 7] Flowchart relating to the subject tracking function in the embodiment [Modes for carrying out the invention]
[0009] The present invention will be described in detail below with reference to the attached drawings, based on exemplary embodiments thereof. Note that the following embodiments do not limit the invention to the claims. Furthermore, while multiple features are described in the embodiments, not all of them are essential to the invention, and the multiple features may be combined arbitrarily. In addition, in the attached drawings, the same or similar configurations are given the same reference numeral, and redundant descriptions are omitted.
[0010] In the following embodiments, the present invention will be described in relation to its implementation using a digital camera. However, the present invention can also be implemented using any electronic device having an imaging function. Such electronic devices include video cameras, computer equipment (personal computers, tablet computers, media players, PDAs, etc.), mobile phones, smartphones, game consoles, robots, drones, and dashcams. These are examples, and the present invention can also be implemented using other electronic devices.
[0011] ●(First Embodiment) Figure 1 is a block diagram showing an example of the functional configuration of a digital camera 101, which is an example of an image processing apparatus according to the first embodiment of the present invention. The optical system 102 is a lens unit having multiple lenses, and it forms an optical image of the subject on the imaging surface of the image sensor 104.
[0012] The image sensor 104 may be, for example, a known CCD or CMOS color image sensor having a primary color Bayer array color filter. The image sensor 104 has a pixel array in which a plurality of pixels are arranged in two dimensions, and peripheral circuits for reading signals from each pixel. Each pixel accumulates charge according to the amount of incident light by photoelectric conversion. By reading signals having a voltage corresponding to the amount of charge accumulated during the exposure period from each pixel, a group of pixel signals (analog image signals) representing the subject image formed on the imaging surface is obtained.
[0013] The control unit 103 has one or more processors (hereinafter referred to as CPUs) capable of executing programs. The control unit 103 controls the operations of each part of the digital camera 101 and realizes the functions of the digital camera 101 by reading the programs stored in the ROM 121 into the RAM 122 and executing them. The control unit 103 is communicably connected to each block that controls the operations.
[0014] The ROM 121 is electrically rewritable, for example, and stores programs executable by the CPU of the control unit 103, setting values of the digital camera 101, data of the GUI to be displayed on the display unit 120, and the like.
[0015] The RAM 122 is used to read the programs executed by the CPU of the control unit 103 and to store the values necessary during the execution of the programs. Note that the tracking memory 109, memory 108, and RAM 122, which will be described later, may be implemented as areas allocated within one memory space.
[0016] The AF sensor 123 generates a signal pair for performing autofocus detection (AF) using the phase difference detection method for a preset focus detection area within the shooting range and outputs it to the control unit 103. The control unit 103 obtains the phase difference of the signal pair acquired from the AF sensor 124 and converts the phase difference into a defocus amount. Then, the control unit 103 controls the position of the focus lens included in the optical system 102 according to the defocus amount to focus the optical system 102 on the focus detection area.
[0017] The AE sensor 124 generates luminance information for the entire shooting range and / or a specific part thereof and outputs it to the control unit 103. The control unit 103 executes AE processing to determine exposure conditions based on the luminance information acquired from the AE sensor 124 and the exposure setting. The exposure conditions are, for example, a combination of shutter speed or exposure time, aperture value, and shooting sensitivity.
[0018] Note that the AF and AE processes described here are merely examples, and the control unit 103 can adjust the focal length of the optical system 102 or determine exposure conditions by any of various known methods. For example, the control unit 103 can set a focus detection area in the subject area or determine exposure conditions so that the subject area has appropriate exposure.
[0019] The analog image signal for one frame of a still image or a moving image, which is captured and read out by the imaging device 104, is supplied to the preprocessing unit 105 and the tracking preprocessing unit 106. This
[0020] The analog image signal read out from the imaging device 104 has one color component corresponding to the color of the color filter provided for each pixel and is called a RAW format image signal. In this embodiment, since the imaging device 104 has a color filter in a primary color Bayer array, the pixel signals constituting the RAW format image signal have one color component of R, G, or B.
[0021] The preprocessing unit 105 applies A / D conversion and color interpolation processing to the RAW format analog image signal read out from the imaging device 104, and generates image data for one frame in which pixel data has all color components (RGB components). The preprocessing unit 105 stores the generated image data in the memory 108. The memory 108 has a capacity capable of storing image data for two or more frames. Thereby, the preprocessing unit 105 and the correction unit 110 can perform processing using image data of a plurality of frames.
[0022] The correction unit 110 applies white balance correction, shading correction, etc. to the image data stored in the memory 108. Also, the correction unit 110 stores the corrected image data in the memory 108. Note that the correction unit 110 may convert the corrected image data from the RGB format to the YUV format and then store it in the memory 108.
[0023] The post-processing unit 113 applies image processing according to the intended use to the corrected image data to generate recording image data and display image data. For example, the post-processing unit 113 applies scaling to a resolution suitable for display, or applies encoding processing according to a predetermined recording format to generate display image data and recording image data. If recording is not performed, the post-processing unit 113 generates only display image data.
[0024] When live view is displayed on the display unit 120, the image sensor 104 records video, and the pre-processing unit 105, correction unit 110, and post-processing unit 113 generate display image data so that the live view image is displayed on the display unit 120 at a predetermined frame rate.
[0025] The post-processing unit 113 outputs the recording image data to the recording unit 117 and the display image data to the tracking frame superimposition unit 119.
[0026] The recording unit 117 records the recording image data converted by the post-processing unit 113 onto the recording medium 118. The recording medium 118 is, for example, a memory card. The recording medium 118 may also be an external storage device.
[0027] The tracking preprocessor 106 generates image data from the analog image signal read from the image sensor 104, similar to the preprocessor 105, and stores it in the tracking memory 109. The memory 109 has a capacity to store image data for two or more frames. This allows the tracking preprocessor 106 and the tracking correction unit 107 to perform processing using image data from multiple frames.
[0028] The tracking correction unit 107 applies white balance correction and shading correction to the image data stored in the tracking memory 109, similar to the correction unit 110. The tracking correction unit 107 also stores the corrected image data in the tracking memory 109. The tracking correction unit 107 converts the corrected image data from RGB format to YUV format before storing it in the tracking memory 109.
[0029] Furthermore, the tracking correction unit 107 may apply image processing to the corrected image data to improve the subject detection accuracy. For example, if the average brightness of the corrected image data is below a threshold, the tracking correction unit 107 can apply image processing to improve the overall brightness of the image data. This image processing may be a process that multiplies the value of individual pixel data by a certain coefficient.
[0030] Furthermore, any processing common to the pre-processing unit 105 and the tracking pre-processing unit 106 may be performed by only one of them, and the processed image data may be passed to the other. The same applies to the correction unit 110 and the tracking correction unit 107. Alternatively, the display image data generated by the post-processing unit 113 may also be output to the memory 109 and used for tracking. In this case, the tracking pre-processing unit 106 is unnecessary. The tracking correction unit 107 applies image processing to the display image data stored in the tracking memory 109 to improve the accuracy of subject detection.
[0031] The subject detection unit 111 detects areas (subject areas) in the image data (tracking image data) for one frame stored in the tracking memory 109 that are thought to contain a specific subject. The subject detection unit 111 can be implemented, for example, using a multi-class classifier using machine learning. The classifier can be implemented using known machine learning algorithms (models) such as multi-class logistic regression, support vector machines, random forests, and neural networks.
[0032] The subject detection unit 111 acquires information for each detected subject area, including the type (class) of the subject, its position in the image, and its size. If multiple subject areas of the same class are detected, the subject detection unit 111 selects one. The subject detection unit 111 can prioritize individual subject areas based on one or more predetermined conditions, such as size, position, and detection confidence, and select the subject area with the highest priority. For example, areas with larger size and areas closer to the focus detection area can be given higher priority.
[0033] The tracking control unit 112 controls the operation of the tracking unit 114, more specifically, the activation and deactivation of the two tracking units that the tracking unit 114 has. Details of the operation of the tracking control unit 112 will be described later.
[0034] The tracking unit 114 performs tracking processing to estimate the location of the area of the subject to be tracked in the image represented by the tracking image data. The tracking unit 114 outputs the results of the tracking processing (e.g., the position and size of the subject area) to the control unit 103 and the tracking frame superimposition unit 119. If the subject detection unit 111 detects multiple types of subject areas, the tracking unit 114 performs tracking processing for each individual subject area.
[0035] In this embodiment, the tracking unit 114 includes an ML tracking unit 115 (first tracking means) that estimates the position of the subject area using machine learning, and a non-ML tracking unit 116 (second tracking means) that estimates the position of the subject area without using machine learning. Here, the ML tracking unit 115 is assumed to estimate the position of the subject area using deep learning (DL) as an example of machine learning. On the other hand, the non-ML tracking unit 116 is assumed to estimate the position of the subject area using a known method that does not use machine learning. For example, the non-ML tracking unit 116 can estimate the position of the subject area in the current frame based on the similarity with the subject area estimated in past frames. Specifically, the position of the subject area can be estimated based on the sum of the differences in pixel values corresponding to the position (pattern matching), the similarity of the histogram of the pixel color, the similarity of distance information, etc.
[0036] Both the ML tracking unit 115 and the non-ML tracking unit 116 can be implemented by hardware, software, or a combination thereof. However, while ML tracking processing (first tracking processing) yields more accurate results than non-ML tracking processing (second tracking processing), it also places a greater processing load on the ML tracking unit. Therefore, performing ML tracking processing at the same frequency as non-ML tracking processing requires high-performance hardware with high power consumption. Furthermore, if non-ML tracking processing and ML tracking processing are performed using equivalent hardware, the time required for ML tracking processing will be longer than for non-ML tracking processing.
[0037] The ML tracking unit 115 and the non-ML tracking unit 116 can operate in parallel. The tracking control unit 112 appropriately controls the operation of the ML tracking unit 115 and the non-ML tracking unit 116 to suppress power consumption while achieving both tracking responsiveness and accuracy.
[0038] The ML tracking unit 115 can be implemented as a multilayer neural network including convolutional layers. Based on parameter settings determined by prior training, the ML tracking unit 115 automatically extracts feature points and the feature quantities contained within those feature points from the tracking image data. The ML tracking unit 115 also has a function to associate feature points between frames based on parameter settings determined by prior training using a feature point dataset and a feature quantity dataset.
[0039] The ML tracking unit 115 estimates the position of the subject region in the tracking image data to be processed by associating feature points and feature quantities automatically extracted from the tracking image data with the feature points and feature quantities of the subject region to be tracked. The ML tracking unit 115 outputs the estimated position and size of the subject region as the result of the tracking process. The position may be, for example, the center or centroid coordinates of the region, and the size may be, for example, the number of pixels in the horizontal and vertical directions.
[0040] As described above, the non-ML tracking unit 116 estimates the position of the subject area in the tracking image data to be processed based on similarity such as color and brightness. Similar to the ML tracking unit 115, the non-ML tracking unit 116 also outputs the estimated position and size of the subject area as the result of the tracking process. The non-ML tracking unit 116 retains or stores in the tracking memory 109 the information necessary for the next tracking process from the results of the tracking process. For example, if the image of the subject area is used as a template, the template is updated with the image of the subject area obtained as a result of the tracking process.
[0041] The control unit 103 can use the position and size of the subject area output by the tracking unit 114 for, for example, AF processing or AE processing. For example, the control unit 103 can set the focus detection area to include the subject area, or determine exposure conditions so that the subject area is properly exposed.
[0042] The tracking frame superimposition unit 119 has video memory for the display unit 120. The tracking frame superimposition unit 119 stores the display image data output by the post-processing unit 113 in its video memory. The tracking frame superimposition unit 119 also generates a frame-shaped image (tracking frame) based on the size of the subject area output by the tracking unit 114. Furthermore, the tracking frame superimposition unit 119 writes the tracking frame data to the video memory based on the position of the subject area output by the tracking unit 114. As a result, the video memory stores display image data with a tracking frame superimposed on it that indicates the subject area detected by the tracking unit 114.
[0043] Furthermore, the tracking frame superimposition section 119 is not limited to the tracking frame; it can also superimpose images such as icons and numerical values representing the status of the digital camera 101 and information to assist in shooting onto the display image data. For example, image data representing the current exposure conditions (sensitivity, aperture value, shutter speed, etc.), operating mode, remaining capacity of the recording medium (number of photos or time that can be recorded, etc.), and battery level can be superimposed onto the display image data.
[0044] The display unit 120 reads the image data stored in the video memory of the tracking frame superimposition unit 119 and displays the image represented by the read image data on a display device, such as an LCD panel or an organic EL panel. As a result, the display unit 120 displays an image with the tracking frame superimposed (live view image).
[0045] Figure 2 shows an example of image 201 based on display image data output by the post-processing unit 113 and image 202 with the tracking frame superimposed by the tracking frame superimposition unit 119. Note that the shape of the tracking frame is just one example, and it may be a frame without breaks. Also, it is not necessarily limited to a frame shape as long as the range of the rectangular area can be roughly grasped. For example, it may be a pattern consisting of four point-like indicators that represent the four vertices of the rectangular area.
[0046] Next, using the time chart shown in Figure 3, the operation of the digital camera 101 to provide subject tracking functionality while displaying live view will be explained. Times t301 to t307 indicate the start timing of each frame period. At times t301 to t307, the shooting process for each frame is executed, and processing of the image data obtained from the shooting also begins. In other words, the interval between adjacent times t301 to t307 corresponds to one frame period. Although Figure 3 does not show the processing of the pre-processing unit 105, correction unit 110, and post-processing unit 113, the operations of these units are executed in parallel, and the post-processing unit 113 outputs display image data to satisfy the display frame rate. In addition, the tracking pre-processing unit 106 and tracking correction unit 107 also generate tracking image data to satisfy the display frame rate and store it in the tracking memory 109.
[0047] Here, we assume that the frame rate for live view display is 120fps, and therefore the display processing is performed at a cycle of 1 / 120 second (8.33 msec). On the other hand, the autofocus processing (optical control processing and lens drive processing) is performed at a cycle of 1 / 60 second (16.66 msec). Furthermore, the results of ML tracking processing, which yield more accurate results than non-ML tracking processing, are used to set the focus detection area in the autofocus processing.
[0048] The frequency (operation cycle) of the autofocus process may be set so that the ML tracking process can keep up. Alternatively, the processing capacity of the ML tracking unit 115 may be determined to meet the frequency of the autofocus process. However, the frequency of the autofocus process will not exceed the frequency of the display process (display unit frame rate), and generally, the frequency of the autofocus process is lower than the frequency of the display process.
[0049] The tracking control unit 112 determines whether to enable or disable the ML tracking unit 115 and the non-ML tracking unit 116 based on the relationship between the time required for tracking processing by the ML tracking unit 115 and the non-ML tracking unit 116 and the duration of one frame. Here, it is assumed that the time required for tracking processing by the ML tracking unit 115 is longer than the duration of one frame and shorter than the duration of two frames, while the time required for tracking processing by the non-ML tracking unit 116 is shorter than the duration of one frame. The tracking control unit 112 then decides to use the processing result of the ML tracking unit 115 for autofocus processing and the processing result of the non-ML tracking unit 116 for superimposing the tracking frame onto the live view image in display processing. Therefore, the tracking control unit 112 controls the operation of the tracking unit 114 so that the non-ML tracking unit 116 is enabled for each frame and the ML tracking unit 115 is enabled every other frame.
[0050] At time t301, the subject detection unit 111 starts subject detection processing 310 on the tracking image data for the first frame. While the subject detection processing 310 is being completed, the display image data for the first frame is generated and output from the post-processing unit 113 to the tracking frame superimposition unit 119. When the subject detection processing 310 is completed, the tracking frame superimposition unit 119 performs tracking frame superimposition processing 311 according to the detection result. Subsequently, the tracking frame superimposition unit 119 performs display processing 312 to the display unit 120 from time t302. No tracking processing is performed on the first frame.
[0051] At time t302, processing of the tracking image data for the second frame begins. For the second frame, when tracking processing is performed for the first time, the tracking control unit 112 enables both the ML tracking unit 115 and the non-ML tracking unit 116 of the tracking unit 114. Then, the ML tracking unit 115 and the non-ML tracking unit 116 execute the ML tracking process 313 and the non-ML tracking process 314 in parallel for the second frame.
[0052] Since the non-ML tracking process 314 is completed within one frame period, the results of the non-ML tracking process 314 can be used for the tracking frame superposition process of the second frame. When the non-ML tracking process 314 is completed, the non-ML tracking unit 116 outputs the position and size information of the subject area in the second frame image to the tracking frame superposition unit 119 as a processing result. Then, the tracking frame superposition unit 119 starts the tracking frame superposition process 315. Based on the processing results obtained from the tracking unit 114 (in this case, the non-ML tracking unit 116), the tracking frame superposition unit 119 determines the size and position of the tracking frame to be superimposed on the live view image. Then, the tracking frame superposition unit 119 superimposes the image data of the tracking frame onto the display image data of the second frame obtained from the post-processing unit 113. After that, at time t303, the tracking frame superposition unit 119 executes the display process 316 to the display unit 120 based on the display image data of the second frame.
[0053] Furthermore, at time t303, the tracking unit 114 starts non-ML tracking processing 317 by the non-ML tracking unit 116 for the tracking image data of the third frame. At time t303, the ML tracking unit 115 is still performing tracking processing for the second frame, so only the non-ML tracking unit 116 performs tracking processing for the third frame. The non-ML tracking processing 317 and the tracking frame superimposition processing 318 based on the result are the same as for the second frame, so their explanation is omitted.
[0054] Between times t303 and t304, the ML tracking process 313 for the second frame is completed. The ML tracking unit 115 outputs the position and size information of the subject area in the second frame image to the control unit 103 as a processing result. Since the tracking frame superposition process for the second frame is complete, the ML tracking unit 115 outputs the processing result only to the control unit 103.
[0055] When the control unit 103 obtains the processing result from the DL tracking unit 115, it executes optical control processing 320 using the processing result. Here, the control unit 103 uses the processing result from the DL tracking unit 115 to set the focus detection area and determine the exposure conditions. Specifically, the control unit 103 sets the focus detection area to include the subject area based on the position and size of the subject area obtained in the tracking process. For example, if the size of the predetermined focus detection area is smaller than the size of the subject area, the control unit 103 sets the focus detection area inside the subject area. Also, if the size of the focus detection area is larger than the size of the subject area, the control unit 103 sets the focus detection area so that the subject area is located in the center of the focus detection area. Note that the focus detection area including the subject area may be set by other methods.
[0056] Furthermore, the control unit 103 can determine exposure conditions so that the subject area is properly exposed. For example, the average brightness value of the subject area can be used as the Ev value, and a combination of aperture value, shutter speed, and shooting sensitivity can be determined using a program diagram. Alternatively, exposure conditions considering the subject area may be determined by other methods.
[0057] The control unit 103 acquires a signal pair for the set focus detection area from the AF sensor 123 and obtains the defocus amount based on the phase difference of the signal pair. The control unit 103 also controls the exposure period and gain value of the image sensor 104 and the aperture value of the optical system 102 according to the determined exposure conditions (AE processing).
[0058] When the optical control process 320 is completed, the control unit 103 executes a lens drive process 321 based on the defocus amount acquired in the optical control process 320. This adjusts the focus distance of the optical system 102 so that it focuses on the focus detection area set in the optical control process 320. Note that the optical control process 320 is executed across the start time t304 of the 4th frame. Therefore, the AF processing based on the ML tracking process for the tracking image data of the 2nd frame is reflected in the shooting of the 5th frame, which starts at time t305.
[0059] While the optical control process 320 is running, at time t304, the display process 319 based on the tracking frame superimposition process 318 and the capture of the fourth frame are performed. Since the ML tracking unit 115 is available at time t304, the tracking control unit 112 enables both the ML tracking unit 115 and the non-ML tracking unit 116 for the tracking image data of the fourth frame, just as it did for the second frame. The ML tracking unit 115 and the non-ML tracking unit 116 then execute the ML tracking process 322 and the non-ML tracking process 323 in parallel for the fourth frame.
[0060] The non-ML tracking process 323 and the tracking frame superimposition process 324 are completed by time t305. At time t305, the display process 325 based on the tracking frame superimposition process 324 and the capture of the 5th frame are performed. At time t305, the ML tracking unit 115 is performing the tracking process for the 4th frame, so only the non-ML tracking process 326 is performed for the 5th frame, and the ML tracking process is not performed. Once the non-ML tracking process 326 is completed, the tracking frame superimposition process 327 is performed.
[0061] Between time t305 and t306, the ML tracking process 322 is completed and the optical control process 329 is started. The optical control process 329 is executed across time t306.
[0062] At time t306, the display processing 328 based on the tracking frame superimposition processing 327 and the capture of the 6th frame are performed. Since the ML tracking unit 115 is available at time t306, the tracking control unit 112 enables both the ML tracking unit 115 and the non-ML tracking unit 116 for the tracking image data of the 6th frame, just as with the 2nd and 4th frames. The ML tracking unit 115 and the non-ML tracking unit 116 then execute the ML tracking processing 331 and the non-ML tracking processing 332 in parallel for the 6th frame.
[0063] Subsequently, lens drive processing 330 based on optical control processing 329 and tracking frame superposition processing 333 based on non-ML tracking processing 332 are executed, resulting in time t307. Processing continues similarly for subsequent frames.
[0064] Thus, the digital camera 101 of this embodiment achieves both responsiveness of the tracking frame display and accurate AF processing by using an ML tracking unit 115 that performs tracking processing using machine learning and a non-ML tracking unit 116 that performs tracking processing without machine learning. In addition, power consumption can be saved by reducing the operating frequency of the ML tracking unit 115 compared to that of the non-ML tracking unit 116.
[0065] (Variation 1) Figure 4 shows another example of a time chart relating to the operation of the digital camera 101 to provide subject tracking functionality while displaying live view. This modified example corresponds to the example shown in Figure 3, where the frequency of autofocus processing is reduced and executed at a cycle of 1 / 30 second (33.33 msec). On the other hand, the frequency of display processing is at a cycle of 1 / 120 second (8.33 msec). In this modified example as well, the post-processing unit 113 and the tracking correction unit 107 generate display image data and tracking image data to satisfy the display frame rate.
[0066] The operation shown in Figure 4 may be performed, for example, when the operating mode of the digital camera 101 is set to an operating mode that reduces power consumption. However, it may also be performed under other conditions. For example, it may be performed when the control unit 103 detects that a predetermined condition for reducing the operating frequency of the ML tracking unit 115 has been met. The condition for reducing the operating frequency of the ML tracking unit 115 may be, for example, when the remaining capacity of the battery supplying power to the digital camera 101 falls to a predetermined threshold, or when the movement of the subject being tracked in the most recent predetermined time is below a threshold.
[0067] In this modified example, times t401 to t410 indicate the start timing of each frame period. Furthermore, the process with the same name as in Figure 3 is identical to the process described in Figure 3, and therefore its explanation is omitted. The processing speeds of the ML tracking unit 115 and the non-ML tracking unit 116 are also the same as in Figure 3.
[0068] Processes 410-423 in Figure 4 are the same as processes 310-321 and 323-324 in Figure 3. In this modified example, the ML tracking process, optical control process, and lens drive process are not executed at time t404 (4th frame). Since the autofocus process is performed once every 4 frames, the first ML tracking process 413 is executed at time t402 (2nd frame), and the second ML tracking process 428 is executed at time t406 (6th frame).
[0069] Then, as autofocus processing for the 9th frame, which is captured at time t409, the optical control processing 435 and the lens drive processing 436 are executed following the second ML tracking processing 428.
[0070] Meanwhile, for each frame from the 5th to the 9th frame, non-ML tracking processing (425, 429, 432, 437, 440), tracking frame superposition processing (426, 430, 433, 438, 441), and display processing (424, 427, 431, 434, 439) are executed. As a result, the tracking frame is superimposed on all live view images displayed at a frame rate of 120fps, allowing the user to check the tracking status in real time.
[0071] In this way, by executing ML tracking processing at the autofocus operation cycle and non-ML tracking processing at the display operation cycle (display frame rate), it is possible to reduce power consumption by suppressing the frequency of ML tracking processing while maintaining the tracking performance of the tracking frame.
[0072] (Modification 2) Figure 5 shows yet another example of a time chart relating to the operation of the digital camera 101 to provide subject tracking functionality while displaying live view. In this modified example, the execution frequency of autofocus processing and tracking frame superposition processing is the same as in the example in Figure 3. This modified example differs from the example in Figure 3 in that the results of ML tracking processing are used for tracking frame superposition processing.
[0073] Processes 510-533 in Figure 5 are identical to processes 310-333 in Figure 3, so the explanation of each process is omitted. This modified example differs from the example shown in Figure 3 in that, when available, the tracking frame superposition process is performed using the results of the ML tracking process. Figure 5 shows an example in which the results of the ML tracking process are used for the tracking frame superposition process performed during the frame period in which the ML tracking process is completed.
[0074] Therefore, the tracking frame superposition process 518, which is executed during the third frame period (time t503~t504) when the ML tracking process 513 is completed, uses the result of the ML tracking process 513, not the result of the non-ML tracking process 517. Similarly, the tracking frame superposition process 527, which is executed during the fifth frame period (time t505~t506) when the ML tracking process 522 is completed, uses the result of the ML tracking process 522, not the result of the non-ML tracking process 526.
[0075] If the tracking frame superposition process cannot be completed between the end of the ML tracking process and the start of the next frame period, the results of the ML tracking process may be used for the tracking frame superposition process executed in the frame period following the frame period in which the ML tracking process was completed. In this case, the results of the non-ML tracking process executed in the same frame period will be used for the tracking frame superposition process executed in the frame period in which the ML tracking process was completed.
[0076] ML tracking processing is more accurate than non-ML tracking processing. Furthermore, the results of ML tracking processing are used in autofocus processing. Therefore, by using the results of ML tracking processing for tracking frame superposition processing, a tracking frame with high consistency with the focus detection area can be displayed. For frames where ML tracking processing results are unavailable, the results of non-ML tracking processing are used to superimpose the tracking frame. This enables the display of a tracking frame with high tracking accuracy.
[0077] Furthermore, the time required for ML tracking processing and tracking frame superposition processing can be statistically predicted. Therefore, it is possible to know in advance which frame's tracking frame superposition processing will utilize ML tracking processing. Accordingly, the tracking control unit 112 may disable the non-ML tracking unit 116 for frame periods in which the results of ML tracking processing are used for tracking frame superposition processing. This reduces power consumption.
[0078] The operation of the digital camera 101 regarding the subject tracking process described in Figures 3 to 5 will be further explained using the flowcharts shown in Figures 6 and 7. The subject tracking process is performed on images captured in a time series (video or still images taken in burst mode). Whether or not the subject tracking process is enabled or disabled may be explicitly set by the user. Alternatively, it may be automatically enabled depending on settings such as the shooting mode.
[0079] In step S601, the digital camera 101 captures one frame and generates tracking image data. Figure 7 is a flowchart detailing step S601.
[0080] In S701, the control unit 103 controls the operation of the image sensor 104 according to the exposure conditions determined by AE processing, causing the image sensor 104 to capture one frame. The image sensor 104 outputs the analog image signal obtained by the capture to the pre-processing unit 105 and the tracking pre-processing unit 106.
[0081] In S702, the tracking preprocessor 106 generates image data from the analog image signal as described earlier. In S703, the tracking preprocessor 106 writes the generated image data to the tracking memory 109. In S704, the tracking correction unit 107 applies correction processing to the image data as described earlier. In S705, the tracking correction unit 107 determines whether all applicable correction processes have been applied. If the tracking correction unit 107 determines that there are still unapplied correction processes, it repeatedly executes S703 and S704 to apply the unapplied image processing. On the other hand, if the tracking correction unit 107 determines that all applicable correction processes have been applied, it writes the corrected image data as tracking image data to the tracking memory 109 and terminates the tracking image data generation process.
[0082] In parallel with the generation of tracking image data, the pre-processing unit 105, correction unit 110, and post-processing unit 113 generate display image data (and further recording image data depending on the shooting application) from the analog image signal obtained in S701.
[0083] Returning to Figure 6, at S602, the subject detection unit 111 applies a subject detection process to the tracking image data stored in the tracking memory 109 to detect the region of a predetermined type of subject. The subject detection unit 111 writes the position and size of the detected subject region, etc., to the tracking memory 109 as the result of subject detection.
[0084] In S603, the tracking control unit 112 determines whether the current frame is a frame in which autofocus (AF) tracking processing should be performed. The frequency of performing AF tracking processing (how often it should be performed) is predetermined. Therefore, the tracking control unit 112 can determine whether the current frame is a frame in which AF tracking processing should be performed based, for example, on the count value of a counter that increases with each frame. For example, in the example shown in Figure 3, where AF tracking processing is performed every two frames, the tracking control unit 112 can determine that the current frame is a frame in which AF tracking processing should be performed if the count value is even.
[0085] The tracking control unit 112 executes S604 if it determines that the current frame is a frame for which tracking processing for AF should be performed, and executes S605 if it does not determine that the current frame is a frame for which tracking processing for AF should be performed. In S604, the tracking control unit 112 enables both the ML tracking unit 115 and the non-ML tracking unit 116 and sets or notifies the tracking unit 114. In S605, the tracking control unit 112 disables the ML tracking unit 115 and enables the non-ML tracking unit 116, and sets or notifies the tracking unit 114.
[0086] In S606, the tracking unit 114, with the help of the ML tracking unit 115 and the non-ML tracking unit 116, performs ML tracking processing and non-ML tracking processing on the tracking image data of the current frame. In S608, the control unit 103 performs optical control processing based on the results of the ML tracking processing performed in S606, and updates the focus detection area. Then, the control unit 103 performs lens driving processing and adjusts the focus distance of the optical system 102 so that it focuses on the updated focus detection area.
[0087] In S607, the tracking unit 114 performs non-ML tracking processing on the tracking image data of the current frame using the non-ML tracking unit 116.
[0088] In S609, the tracking frame superimposing section 119 updates the position and size of the tracking frame based on the results of the non-ML tracking process in S606 or S607, or the results of the ML tracking process in S606.
[0089] In S610, the tracking frame superimposition unit 119 generates updated tracking frame image data and superimposes it onto the display image data output by the post-processing unit 113 for display on the display unit 120. This completes the subject tracking process for one frame. Thereafter, the digital camera 101 repeatedly performs the same process for each frame.
[0090] As described above, according to this embodiment, in an image processing apparatus having an ML tracking unit that performs tracking processing using machine learning and a non-ML tracking unit that performs tracking processing without using machine learning, the operating frequency of the ML tracking unit is made lower than that of the non-ML tracking unit. Therefore, the increase in power consumption due to operating the ML tracking unit can be suppressed. Furthermore, by using the processing results of the ML tracking unit in the autofocus processing, autofocus based on highly accurate tracking results can be realized.
[0091] (Other embodiments) In the examples shown in Figures 3 to 7, the frequency of ML tracking processing or autofocus processing was determined based on the assumption that ML tracking processing was used for autofocus processing, in order to facilitate explanation and understanding. However, the essence of the present invention is to have a non-ML tracking unit in addition to the ML tracking unit, and to use the results of the non-ML tracking processing for frames in which the ML tracking processing cannot keep up. This eliminates the need to complete the ML tracking processing within a single frame period, thereby reducing power consumption. Furthermore, for frames in which the results of ML tracking processing cannot be used, processing using the results of the non-ML tracking processing becomes possible, enabling, for example, tracking frame superposition processing that has good tracking performance for high frame rate displays.
[0092] Therefore, it is also possible to configure the system so that both the ML tracking unit and the non-ML tracking unit are always enabled, and for frames where ML tracking processing is not completed in time, the results of the non-ML tracking process are used, and for frames where the results of the ML tracking process are available, the results of the ML tracking process are used. In this case, the process using the results of the non-ML tracking process and the process using the results of the ML tracking process may be the same or different.
[0093] The present invention can also be realized by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit (e.g., an ASIC) that implements one or more functions.
[0094] This embodiment includes the following image processing apparatus, imaging apparatus, image processing method, and program. (Item 1) A first tracking means that applies a first tracking process using machine learning to an image, A second tracking means that applies a second tracking process that does not use machine learning to an image, The system includes control means for controlling the operation of the first tracking means and the second tracking means, The image processing apparatus is characterized in that the control means controls the operating frequency of the first tracking means to be lower than the operating frequency of the second tracking means. (Item 2) The image processing apparatus according to item 1, characterized in that the aforementioned image is an image for display, and the result of the second tracking process is used for live view display. (Item 3) The image processing apparatus according to item 2, characterized in that the result of the second tracking process is used to superimpose an image representing the tracked subject onto the live view display. (Item 4) The image processing apparatus according to item 2 or 3, characterized in that, if the result of the first tracking process can be used, the result of the first tracking process is used for the live view display. (Item 5) The image processing apparatus according to any one of items 1 to 4, characterized in that the image is a frame of a video, the second tracking means applies the second tracking process to each frame, and the first tracking means applies the first tracking process to multiple frames at a time. (Item 6) The image processing apparatus according to item 5, characterized in that the second tracking process is not applied to frames for which the results of the first tracking process can be used. (Item 7) The image processing apparatus according to item 5 or 6, characterized in that the number of frames when the image processing apparatus satisfies a predetermined condition is greater than the number of frames when the image processing apparatus satisfies a predetermined condition. (Item 8) The image processing apparatus according to item 7, characterized in that the predetermined condition is that the image processing apparatus is set to an operating mode that reduces power consumption. (Item 9) The image processing apparatus according to any one of items 1 to 8, characterized in that the operating frequency of the first tracking means is based on the operating frequency of a process that uses the results of the first tracking process. (Item 10) The image processing apparatus according to item 9, characterized in that the process using the result of the first tracking process is an autofocus process. (Item 11) The image processing apparatus according to item 10, characterized in that the autofocus processing uses a focus detection region set according to the result of the first tracking processing. (Item 12) The image processing apparatus according to any one of items 1 to 11, characterized in that the machine learning is deep learning. (Item 13) The image processing apparatus according to any one of items 1 to 12, characterized in that the second tracking process is based on pattern matching. (Item 14) An acquisition means for acquiring an image having a predetermined frame rate, A first tracking means that applies a first tracking process using machine learning to the aforementioned image, A second tracking means that applies a second tracking process that does not use machine learning to the aforementioned image, The system includes processing means for performing a process using the results of the first tracking process or the second tracking process, The time required for the first tracking process is longer than the duration of one frame, and the time required for the second tracking process is shorter than the duration of one frame. The processing means is characterized in that, for frames where the result of the first tracking process cannot be used, the result of the second tracking process is used. (Item 15) Image sensor and An image processing apparatus according to any one of items 1 to 14 that uses the image obtained from the image sensor, An imaging device characterized by having the following features. (Item 16) A first tracking means that applies a first tracking process using machine learning to an image, An image processing method performed by an image processing apparatus having a second tracking means that applies a second tracking process that does not use machine learning to an image, An image processing method characterized by controlling the operating frequency of the first tracking means to be lower than the operating frequency of the second tracking means. (Item 17) A program for causing a computer to function as one of the means of an image processing device described in any one of items 1 through 14.
[0095] The present invention is not limited to the embodiments described above, and various modifications and variations are possible without departing from the spirit and scope of the invention. Accordingly, claims are attached to disclose the scope of the invention. [Explanation of symbols]
[0096] 103...Control unit, 104...Image sensor, 106...Pre-processing unit for tracking, 107...Correction unit for tracking, 109...Memory for tracking, 111...Subject detection unit, 112...Tracking control unit, 114...Tracking unit, 115...ML tracking unit, 116...Non-ML tracking unit, 119...Tracking frame superimposition unit, 120...Display unit
Claims
1. a first tracking means for applying a first tracking process using machine learning to an image; a second tracking means for applying a second tracking process that does not use machine learning to an image; a control means for controlling the operation of the first tracking means and the second tracking means, The image processing device according to claim 1, wherein the control means controls the first tracking means so that the operation frequency of the first tracking means is lower than the operation frequency of the second tracking means.
2. The image processing device according to claim 1 , wherein the image is an image for display, and the result of the second tracking process is used for live view display.
3. 3. The image processing device according to claim 2, wherein the result of the second tracking process is used to superimpose an image representing the tracked subject on the live view display.
4. 3. The image processing device according to claim 2, wherein, when a result of the first tracking process can be used, the result of the first tracking process is used for the live view display.
5. 2. The image processing device according to claim 1, wherein the images are frames of a video, the second tracking means applies the second tracking process to each frame, and the first tracking means applies the first tracking process to every multiple frames.
6. The image processing device according to claim 5 , wherein the second tracking process is not applied to a frame for which the result of the first tracking process can be used.
7. 6. The image processing device according to claim 5, wherein the number of the plurality of frames when the image processing device satisfies a predetermined condition is greater than the number of the plurality of frames when the image processing device satisfies a predetermined condition.
8. 8. The image processing device according to claim 7, wherein the predetermined condition is that the image processing device is set to an operation mode that reduces power consumption.
9. 2. The image processing device according to claim 1, wherein the operation frequency of said first tracking means is based on the operation frequency of a process that uses the result of said first tracking process.
10. 10. The image processing device according to claim 9, wherein the process using the result of the first tracking process is an autofocus process.
11. 11. The image processing device according to claim 10, wherein the autofocus process uses a focus detection area set in accordance with the result of the first tracking process.
12. The image processing device according to claim 1 , wherein the machine learning is deep learning.
13. 2. The image processing device according to claim 1, wherein the second tracking process is based on pattern matching.
14. an acquisition means for acquiring an image having a predetermined frame rate; a first tracking means for applying a first tracking process using machine learning to the image; A second tracking means for applying a second tracking process that does not use machine learning to the image; processing means for executing a process using a result of the first tracking process or the second tracking process; a time required for the first tracking process is longer than one frame period, and a time required for the second tracking process is shorter than one frame period; The image processing device according to claim 1, wherein the processing means uses the result of the second tracking process for frames for which the result of the first tracking process cannot be used.
15. An imaging element; an image processing device according to any one of claims 1 to 14, which uses an image obtained by the imaging element; An imaging device comprising:
16. a first tracking means for applying a first tracking process using machine learning to an image; and a second tracking means for applying a second tracking process that does not use machine learning to an image, 11. An image processing method comprising: controlling an operation frequency of said first tracking means to be lower than an operation frequency of said second tracking means.
17. A program for causing a computer to function as each of the means included in the image processing device according to any one of claims 1 to 14.