Image processing apparatus and control method thereof
The image processing apparatus employs dual tracking methods to manage power consumption by switching between deep learning and non-deep learning processes based on feature points, addressing the inefficiencies of existing neural network-based tracking systems.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- CANON KK
- Filing Date
- 2021-12-10
- Publication Date
- 2026-05-11
AI Technical Summary
Existing image processing technologies for subject tracking, particularly those using neural networks, face high computational loads and power consumption, especially in live view display, leading to battery drain and inefficiency.
An image processing apparatus with dual tracking mechanisms: a first tracking method using deep learning (DL) and a second with lower computational load, controlled by a switching mechanism based on feature points detected in the image, enabling/disabling these methods to optimize power usage.
The system achieves effective subject tracking while significantly reducing power consumption by dynamically switching between DL and non-DL tracking methods, ensuring good performance without excessive energy use.
Smart Images

Figure 0007856417000003 
Figure 0007856417000004 
Figure 0007856417000005
Abstract
Description
Technical Field
[0001] The present invention relates to an image processing apparatus for subject tracking processing and a control method thereof.
Background Art
[0002] Some imaging devices such as digital cameras have a function of tracking a feature area (subject tracking function) by applying detection of a feature area such as a face area over time. Also, a device that tracks a subject using a learned neural network is known (Patent Document 1).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In technologies for subject recognition and subject tracking using images, the accuracy of subject tracking may be improved by using machine learning (neural network, deep learning) compared to using correlations and similarities between image regions. However, processing using a neural network has a large amount of computation, requires a high-speed processor or a large-scale circuit, and thus has a problem of high power consumption. For example, when applying subject tracking using a neural network to a moving image for live view display, battery consumption due to live view display becomes a problem. Also, even among subject tracking using a circuit of a machine learning model, there may be a difference in computation load and power consumption depending on the learning model.
[0005] The present invention has been made in view of the above problems, and an object thereof is to provide an image processing apparatus and an image processing method having a subject tracking function that realizes good performance while suppressing power consumption. [Means for solving the problem]
[0006] To solve the above problems, the image processing apparatus of the present invention is characterized by comprising: a first tracking means that performs subject tracking using an image acquired by an imaging means; a second tracking means that performs subject tracking using an image acquired by the imaging means and has a lower computational load than the first tracking means; and a control means that switches between enabling both the first tracking means and the second tracking means or disabling one of them based on feature points detected from the image acquired by the imaging means. [Effects of the Invention]
[0007] According to the present invention, it is possible to realize a subject tracking function that achieves good performance while suppressing power consumption. [Brief explanation of the drawing]
[0008] [Figure 1] Block diagram showing an example of the functional configuration of the imaging device according to the first embodiment. [Figure 2] Operation flow diagram of the tracking control unit 113 in the imaging device according to the first embodiment. [Figure 3] This figure shows the live view display in the subject tracking process according to the first embodiment. [Figure 4] Operational flowchart of the control unit 102 in the imaging device according to the second embodiment. [Figure 5] Table showing the relationship between the shooting scene and the operating modes of the detection unit 110 and the tracking unit 105 according to the second embodiment. [Figure 6] Table showing the operating modes of the detection unit 110 and the tracking unit 105 according to the second embodiment. [Figure 7] Operation flow diagram of the control unit 102 in the third embodiment [Figure 8] Flowchart of the feature point detection process performed by the feature point detection unit 201 of the third embodiment. [Modes for carrying out the invention]
[0009] Preferred embodiments of the present invention will be described below with reference to the attached drawings.
[0010] (First embodiment) The present invention will be described in detail below with reference to the attached drawings, based on exemplary embodiments thereof. Note that the following embodiments do not limit the invention to the claims. Furthermore, while multiple features are described in the embodiments, not all of them are essential to the invention, and the multiple features may be combined arbitrarily. In addition, in the attached drawings, the same or similar configurations are given the same reference numeral, and redundant descriptions are omitted.
[0011] In the following embodiments, the present invention will be described in relation to cases where it is implemented using an imaging device such as a digital camera. However, the present invention can also be implemented with any electronic device having an imaging function. Such electronic devices include computer equipment (personal computers, tablet computers, media players, PDAs, etc.), mobile phones, smartphones, game consoles, robots, drones, and dashcams. These are examples, and the present invention can also be implemented with other electronic devices.
[0012] Figure 1 is a block diagram showing an example of the functional configuration of an imaging device 100 as an example of an image processing device according to the first embodiment.
[0013] The optical system 101 has multiple lenses, including a movable lens such as a focusing lens, and forms a clear image of the shooting range on the image-forming plane of the image sensor 103.
[0014] The control unit 102 has a CPU and, for example, loads a program stored in ROM 123 into RAM 122 and executes it. The control unit 102 realizes the functions of the imaging device 100 by controlling the operation of each functional block. ROM 123 is, for example, a rewritable non-volatile memory that stores programs that the CPU of the control unit 102 can execute, setting values, GUI data, etc. RAM 122 is a system memory used to load programs executed by the CPU of the control unit 102 and to save values necessary during program execution. Although not shown in Figure 1, the control unit 102 is connected to each functional block in a communicative manner.
[0015] The image sensor 103 may be, for example, a CMOS image sensor having a primary color Bayer array color filter. The image sensor 103 has multiple pixels arranged in two dimensions, each having a photoelectric conversion area. The image sensor 103 converts the optical image formed by the optical system 101 into a group of electrical signals (analog image signals) using the multiple pixels. The analog image signals are converted into digital image signals (image data) by an A / D converter in the image sensor 103 and output. The A / D converter may be provided outside the image sensor 103.
[0016] The evaluation value generation unit 124 generates signals and evaluation values used for autofocus detection (AF) and calculates evaluation values used for automatic exposure control (AE) from image data obtained from the image sensor 103. The evaluation value generation unit 124 outputs the generated signals and evaluation values to the control unit 102. Based on the signals and evaluation values obtained from the evaluation value generation unit 124, the control unit 102 controls the focus lens position of the optical system 101 and determines shooting conditions (exposure time, aperture value, ISO sensitivity, etc.). The evaluation value generation unit 124 may also generate signals and evaluation values from display image data generated by the post-processing unit 114, which will be described later.
[0017] The first preprocessing unit 104 applies color interpolation processing to the image data obtained from the image sensor 103. The color interpolation processing, also called demosaicing processing, etc., is a process that makes each pixel data constituting the image data have values of R component, G component, and B component. Further, the first preprocessing unit 104 may apply reduction processing for reducing the number of pixels as necessary. The first preprocessing unit 104 stores the image data to which the processing has been applied in the display memory 107.
[0018] The first image correction unit 109 applies correction processing such as white balance correction processing and shading correction processing, and conversion processing from RGB format to YUV format, etc. to the image data stored in the display memory 107. Note that when applying the correction processing, the first image correction unit 109 may use image data of one or more frames different from the frame to be processed among the image data stored in the display memory 107. The first image correction unit 109 can use, for example, image data of frames before and / or after in time series from the frame to be processed for the correction processing. The first image correction unit 109 outputs the image data to which the processing has been applied to the post-processing unit 114.
[0019] The post-processing unit 114 generates recording image data and display image data from the image data supplied from the first image correction unit 109. The post-processing unit 114, for example, applies encoding processing to the image data and generates a data file for storing the encoded image data as the recording image data. The post-processing unit 114 supplies the recording image data to the recording unit 118.
[0020] Also, the post-processing unit 114 generates display image data for display on the display unit 121 from the image data supplied from the first image correction unit 109. The display image data has a size corresponding to the display size on the display unit 121. The post-processing unit 114 supplies the display image data to the information superimposing unit 120.
[0021] The recording unit 118 records the recording image data converted by the post-processing unit 114 onto the recording medium 119. The recording medium 119 may be, for example, a semiconductor memory card or an internal non-volatile memory.
[0022] The second preprocessor 105 applies color interpolation processing to the image data output by the image sensor 103. The second preprocessor 105 stores the processed image data in the tracking memory 108. The tracking memory 108 and the display memory 107 may be implemented as separate address spaces within the same memory space. The second preprocessor 105 may also apply a reduction process to reduce the number of pixels as needed to reduce the processing load. Although the first preprocessor 104 and the second preprocessor 105 are described here as separate functional blocks, a configuration using a common preprocessor is also possible.
[0023] The second image correction unit 106 applies correction processing such as white balance correction processing and shading correction processing, as well as conversion processing from RGB format to YUV format, to the image data stored in the tracking memory 108. The second image correction unit 106 may also apply image processing suitable for subject detection processing to the image data. For example, if the representative brightness of the image data (e.g., the average brightness of all pixels) is below a predetermined threshold, the second image correction unit 106 may multiply the entire image data by a certain coefficient (gain) so that the representative brightness becomes above the threshold.
[0024] Furthermore, when applying the correction process, the second image correction unit 106 may use one or more image data frames from the tracking memory 108 that are different from the frame being processed. For example, the second image correction unit 106 may use image data from frames that are chronologically earlier and / or later than the frame being processed for the correction process. The second image correction unit 106 stores the processed image data in the tracking memory 108.
[0025] Furthermore, functional blocks related to the subject tracking function, such as the second preprocessing unit 105 and the second image correction unit 106, do not need to operate when the subject tracking function is not being performed. The image data to which the subject tracking function is applied is video data captured for live view display or recording. The video data has a predetermined frame rate, such as 30fps, 60fps, or 120fps.
[0026] The detection unit 110 detects one or more predetermined candidate subject regions (candidate regions) from the image data of one frame. The detection unit 110 also associates each detected region with its position and size within the frame, an object class indicating the type of candidate subject (car, airplane, bird, insect, human body, head, pupil, cat, dog, etc.), and its confidence level. It also counts the number of detected regions for each object class.
[0027] The detection unit 110 can detect candidate regions using known techniques for detecting feature regions such as the facial regions of people and animals. For example, the detection unit 110 may be configured as a class classifier trained using training data. There are no particular restrictions on the classification algorithm. The detection unit 110 can be realized by training a classifier that implements multi-class logistic regression, support vector machines, random forests, neural networks, etc. The detection unit 110 stores the detection results in the tracking memory 108.
[0028] The target determination unit 111 determines the subject area to be tracked (main subject area) from the candidate areas detected by the detection unit 110 or the tracking results from the tracking unit 115, which will be described later. The subject area to be tracked can be determined based on a priority order assigned in advance for each item included in the detection results, such as object class and area size. Specifically, the sum of the priorities may be calculated for each candidate area, and the candidate area with the smallest sum may be determined as the subject area to be tracked. Alternatively, among the candidate areas belonging to a specific object class, the candidate area closest to the center or focus detection area of the image, or the largest candidate area, may be determined as the subject area to be tracked. The target determination unit 111 stores information identifying the determined subject area in the tracking memory 108.
[0029] The difficulty determination unit 112 calculates a difficulty score, which is an evaluation value indicating the difficulty of tracking, for the subject area to be tracked, as determined by the target determination unit 111. For example, the difficulty determination unit 112 can calculate the difficulty score by considering one or more factors that affect the difficulty of tracking. Examples of factors that affect the difficulty of tracking include, but are not limited to, the size of the subject area, the object class (type) of the subject, the total number of areas belonging to the same object class, and the position within the image. A specific example of the method for calculating the difficulty score will be described later. The difficulty determination unit 112 outputs the calculated difficulty score to the tracking control unit 113.
[0030] The tracking control unit 113 determines whether to enable or disable each of the multiple tracking units of the tracking unit 115 based on the difficulty score calculated by the difficulty determination unit 112. In this embodiment, the tracking unit 115 has multiple tracking units with different computational loads and tracking accuracy. Specifically, the tracking unit 115 has a DL tracking unit 116 that performs subject tracking using deep learning (DL) and a non-DL tracking unit 117 that performs subject tracking without using DL. The DL tracking unit 116 has higher processing accuracy than the non-DL tracking unit 117, but has a higher computational load than the non-DL tracking unit 117.
[0031] In this case, the tracking control unit 113 decides whether to enable or disable the DL tracking unit 116 and the non-DL tracking unit 117, respectively. The tracking control unit 113 also decides the operating frequency for the enabled tracking units. The operating frequency is the frequency (fps) at which the tracking process is applied.
[0032] The tracking unit 115 estimates the subject area to be tracked from the image data of the frame to be processed (the current frame) stored in the tracking memory 108, and determines the position and size of the estimated subject area within the frame as the tracking result. For example, the tracking unit 115 estimates the subject area to be tracked within the current frame using the image data of the current frame and the image data of a past frame taken before the current frame (for example, the previous frame). The tracking unit 115 outputs the tracking result to the information superposition unit 120.
[0033] Here, the tracking unit 115 estimates the region within the frame to be processed that corresponds to the region of the subject being tracked in past frames. In other words, the region of the subject being tracked determined by the target determination unit 111 for the frame to be processed is not the region of the subject being tracked in the tracking process for the frame to be processed. The region of the subject being tracked in the tracking process for the frame to be processed is the region of the subject being tracked in past frames. The region of the subject being tracked determined by the target determination unit 111 for the frame to be processed is used in the tracking process for the next frame if the subject being tracked switches to a different subject.
[0034] The tracking unit 115 includes a DL tracking unit 116 that performs subject tracking using deep learning (DL), and a non-DL tracking unit 117 that performs subject tracking without using DL. The tracking unit, which has been enabled by the tracking control unit 113, outputs tracking results at the operating frequency set by the tracking control unit 113.
[0035] The DL tracking unit 116 estimates the position and size of the subject region to be tracked using a pre-trained multilayer neural network including convolutional layers. More specifically, the DL tracking unit 116 has the function of extracting feature points and feature quantities contained in the subject region for each possible object class, and the function of associating the extracted feature points between frames. Therefore, the DL tracking unit 116 can estimate the position and size of the subject region to be tracked in the current frame from the feature points of the current frame that are associated with the feature points of the subject region to be tracked in past frames.
[0036] The DL tracking unit 116 outputs the position, size, and confidence score of the subject area to be tracked estimated for the current frame. The confidence score indicates the reliability of the correspondence between feature points between frames, that is, the reliability of the estimation result of the subject area to be tracked. If the confidence score indicates that the reliability of the correspondence between feature points between frames is low, it indicates that the subject area estimated in the current frame may be a different subject area from the subject area to be tracked in past frames.
[0037] On the other hand, the non-DL tracking unit 117 estimates the subject area to be tracked in the current frame using a method that does not employ deep learning. Here, the non-DL tracking unit 117 estimates the subject area to be tracked based on the similarity of color composition. However, other methods may be used, such as pattern matching using the subject area to be tracked in past frames as a template. The non-DL tracking unit 117 outputs the position, size, and confidence score of the subject area to be tracked estimated for the current frame.
[0038] Here, we will explain the similarity of color composition. For the sake of clarity and ease of explanation, we will assume that the shape and size of the subject area being tracked are the same in past frames and current frames. Furthermore, we will assume that the image data has a depth of 8 bits (values from 0 to 255) for each RGB color component.
[0039] The non-DL tracking unit 117 divides the range of possible values (0 to 255) for a certain color component (for example, the R component) into multiple regions. Then, the non-DL tracking unit 117 classifies the pixels included in the subject region to be tracked according to the region to which the R component value belongs (frequency for each value range), and uses this result as the color composition of the subject region to be tracked.
[0040] As the simplest example, let's assume that the range of possible values for the R component (0 to 255) is divided into Red1 (0 to 127) and Red2 (128 to 255). Let's also assume that the color composition of the subject area being tracked in past frames was 50 pixels for Red1 and 70 pixels for Red2. And let's assume that the color composition of the subject area being tracked in the current frame is 45 pixels for Red1 and 75 pixels for Red2.
[0041] In this case, the non-DL tracking unit 117 can calculate a score representing the similarity of the color configuration (similarity score) based on the difference in the number of pixels classified within the same value range, as follows. Similarity score = |50-45| + |70-75| = 10 Assuming that the color composition of the subject area being tracked in the current frame is 10 pixels for Red1 and 110 pixels for Red2, the similarity score would be: Similarity score = |50-10| + |70-110| = 80 Thus, the lower the similarity of the color compositions, the higher the similarity score. Conversely, a smaller similarity score indicates a higher similarity of the color compositions.
[0042] The information overlay unit 120 generates a tracking frame image based on the size of the subject area included in the tracking result output by the tracking unit 115. For example, the tracking frame image may be a frame-shaped image representing a rectangular outline circumscribing the subject area. The information overlay unit 120 then generates composite image data by overlaying the tracking frame image onto the display image data output by the post-processing unit 114 so that the tracking frame is displayed at the position of the subject area included in the tracking result. The information overlay unit 120 may also generate images representing the current settings and status of the imaging device 100, and overlay these images onto the display image data output by the post-processing unit 114 so that they are displayed at predetermined positions. The information overlay unit 120 outputs the composite image data to the display unit 121.
[0043] The display unit 121 may be, for example, a liquid crystal display or an organic EL display. The display unit 121 displays an image based on the composite image data output by the information superposition unit 120. In this manner, a live view display for one frame is performed.
[0044] The evaluation value generation unit 124 generates signals and evaluation values used for autofocus detection (AF) and calculates evaluation values (luminance information) used for automatic exposure control (AE) from image data obtained from the image sensor 103. The luminance information is generated by color conversion from the integrated values (red, blue, green) obtained by integrating each color filter pixel (red, blue, green). Note that other methods may be used to generate the luminance information. In addition, evaluation values (integrated values for each color (red, blue, green)) used for automatic white balance (AWB) are calculated in the same way as when generating the luminance information. The control unit 102 identifies the light source from these integrated values for each color and calculates a correction value for the pixels so that white objects become white. White balance is performed by multiplying each pixel by this correction value in the first image correction unit 109 and the second image correction unit 106, which will be described later. In addition, evaluation values (motion vector information) used for camera shake detection for image stabilization are calculated from reference image data using two or more image data images to obtain a motion vector. The evaluation value generation unit 124 outputs the generated signal and evaluation value to the control unit 102. Based on the signal and evaluation value obtained from the evaluation value generation unit 124, the control unit 102 controls the focus lens position of the optical system 101 and determines the shooting conditions (exposure time, aperture value, ISO sensitivity, etc.). The evaluation value generation unit 124 may also generate the signal and evaluation value from display image data generated by the post-processing unit 114, which will be described later.
[0045] The selection unit 125 adopts either the tracking result of the DL tracking unit 116 or the non-DL tracking unit 117 based on the confidence score output by the DL tracking unit 116 and the similarity score output by the non-DL tracking unit 117. For example, if the confidence score is below a predetermined confidence score threshold and the similarity score is below a predetermined similarity score threshold, the selection unit 125 adopts the tracking result of the non-DL tracking unit 117; otherwise, it adopts the tracking result of the DL tracking unit 116. The selection unit 125 outputs the adopted tracking result to the information superposition unit 120 and the control unit 102.
[0046] In this example, the decision of whether to use the tracking results from the DL tracking unit 116 or the non-DL tracking unit 117 was made based on the confidence score and similarity score. However, other methods may be used to make this decision. For example, since the accuracy of the DL tracking unit 116 tends to be higher than that of the non-DL tracking unit 117, the tracking results of the DL tracking unit 116 may be given priority. Specifically, if tracking results from the DL tracking unit 116 are available, those results may be used; otherwise, the tracking results of the non-DL tracking unit 117 may be used.
[0047] The image sensor motion detection unit 126 detects the movement of the image sensor 100 itself and is composed of a gyro sensor and the like. The image sensor motion detection unit 126 outputs the detected motion information of the image sensor to the control unit 102. Based on the motion information of the image sensor, the control unit 102 detects camera shake and the image sensor swinging in a certain direction and makes a determination of whether it is panning. In addition, the accuracy of the panning determination can be improved by combining the result of the image sensor motion detection unit 126 and the motion vector of the evaluation value generation unit 124, and observing that the image sensor is swinging in a certain direction but the motion vector of the subject is almost zero.
[0048] Next, using Figure 2, the operation flow of the tracking control unit 113 for tracking a subject when the imaging device 100 performs an imaging operation will be explained. In this embodiment, DL tracking and non-DL tracking are controlled depending on whether or not it is a scene for panning photography and whether or not it is a scene with low brightness. However, only one of these scenes may be used for determination, or the system may be configured to perform other scene determinations and then determine whether to perform DL tracking or non-DL tracking.
[0049] In S201, the tracking control unit 113 acquires motion information of the imaging device itself detected by the imaging device motion detection unit 126, and proceeds to S202.
[0050] In S202, the tracking control unit 113 determines whether panning is occurring based on the movement information of the imaging device itself, by checking whether the imaging device is moving in a certain direction. If it determines that panning is occurring, the process proceeds to S205; otherwise, the process proceeds to S203.
[0051] In S203, the tracking control unit 113 acquires the brightness information generated by the evaluation value generation unit 124 and proceeds to S204.
[0052] In S204, the tracking control unit 113 compares the acquired brightness information with a threshold. If the brightness is less than the threshold, it proceeds to S205; if it is greater than or equal to the threshold, it proceeds to S206. Specifically, this means that depending on the brightness of the image data, if it is dark, it proceeds to S205; if it is bright, it proceeds to S206. In this embodiment, the determination is made based on the brightness information of only one frame, but it may also be possible to compare the brightness information with the threshold over multiple frames and proceed to S205 if the brightness is less than the threshold over multiple frames.
[0053] In S205, the tracking control unit 113 decides to disable the DL tracking unit 116 and enable the non-DL tracking unit 117, and then finishes processing. This is because, in panning shots, the imaging device does not track the target subject, i.e., the moving subject; rather, the user captures the subject and moves the imaging device itself in a certain direction. Therefore, it can be considered a scene where tracking performance is not required, and the frequency of operation of the non-DL tracking unit 117 can be reduced. Similarly, when the image data is dark, i.e., night photography is expected, it can be considered a scene where tracking performance is not required, and the frequency of operation of the non-DL tracking unit 117 can be reduced.
[0054] In S206, the tracking control unit 113 decides to enable the DL tracking unit 116 and disable the non-DL tracking unit 117, and then terminates the process.
[0055] (Display processing by display unit 121) Figure 3 shows an example of live view display. Figure 3(a) shows image 300 represented by display image data output by the post-processing unit 114. Figure 2(b) shows image 302 represented by composite image data obtained by superimposing the image of the tracking frame 303 onto the display image data. In this case, since there is only one candidate subject 301 in the shooting range, candidate subject 301 is selected as the subject to be tracked. The tracking frame 303 is superimposed so as to surround the candidate subject 301. In the example of Figure 2(b), the tracking frame 303 is composed of a combination of four hollow hook shapes, but other forms of tracking frame 303 may be used, such as a combination of non-hollow hook shapes, a continuous frame, a combination of rectangles, or a combination of triangles. The form of the tracking frame 303 may also be selectable by the user.
[0056] Figure 4 is a flowchart illustrating the operation of the subject tracking function in a series of imaging operations by the imaging device 100. Each step is executed by the control unit 102 or by instructions from the control unit 102.
[0057] In S400, the control unit 102 controls the image sensor 103 to capture one frame and acquire image data.
[0058] In S401, the first preprocessing unit 104 applies preprocessing to the image data read from the image sensor 103.
[0059] In S402, the control unit 102 stores the preprocessed image data in the display memory 107.
[0060] In S403, the first image correction unit 109 starts applying a predetermined image correction process to the image data read from the display memory 107.
[0061] In S404, the control unit 102 determines whether all image correction processing to be applied has been completed. If it determines that all processing has been completed, it outputs the image data to which the image correction processing has been applied to the post-processing unit 114 and proceeds to S405. If the first image correction unit 109 determines that all image correction processing has not been completed, it continues the image correction processing.
[0062] In S405, the post-processing unit 114 generates display image data from the image data to which image correction processing has been applied by the first image correction unit 109, and outputs it to the information overlay unit 120.
[0063] In S406, the information overlay unit 120 uses the display image data generated by the post-processing unit 114, the tracking frame image data, and the image data indicating other information to generate composite image data in which the tracking frame and other information images are superimposed on the captured image. The information overlay unit 120 outputs the composite image data to the display unit 121.
[0064] In S407, the display unit 121 displays the composite image data generated by the information overlay unit 120. This completes the live view display for one frame.
[0065] As described above, in this embodiment, an image processing apparatus using a first tracking means and a second tracking means having a lower computational load than the first tracking means controls the enabling and disabling of the first and second tracking means based on at least one of the movement of the imaging device and the brightness of the image data. Therefore, power consumption can be reduced by disabling the first tracking means in scenes where there is little need to obtain good tracking results.
[0066] In this embodiment, when controlling the enabling / disabling of the DL / non-DL tracking unit based on the movement of the imaging device itself and the brightness of the image data, an example of exclusive control was shown where the non-DL tracking unit 117 is disabled when the DL tracking unit 116 is enabled. However, this is not limited to this, and in difficult panning scenes or low brightness values, both the DL tracking unit 116 and the non-DL tracking unit 117 may be enabled depending on the panning speed and the low brightness value. In other words, in this case, the tracking process may be controlled to perform tracking processing based on the tracking results of both. In the above embodiment, an example was shown where the control of enabling and disabling the DL tracking unit 116 and the non-DL tracking unit 117 was switched in a binary manner. However, this is not limited to this, and it may be switched in multiple stages depending on the brightness of the image and the movement of the subject. In other words, multiple levels of computational load may be provided for enabling the DL tracking unit 116 and the non-DL tracking unit 117, and it may be switched to perform processing with a higher computational load when it is more effective.
[0067] Furthermore, in this embodiment, the inactivation of the DL tracking unit 116 and the non-DL tracking unit 117 is shown as an example where all calculation processing performed by the L tracking unit 116 is omitted and not executed. However, this is not the only example, and it may also include omitting or not executing at least a part of the tracking calculation processing and tracking result output processing that are performed when the units are active, such as pre-processing for tracking processing and calculations for the main tracking processing.
[0068] (Second embodiment) Next, a second embodiment of the present invention will be described. Here, only the parts that differ from the first embodiment described above will be described, and the same parts will be denoted by the same reference numerals and detailed descriptions will be omitted. In the second embodiment, the imaging device automatically recognizes the shooting scene based on at least one of the following: the captured image, shooting parameters, the orientation of the imaging device, etc., and uses the result to control the DL tracking unit 116, the non-DL tracking unit 117, and the detection unit 110. The following explanation will be given with reference to Figures 5 and 6.
[0069] Figure 5 shows the operation flow of the control unit 102 in the second embodiment.
[0070] In S501, the control unit 102 determines the shooting scene shown in Figure 4 (explained later) and proceeds to S502. To determine the shooting scene in Figure 4, the brightness / darkness of the background is determined from the brightness information acquired by the evaluation value generation unit 124, and the blue sky / evening scene of the background is determined from the light source information obtained in the process of calculating the white balance correction value and the brightness information. In addition, whether the subject is a person or not is determined from the results of the detection unit 110, and whether it is a moving object or a non-moving object is determined by the tracking unit 115. Not limited to these determination methods, any known processing sequence that can determine the shooting scene from information obtained from images, gyro sensors, infrared sensors, ToF (time of flight) sensors, etc., can be applied. Panning detection is performed in the same way as in the first embodiment.
[0071] In S502, the control unit 102 controls the system to enter the operating mode shown in Figure 5 (explained later) corresponding to the shooting scene shown in Figure 4, and then completes the process. Specifically, the control unit 102 controls the detection unit 110 and notifies the tracking control unit 113 according to the operating mode in the shooting scene table in Figure 4. The tracking control unit 113, upon receiving the notification, controls the tracking unit 116.
[0072] Figure 6(a) is a table showing the relationship between the shooting scene and the operating modes of the detection unit 110 and the tracking unit 105. The horizontal axis represents the subject determination, such as whether it is a person, something other than a person, and whether it is a moving or stationary object, or whether it is a panning scene. The vertical axis represents the brightness of the background and whether it is a blue sky or a sunset scene. In other words, this table determines the operating mode by determining the subject and background. Note that the shooting scene in Figure 4 is just one example, and other shooting scenes may be added to determine the operating mode.
[0073] Figure 6(b) is a table showing the operating modes of the detection unit 110 and the tracking unit 105.
[0074] In operation mode 1, the DL tracking unit 116 is disabled, the non-DL tracking unit 117 is enabled, the detection unit 110 operates by detecting only people and non-people, and the operation cycle of the detection unit 110 is set to, for example, less than half of the shooting frame rate.
[0075] In operation mode 2, the DL tracking unit 116 is disabled, the non-DL tracking unit 117 is enabled, and the detection unit 110 is set to detect people and non-people, as well as inanimate objects such as buildings, roads, the sky, trees, etc., and the operation cycle is set to, for example, less than half of the shooting frame rate. The recognition results of inanimate objects are used, for example, for light source identification in white balance, or for image processing that distinguishes between man-made and non-man-made objects in the correction processing of the first image correction unit 109 and the second image correction unit 106.
[0076] In operation mode 3, the DL tracking unit 116 is enabled, the non-DL tracking unit 117 is enabled, the detection unit 110 operates by distinguishing between people and non-people, and the operation cycle is set, for example, to the same frame rate as the shooting for people and to less than half the frame rate for non-people.
[0077] In operation mode 4, the DL tracking unit 116 is turned ON, the non-DL tracking unit 117 is enabled, and the detection unit 110 operates by detecting people, non-people, and inanimate objects. Furthermore, the operation cycle is set, for example, to the same frame rate as the shooting for people, and to less than half the frame rate for non-people and inanimate objects.
[0078] In operation mode 5, the DL tracking unit 116 is disabled, the non-DL tracking unit 117 is enabled, the detection unit 110 operates by detecting people and non-people separately, and the operation cycle is set, for example, to the same frame rate as the shooting for people and to less than half the frame rate for non-people.
[0079] In operation mode 6, the DL tracking unit 116 is enabled, the non-DL tracking unit 117 is enabled, the detection unit 110 operates by distinguishing between people and non-people, and the operation cycle is set to, for example, less than half the shooting frame rate for people and the same as the shooting frame rate for non-people.
[0080] In operation mode 7, the DL tracking unit 116 is enabled, the non-DL tracking unit 117 is enabled, and the detection unit 110 operates by detecting people, non-people, and inanimate objects. In addition, the operation cycle is set to, for example, less than half the shooting frame rate for people and inanimate objects, and the same as the shooting frame rate for non-people.
[0081] In operation mode 8, the DL tracking unit 116 is disabled, the non-DL tracking unit 117 is enabled, the detection unit 110 operates by detecting people and non-people separately, and the operation cycle is set to, for example, less than half the shooting frame rate for people and the same as the shooting frame rate for non-people.
[0082] Note that the operating modes in Figure 6(b) are just one example of operating modes corresponding to the scene shown in Figure 6(a), and the operating modes may be changed. In this embodiment, the non-DL tracking unit 116 is used to determine whether the subject is moving or stationary as part of the shooting scene determination, so the non-DL tracking unit 116 is enabled in all operating modes. However, motion determination of the subject may also be performed by monitoring the position of the subject detected by the detection unit 110 over multiple frames. In that case, the non-DL tracking unit 116 may be disabled when the subject is stationary (determined not to be moving).
[0083] As described above, in this embodiment, in an image processing device using a first tracking means and a second tracking means that has a lower computational load than the first tracking means, the enabling and disabling of the first and second tracking means are controlled based on the scene in which the image was captured. Furthermore, the object to be detected by the detection unit from the image and the operation cycle are changed based on the scene in which the image was captured. Therefore, power consumption can be suppressed in scenes where there is little need to obtain good tracking results.
[0084] (Third embodiment) Next, a third embodiment of the present invention will be described. In the third embodiment, an image processing device that detects so-called "feature points" in multiple regions within an captured image controls the DL tracking unit 116, the non-DL tracking unit 117, and the detection unit 110 based on the detection results of these feature points. The following explanation will be given with reference to Figures 7 and 8.
[0085] Figure 7 shows the operation flow of the control unit 102 in the third embodiment. This flow is intended to operate when the imaging device 100 is powered on, an imaging mode is selected from the menu, the tracking subject to be tracked is determined for the images sequentially acquired from the image sensor 103, and the tracking process is performed. If there is an ON / OFF setting for tracking control, this flow may be controlled to start only when tracking control is set to ON.
[0086] In S701, the control unit 102 acquires the captured image output from the image sensor 103 or stored in the detection and tracking memory 108.
[0087] In S702, at the instruction of the control unit 102, the evaluation value generation unit 124 performs a detection process to analyze the captured image obtained in S601 and detect feature points within the image. Details of the feature point detection process will be described later.
[0088] In S703, the control unit 102 acquires the feature point intensity information calculated when detecting each feature point in S602.
[0089] In S704, the control unit 102 performs a determination process on feature points detected within the tracking subject region, which has been determined to contain the subject to be tracked by the DL tracking unit 116, the non-DL tracking unit 117, or other subject detection processes (such as face detection) up to the previous frame. Specifically, it determines whether the number of feature points in the tracking subject region whose feature point intensity is equal to or greater than a first threshold is equal to or greater than a second threshold. If the number of feature points whose feature point intensity is equal to or greater than a first threshold is equal to or greater than a second threshold, the process proceeds to S705. If the number of feature points whose feature point intensity is equal to or greater than a first threshold is less than a second threshold, the process proceeds to S706.
[0090] In S705, the control unit 102 performs a determination process on feature points detected outside the area determined to be the tracking subject area in the previous frame within the captured image. Specifically, it determines whether the number of feature points with a feature point intensity of 3 or greater than or equal to the third threshold is 4 or greater than or equal to the fourth threshold. If the number of feature points with a feature point intensity of 3 or greater than or equal to the fourth threshold is 4 or greater, the process proceeds to S707. If the number of feature points with a feature point intensity of 3 or greater than or equal to the fourth threshold is less than or equal to the fourth threshold, the process proceeds to S708.
[0091] In S706, the control unit 102 performs a determination process on feature points detected outside the area determined to be the tracking subject area in the previous frame within the captured image. Specifically, it determines whether the number of feature points with a feature point intensity of 3 or greater than or equal to the third threshold is 4 or greater than or equal to the fourth threshold. If the number of feature points with a feature point intensity of 3 or greater than or equal to the fourth threshold is 4 or greater, the process proceeds to S709. If the number of feature points with a feature point intensity of 3 or greater than or equal to the fourth threshold is less than or equal to the fourth threshold, the process proceeds to S710.
[0092] In S707, the tracking control unit 113, at the instruction of the control unit 102, enables both the DL tracking unit 116 and the non-DL tracking unit 117, and sets the operating rate of the DL tracking process higher than the operating rate of the non-DL tracking process. Because there are many subjects with complex textures both inside and outside the tracking subject area, making tracking difficult, tracking accuracy can be maintained by performing both tracking processes at a high rate.
[0093] In S708, the tracking control unit 113 disables the DL tracking unit 116 and enables the non-DL tracking unit 117, as instructed by the control unit 102. In this embodiment, the operating rate of the non-DL tracking process at this time is higher than the operating rate of the non-DL tracking process set in S707. Because it is easy to distinguish between areas inside and outside the tracking subject area, power consumption can be suppressed while maintaining tracking accuracy by performing the tracking process using only non-DL tracking.
[0094] In S709, the tracking control unit 113, instructed by the control unit 102, enables the DL tracking unit 116 and disables the non-DL tracking unit 117. In this embodiment, the operating rate of the DL tracking process at this time is set to the highest operating rate among those set for the DL tracking unit 116 in S707 to S710. The fact that there are few feature points within the tracking subject area and many feature points outside the tracking subject area makes tracking more difficult, and non-DL tracking processes that perform tracking based on edge parts in the image, such as feature point detection processing, are more likely to output incorrect results. Therefore, tracking is performed using only DL tracking processing to suppress a decrease in tracking accuracy.
[0095] In S710, the tracking control unit 113, instructed by the control unit 102, enables both the DL tracking unit 116 and the non-DL tracking unit 117, and sets the operating rates of the DL tracking process and the non-DL tracking process to be lower than the operating rates set in S707. In situations where there are few detectable feature points in any area, both inside and outside the tracking subject area, accuracy is poor in both the DL tracking process and the non-DL tracking process, which may cause the results to fluctuate across various areas. If these are reflected at a high rate, it can cause image flickering, so by enabling both tracking processes and lowering the operating rates, the decrease in visibility due to flickering of the tracking results is suppressed.
[0096] (Feature point detection process) Figure 8 is a flowchart of the feature point detection process performed by the feature point detection unit 201. In S800, the control unit 102 generates a horizontal first derivative image by performing a horizontal first derivative filter process on the region of the tracked subject. In S802, the control unit 102 generates a horizontal second derivative image by performing a horizontal first derivative filter process on the horizontal first derivative image obtained in S800.
[0097] In S801, the control unit 102 generates a vertical first derivative image by performing a vertical first derivative filter process on the region of the tracked subject.
[0098] In S804, the control unit 102 generates a horizontal second derivative image by further applying a vertical first derivative filter to the vertical first derivative image obtained in S801.
[0099] In S803, the control unit 102 generates horizontal and vertical first derivative images by further applying a vertical first derivative filter to the horizontal first derivative image obtained in S800.
[0100] In S805, the control unit 102 calculates the determinant Det of the Hessian matrix H obtained in S802, S803, and S804. When the horizontal second derivative obtained in S802 is Lxx, the vertical second derivative obtained in S804 is Lyy, and the horizontal first derivative and vertical first derivative obtained in S803 are Lxy, the Hessian matrix H is expressed by equation (1) and the determinant Det is expressed by equation (2).
[0101]
number
[0102]
number
[0103] In S806, the control unit 102 determines whether the determinant Det obtained in S805 is 0 or greater. If the determinant Det is 0 or greater, the process proceeds to S807. If the determinant Det is less than 0, the process proceeds to S808.
[0104] In S807, the control unit 102 detects points where the determinant Det is 0 or greater as feature points.
[0105] In S808, if the control unit 102 determines that it has processed all of the input subject area, it terminates the feature point detection process. If processing is not yet complete, it repeats the processes from S800 to S807 to continue the feature point detection process.
[0106] As described above, in this embodiment, an image processing device using a first tracking means and a second tracking means that has a lower computational load than the first tracking means is configured to control the enabling and disabling of the first and second tracking means based on the feature quantities of the image. Therefore, power consumption can be suppressed in scenes where there is little need to obtain good tracking results.
[0107] Although the present invention has been specifically described above based on examples, it goes without saying that the present invention is not limited to the above examples, and various modifications are possible without departing from the spirit of the invention. [Explanation of symbols]
[0108] 102 Control Unit 110 Detection unit 113 Tracking Control Unit 115 Tracking part 116 DL tracking section 117 Non-DL tracking section 124 Evaluation Value Generation Unit 126 Motion detection unit of imaging device
Claims
1. A first tracking means that performs subject tracking using an image acquired by an imaging means, A second tracking means, which performs subject tracking using the image acquired by the imaging means, has a lower computational load than the first tracking means, The system includes a control means that switches between enabling both the first tracking means and the second tracking means, or disabling one of them, based on feature points detected from the image acquired by the imaging means. The control means is Within the region of the tracked subject, the number of feature points detected from the image is greater than the second threshold. If, outside the area of the tracked subject, the number of feature points detected from the image is less than the fourth threshold, An image processing apparatus characterized by disabling the first tracking means and enabling the second tracking means.
2. The control means is Within the region of the tracked subject, the number of feature points detected from the image is greater than the second threshold. If, outside the area of the tracked subject, the number of feature points detected from the image is greater than the fourth threshold, The image processing apparatus according to claim 1, characterized in that both the first tracking means and the second tracking means are enabled.
3. The control means is Within the region of the tracked subject, the number of feature points detected from the image is less than the second threshold. If, outside the area of the tracked subject, the number of feature points detected from the image is greater than the fourth threshold, The image processing apparatus according to claim 1 or 2, characterized in that the first tracking means is enabled and the second tracking means is disabled.
4. The control means is Within the region of the tracked subject, the number of feature points detected from the image is less than the second threshold. If, outside the area of the tracked subject, the number of feature points detected from the image is less than the fourth threshold, The image processing apparatus according to any one of claims 1 to 3, characterized in that both the first tracking means and the second tracking means are enabled, while the operating rate is lower than in other cases.
5. The image processing apparatus according to any one of claims 1 to 4, characterized in that the feature points are detected by performing horizontal and vertical differential filtering on the image acquired by the imaging means.
6. The image processing apparatus according to any one of claims 1 to 4, characterized in that the first tracking means performs subject tracking using a multilayer neural network trained using deep learning.
7. The image processing apparatus according to any one of claims 1 to 4, characterized in that the second tracking means performs subject tracking based on the similarity of color configuration or pattern matching.
8. A first tracking step in which a first tracking means performs tracking of a subject using an image acquired by an imaging means, A second tracking step in which a second tracking means performs tracking of a subject using an image acquired by the imaging means, and the second tracking means has a lower computational load than the first tracking means, The control step includes switching between enabling both the first tracking means and the second tracking means, or disabling one of them, based on feature points detected from the image acquired by the imaging means. In the control step described above, Within the region of the tracked subject, the number of feature points detected from the image is greater than the second threshold. If, outside the area of the tracked subject, the number of feature points detected from the image is less than the fourth threshold, The first tracking means is disabled, and the second tracking means is enabled. An image processing method characterized by the following: