Imaging device, method for driving the imaging device, and program
By employing two models with different scales and dynamic selection based on frame rate and subject characteristics, the technology achieves accurate and real-time subject tracking, addressing the limitations of existing imaging technologies.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-12-25
- Publication Date
- 2026-03-25
AI Technical Summary
Existing imaging technologies face challenges in achieving both high accuracy and real-time performance in subject tracking, particularly in scenarios with varying frame rates and subject dynamics.
The implementation of two machine-learned models, a smaller first model for high-speed tracking and a larger second model for high accuracy, with dynamic selection based on frame rate and subject characteristics, along with the creation of reference images tailored to these models, ensures accurate and real-time subject tracking.
This approach maintains both accuracy and real-time performance by adaptively selecting the appropriate model and reference image resolution, ensuring consistent tracking regardless of frame rate or subject changes.
Smart Images

Figure 2026053644000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to an estimation device, a driving method of the estimation device, and a program.
Background Art
[0002] Japanese Patent Application Laid-Open No. 2020-038410 discloses a solid-state imaging device including a DNN processing unit that executes a DNN on an input image based on a DNN (Deep Neural Network) model, and a DNN control unit that receives control information generated based on evaluation information of the execution result of the DNN and changes the DNN model based on the control information.
[0003] Japanese Patent Application Laid-Open No. 2019-118097 discloses a selection step of performing a process of selecting any one of a plurality of learning models learned with criteria for recording an image generated by an imaging element, a determination step of performing a determination process of whether or not an image generated by the imaging element satisfies the criteria using the selected learning model, and a recording step of causing the image generated by the imaging element to be recorded in a memory when it is determined in the determination process that the image generated by the imaging element satisfies the criteria. The process of selecting any one of the learning models is performed based on at least any one of a shooting instruction by a user, an evaluation result of an image by the user, an environment when an image is generated by the imaging element, and a score of an image generated by the imaging element for the plurality of learning models.
Summary of the Invention
Problems to be Solved by the Invention
[0004] One embodiment of the technology according to the present disclosure provides an estimation device, a driving method of the estimation device, and a program that enable both the accuracy and real-time performance of subject tracking.
Means for Solving the Problems
[0005] To achieve the above objective, the estimation device of this disclosure comprises a memory storing a first model and a second model that have been machine-learned for subject tracking, and a processor that receives an imaging signal from an image sensor, wherein the processor is configured to perform a decision process to determine a subject to be tracked, a first creation process to create a first reference image for the first model including the subject to be tracked and a second reference image for the second model including the subject to be tracked based on the imaging signal, a selection process to select one of the first model and the second model as the selected model based on factor information, an input process to input the imaging image represented by the imaging signal into the selected model, and an estimation process to estimate the position of the subject to be tracked from the imaging image using the selected model and the reference image for the selected model from the first reference image and the second reference image.
[0006] The second model preferably has more layers or larger layers than the first model.
[0007] The second reference image is preferably of higher resolution than the first reference image.
[0008] The factor information preferably includes the type of subject being tracked, the speed at which the subject is moving, or the degree of change in the form of the subject being tracked.
[0009] The factor information is preferably the frame rate value of the captured image input to the selected model.
[0010] The processor is configured to be able to perform a second creation process, which creates a first reference image but does not create a second reference image, in place of the first creation process, and it is preferable to select either the first or second creation process based on the frame rate value.
[0011] It is preferable that the processor is configured to perform a first update process to update the first and second reference images when the selected model switches from one of the first and second models to the other during the selection process.
[0012] Preferably, the processor is configured to perform a second update process that updates the first and second reference images based on the change in the size of the captured image of the tracked subject within the field of view.
[0013] The processor is preferably configured to perform a second update process based on a change in the imaging magnification of an imaging device having an image sensor.
[0014] The method for driving an estimation device according to the present disclosure is a method for driving an estimation device comprising: a memory storing a first model and a second model that have been machine-learned for subject tracking, the method comprising: a reception step of receiving an imaging signal from an image sensor; a determination step of determining a subject to be tracked; a first creation step of creating a first reference image for the first model including the subject to be tracked and a second reference image for the second model including the subject to be tracked, based on the imaging signal; a selection step of selecting one of the first model and the second model as a selected model based on factor information; an input step of inputting the imaging image represented by the imaging signal into the selected model; and an estimation step of estimating the position of the subject to be tracked from the imaging image using the selected model and the reference image for the selected model from the first reference image and the second reference image.
[0015] The program of this disclosure is a program for operating an estimation device comprising a memory storing a first model and a second model that have been machine-learned for subject tracking, and causes the estimation device to execute: an acceptance process for receiving an imaging signal from an image sensor; a determination process for determining a subject to be tracked; a first creation process for creating a first reference image for the first model including the subject to be tracked and a second reference image for the second model including the subject to be tracked, based on the imaging signal; a selection process for selecting one of the first model and the second model as the selected model based on factor information; an input process for inputting the imaging image represented by the imaging signal into the selected model; and an estimation process for estimating the position of the subject to be tracked from the imaging image using the selected model and the reference image for the selected model from the first and second reference images. [Brief explanation of the drawing]
[0016] [Figure 1] It is a diagram showing an example of the internal configuration of an imaging device. [Figure 2] It is a block diagram showing an example of the functional configuration of a processor. [Figure 3] It is a diagram conceptually showing an example of the determination process of a tracking subject and the creation process of a reference image. [Figure 4] It is a diagram showing an example of the configuration of a first model. [Figure 5] It is a diagram showing an example of the configuration of a second model. [Figure 6] It is a diagram showing an example of teacher data used for machine learning of a first model. [Figure 7] It is a diagram showing an example of teacher data used for machine learning of a second model. [Figure 8] It is a diagram showing an example of a score map. [Figure 9] It is a flowchart explaining the processing procedure of a subject tracking function. [Figure 10] It is a diagram showing an example of using an image cut out from a captured image as a search image. [Figure 11] It is a flowchart showing the creation process of a reference image according to a modification example. [Figure 12] It is a flowchart explaining the processing procedure of a subject tracking function according to a modification example. [Figure 13] It is a flowchart showing an example of the processing procedure of a first update process. [Figure 14] It is a flowchart showing another example of the processing procedure of a first update process. [Figure 15] It is a flowchart showing an example of the processing procedure of a second update process.
Embodiments for Carrying Out the Invention
[0017] An example of an embodiment according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0018] [[ID=�5]]First, the language used in the following description will be explained.
[0019] In the following explanation, "IC" is an abbreviation for "Integrated Circuit". "CPU" is an abbreviation for "Central Processing Unit". "ROM" is an abbreviation for "Read Only Memory". "RAM" is an abbreviation for "Random Access Memory". "CMOS" is an abbreviation for "Complementary Metal Oxide Semiconductor".
[0020] FPGA is an abbreviation for Field Programmable Gate Array. PLD is an abbreviation for Programmable Logic Device. ASIC is an abbreviation for Application Specific Integrated Circuit. OVF is an abbreviation for Optical View Finder. EVF is an abbreviation for Electronic View Finder. JPEG is an abbreviation for Joint Photographic Experts Group. CNN is an abbreviation for Convolutional Neural Network.
[0021] As one embodiment of the imaging device, the technology of this disclosure will be explained using a lens-interchangeable digital camera as an example. However, the technology of this disclosure is not limited to lens-interchangeable cameras, but can also be applied to lens-integrated digital cameras.
[0022] Figure 1 shows an example of the configuration of the imaging device 10. The imaging device 10 is a lens-interchangeable digital camera. The imaging device 10 consists of a main body 11 and an imaging lens 12 that is interchangeably attached to the main body 11. The imaging lens 12 is attached to the front side of the main body 11 via a camera-side mount 11A and a lens-side mount 12A.
[0023] The main unit 11 is provided with an operation unit 13, which includes a dial, a shutter release button, and the like. The operating modes of the imaging device 10 include, for example, a still image capture mode, a video capture mode, and an image display mode. The operation unit 13 is operated by the user when setting the operating mode. The operation unit 13 is also operated by the user when starting still image capture or video capture. The operation unit 13 includes a touch panel provided on the display 15, which will be described later.
[0024] Furthermore, the imaging device 10 is equipped with a subject tracking function that tracks a subject specified by the user as the tracking target in video recording mode. The subject tracking function can also be activated when displaying a live view image, which is performed before still image capture or video recording. The imaging device 10 is an example of an "estimation device" related to the technology of this disclosure.
[0025] Furthermore, the main body 11 is equipped with a viewfinder 14. Here, the viewfinder 14 is a hybrid viewfinder (registered trademark). A hybrid viewfinder refers to a viewfinder in which, for example, an optical viewfinder (hereinafter referred to as "OVF") and an electronic viewfinder (hereinafter referred to as "EVF") are selectively used. The user can observe the optical image or live view image of the subject projected by the viewfinder 14 through the viewfinder eyepiece (not shown).
[0026] Furthermore, a display 15 is provided on the back of the main unit 11. The display 15 shows images based on the image signal obtained by imaging, as well as various menu screens, etc. The user can also observe the live view image displayed on the display 15 instead of the viewfinder 14.
[0027] The main unit 11 and the imaging lens 12 are electrically connected by contact between an electrical contact 11B provided on the camera-side mount 11A and an electrical contact 12B provided on the lens-side mount 12A.
[0028] The imaging lens 12 includes an objective lens 30, a focusing lens 31, a rear-end lens 32, and an aperture 33. Each component is arranged along the optical axis A of the imaging lens 12, from the objective side, in the order of objective lens 30, aperture 33, focusing lens 31, and rear-end lens 32. The objective lens 30, focusing lens 31, and rear-end lens 32 constitute the imaging optical system. The type, number, and arrangement order of the lenses constituting the imaging optical system are not limited to the example shown in Figure 1.
[0029] Furthermore, the imaging lens 12 has a lens drive control unit 34. The lens drive control unit 34 is composed of, for example, a CPU, RAM, and ROM. The lens drive control unit 34 is electrically connected to the processor 40 in the main unit 11 via electrical contacts 12B and 11B.
[0030] The lens drive control unit 34 drives the focus lens 31 and aperture 33 based on control signals transmitted from the processor 40. The lens drive control unit 34 controls the drive of the focus lens 31 based on focus control control signals transmitted from the processor 40 in order to adjust the focus position of the imaging lens 12. The processor 40 may also perform focus control based on an estimated result R representing the position of the tracked subject, which will be described later.
[0031] The aperture 33 has an aperture whose diameter is variable around the optical axis A. The lens drive control unit 34 controls the drive of the aperture 33 based on an aperture adjustment control signal transmitted from the processor 40 in order to adjust the amount of light incident on the light-receiving surface 20A of the image sensor 20.
[0032] Furthermore, the main unit 11 houses an image sensor 20, a processor 40, and a memory 42. The image sensor 20, memory 42, operation unit 13, viewfinder 14, and display 15 are all controlled by the processor 40.
[0033] The processor 40 is composed of, for example, a CPU, RAM, and ROM. In this case, the processor 40 performs various processes based on a program 43 stored in memory 42. The processor 40 may also be composed of an assembly of multiple IC chips.
[0034] Furthermore, memory 42 stores the first model M1 and the second model M2, which have been machine-trained for subject tracking. As will be explained in more detail later, the first model M1 and the second model M2 are composed of neural networks, and the second model M2 is larger in scale than the first model M1. Larger scale means that there are more layers (convolutional layers, pooling layers, fully connected layers, etc.) that make up the neural network, and / or the size of the layers (the number of neurons that make up the layer) is large. The first model M1 is small in scale, so the estimation process of the subject to be tracked is fast, but the estimation accuracy is low. Conversely, the second model M2 is large in scale, so the estimation process of the subject to be tracked is slow, but the accuracy of subject tracking is high.
[0035] The imaging sensor 20 is, for example, a CMOS type image sensor. The imaging sensor 20 is positioned such that the optical axis A is perpendicular to the light-receiving surface 20A and the optical axis A is located at the center of the light-receiving surface 20A. Light (subject image) that has passed through the imaging lens 12 is incident on the light-receiving surface 20A. Multiple pixels are formed on the light-receiving surface 20A, which generate an image signal by performing photoelectric conversion. The imaging sensor 20 generates and outputs an image signal by performing photoelectric conversion on the light incident on each pixel. Note that the imaging sensor 20 is an example of an "image sensor" related to the technology of this disclosure.
[0036] Furthermore, a Bayer-arranged color filter array is positioned on the light-receiving surface of the image sensor 20, with one of the R (red), G (green), or B (blue) color filters positioned opposite each pixel. Some of the multiple pixels arranged on the light-receiving surface of the image sensor 20 may be phase-difference pixels used for focus control.
[0037] Figure 2 shows an example of the functional configuration of the processor 40. The processor 40 implements various functional units by executing processing according to the program 43 stored in the memory 42. As shown in Figure 2, for example, the processor 40 implements a main control unit 50, an imaging control unit 51, an image processing unit 52, a tracking target determination unit 53, a reference image creation unit 54, a model selection unit 55, an image input unit 56, an estimation unit 57, and a display control unit 58.
[0038] The main control unit 50 comprehensively controls the operation of the imaging device 10 based on instruction signals input from the operation unit 13. The imaging control unit 51 controls the imaging sensor 20 to perform imaging operations. The imaging control unit 51 drives the imaging sensor 20 in still image mode or video imaging mode. The imaging sensor 20 outputs an imaging signal RD generated by the imaging operation. The imaging signal RD is so-called RAW data.
[0039] The image processing unit 52 performs reception processing to receive the imaging signal RD output from the imaging sensor 20. The image processing unit 52 also generates the captured image PD by performing image processing, including demosaicing, on the received imaging signal RD. For example, the captured image PD is a color image in which each pixel is represented by the three primary colors R, G, and B. More specifically, for example, the captured image PD is a 24-bit color image in which each of the R, G, and B signals contained in one pixel is represented by 8 bits.
[0040] The tracking target determination unit 53 performs a determination process to determine the subject specified by the user as the tracking target. For example, the user uses the operation unit 13 to specify the subject they want to track from within the captured image PD displayed on the display 15. The tracking target determination unit 53 determines the subject specified by the user from within the captured image PD as the tracking target.
[0041] Furthermore, if the imaging device 10 has a subject detection function that detects a subject based on the captured image PD, the tracking target determination unit 53 may determine a specific subject detected by the subject detection function as the tracking subject.
[0042] The reference image creation unit 54 performs a creation process to create a first reference image T1 for a first model that includes the tracking subject determined by the tracking target determination unit 53, and a second reference image T2 for a second model that includes the said tracking subject, based on the captured image PD. The creation process in this embodiment corresponds to the "first creation process" related to the technology of this disclosure.
[0043] The reference image creation unit 54 creates the first reference image T1 and the second reference image T2 by cutting out the region containing the tracked subject from the captured image PD. The second reference image T2 is a reference image for the second model M2, which is larger in scale than the first model M1, and therefore has a higher resolution than the first reference image T1. High resolution means that the image has a large number of pixels, a large amount of high-frequency component data, and a large number of bits for each pixel that makes up the image, resulting in a large amount of information. Hereafter, when the first reference image T1 and the second reference image T2 are not distinguished, they will simply be referred to as reference images. A reference image is a so-called template.
[0044] The model selection unit 55 performs a selection process to select one of the first model M1 and the second model M2 stored in the memory 42 as the selected model, based on the factor information. In this embodiment, the model selection unit 55 performs the selection process using the frame rate value as factor information. The frame rate refers to the reciprocal of the repetition period of the imaging operation by the imaging sensor 20.
[0045] The frame rate value can be changed, for example, by user operation using the control unit 13. Furthermore, the frame rate value may decrease if a composite mode is selected that combines multiple frames to increase image brightness.
[0046] The first model M1 has high estimation processing speed but low estimation accuracy, making it suitable for tracking subjects with small shape changes or blur between frames. When the frame rate is high, high-speed subject tracking processing is required, and since the time difference between frames is small and the shape changes or blur of the subject is small, the model selection unit 55 selects the first model M1 as the selected model.
[0047] On the other hand, the second model M2 has a slow estimation process but high estimation accuracy, making it suitable for tracking subjects with large shape changes or blurs between frames. When the frame rate is low, high-speed subject tracking processing is not necessary, but the time difference between frames is large and the shape changes or blurs of the subject are large, so the model selection unit 55 selects the second model M2 as the selected model.
[0048] The image input unit 56 performs input processing to input the captured image PD, represented by the imaging signal RD, to the selected model selected by the model selection unit 55. In this embodiment, the captured image PD that the model selection unit 55 inputs to the selected model is a search image for searching for a tracking subject included in the reference image.
[0049] Furthermore, the image input unit 56 changes the resolution of the captured image PD input to the selected model according to the resolution of the reference image input to the selected model. When the selected model is the second model M2, the image input unit 56 increases the resolution of the captured image PD compared to when the selected model is the first model M1.
[0050] The estimation unit 57 performs estimation processing to estimate the position of the tracked subject from the captured image PD using the selected model selected by the model selection unit 55 and the reference image for the selected model. Specifically, if the model selection unit 55 selects the first model M1, the estimation unit 57 inputs the first reference image T1 to the selected model. On the other hand, if the model selection unit 55 selects the second model M2, the estimation unit 57 inputs the second reference image T2 to the selected model.
[0051] The selected model outputs a score map SM that represents the similarity between each region in the captured image PD and the reference image. The estimation unit 57 outputs the information of the position with the highest score (i.e., the highest similarity) in the score map SM to the display control unit 58 as the estimated position R of the tracked subject.
[0052] The display control unit 58 displays the estimated result R on the display 15 along with the captured image PD. Specifically, the display control unit 58 displays the position of the tracked subject within the captured image PD in a recognizable manner based on the estimated result R. For example, the display control unit 58 displays a rectangular frame surrounding the tracked subject within the captured image PD.
[0053] Figure 3 conceptually illustrates an example of the process for determining the tracking subject and creating a reference image. In Figure 3, region S is the region designated by the user as the tracking target from within the captured image PD using the operation unit 13. The tracking target determination unit 53 determines the subject included in the designated region S as the tracking subject H.
[0054] The reference image creation unit 54 creates a first reference image T1 by cutting out the region containing the tracking subject H from the captured image PD and reducing the resolution of the cut-out image (in other words, the resolution of the second reference image T2 is higher than the resolution of the first reference image T1). The reference image creation unit 54 also creates a second reference image T2 by cutting out the region containing the tracking subject H from the captured image PD.
[0055] Figure 4 shows an example of the configuration of the first model M1. The first model M1 consists of a first convolutional network (hereinafter referred to as the first CNN) 61A, a second convolutional network (hereinafter referred to as the second CNN) 62A, and a convolutional operation unit 63A.
[0056] The first CNN61A is composed of multiple convolutional layers and multiple pooling layers. Similarly, the second CNN62A is composed of multiple convolutional layers and multiple pooling layers. The convolutional operation unit63A is composed of multiple fully connected layers.
[0057] The first CNN 61A receives the first reference image T1 as input. The second CNN 62A receives the captured image PD as input. The first CNN 61A converts the input first reference image T1 into a feature map FM1 and outputs it. The second CNN 62A converts the input captured image PD into a feature map FM2 and outputs it. Feature maps FM1 and FM2 are input to the convolution unit 63A.
[0058] The first CNN61A and the second CNN62A have similar configurations, but the input layer into which the image is input is sized according to the size of the input image (number of neurons). In other words, the size of the input layer differs between the first CNN61A and the second CNN62A.
[0059] The convolution unit 63A generates a score map SM by convolving the feature map FM2 with the feature map FM1 as the kernel, and outputs the generated score map SM to the estimation unit 57. The score map SM is an image that represents the similarity of each region in the captured image PD with the first reference image T1. The higher the similarity, the higher the score.
[0060] Figure 5 shows an example of the configuration of the second model M2. The second model M2 consists of a first CNN 61B, a second CNN 62B, and a convolution unit 63B. The first CNN 61B, the second CNN 62B, and the convolution unit 63B each have more layers than the first CNN 61A, the second CNN 62A, and the convolution unit 63A. The layer sizes of the first CNN 61B, the second CNN 62B, and the convolution unit 63B may also be larger than those of the first CNN 61A, the second CNN 62A, and the convolution unit 63A.
[0061] The second model M2 has the same configuration as the first model M1, except that it has a larger number of layers and / or larger layer sizes. A larger number of layers means a larger number of convolutional or pooling layers. Furthermore, a larger layer size means a larger number of operations or computational complexity in the convolutional or pooling layers.
[0062] The first CNN61B receives the second reference image T2 as input. The second CNN62B receives the captured image PD as input. The first CNN61B converts the input second reference image T2 into a feature map FM1 and outputs it. The second CNN62B converts the input captured image PD into a feature map FM2 and outputs it. Feature maps FM1 and FM2 are input to the convolution unit 63B.
[0063] The convolution unit 63B generates a score map SM by convolving the feature map FM2 using the feature map FM1 as a kernel, and outputs the generated score map SM to the estimation unit 57.
[0064] Figure 6 shows an example of training data used for machine learning of the first model M1. Machine learning of the first model M1 is performed using two frames selected from the video as a pair. Specifically, machine learning is performed by inputting training data consisting of a first reference image T1 generated from the first frame and an captured image PD generated from the second frame into the first model M1. For machine learning of the first model M1, it is preferable to use two frames that have a small time difference and small changes in the shape of the subject.
[0065] Figure 7 shows an example of training data used for machine learning of the second model M2. Machine learning of the second model M2 is performed using two frames selected from the video as a pair. Specifically, machine learning is performed by inputting training data consisting of a second reference image T2 generated from the first frame and an captured image PD generated from the second frame into the second model M2. For machine learning of the second model M2, it is preferable to use two frames that have a large time difference and show significant changes in the shape of the subject.
[0066] Figure 8 shows an example of a score map SM. As shown in Figure 8, the estimation unit 57 identifies, for example, a region U in the score map SM that contains the highest score, and outputs the location information of the identified region U as an estimation result R to the display control unit 58.
[0067] Figure 9 is a flowchart illustrating the processing procedure for the subject tracking function during video capture or live view image display.
[0068] The main control unit 50 determines whether or not a user has issued an instruction to start video capture or live view image display by operating the operation unit 13 (step S10). If an instruction to start is received (step S10: YES), the main control unit 50 controls the imaging control unit 51 to cause the imaging sensor 20 to perform an imaging operation and acquires the imaging signal RD output from the imaging sensor 20 (step S11). The display control unit 58 displays the captured image PD generated by the image processing unit 52 based on the imaging signal RD on the display 15 (step S12).
[0069] The main control unit 50 determines whether the user has specified a region to be tracked from within the captured image PD using the operation unit 13 (step S13). If the user has not specified a region (step S13: NO), the main control unit 50 returns to step S11 and causes the imaging sensor 20 to perform an imaging operation. The processes in steps S11 to S12 are repeatedly executed until it is determined in step S13 that the user has specified a region.
[0070] If the user specifies an area (step S13: YES), the main control unit 50 causes the tracking target determination unit 53 to determine the tracking target (step S14). In step S14, the tracking target determination unit 53 determines the subject included in the specified area as the tracking subject H.
[0071] The reference image creation unit 54 cuts out the region containing the tracked subject H from the captured image PD to create the first reference image T1 and the second reference image T2 (step S15). Here, the second reference image T2 has a higher resolution than the first reference image T1.
[0072] The model selection unit 55 uses the frame rate value as factor information to select either the first model M1 or the second model M2 as the selected model (step S16). In step S16, the model selection unit 55 selects the first model M1 as the selected model if the frame rate value is above a certain value, and selects the second model M2 as the selected model if the frame rate value is below a certain value.
[0073] The main control unit 50 controls the imaging control unit 51 to cause the imaging sensor 20 to perform an imaging operation and acquires the imaging signal RD output from the imaging sensor 20 (step S17). The image input unit 56 inputs the captured image PD generated by the image processing unit 52 based on the imaging signal RD to the selected model as a resolution corresponding to the selected model selected by the model selection unit 55 (step S18).
[0074] The estimation unit 57 inputs the reference image for the selected model, selected by the model selection unit 55 from the first reference image T1 and the second reference image T2, into the selected model. Based on the score map SM output from the selected model, it estimates the position of the tracked subject from the imaging signal RD and outputs the estimation result R to the display control unit 58 (step S19). The display control unit 58 displays the estimation result R together with the imaging image PD on the display 15 (step S20).
[0075] The main control unit 50 determines whether a predetermined termination condition is met (step S21). The termination condition is, for example, that the user has performed an operation to stop video capture using the operation unit 13. If the termination condition is not met (step S21: NO), the main control unit 50 returns to step S17 and causes the imaging sensor 20 to perform the imaging operation. The processes in steps S17 to S20 are repeatedly executed until it is determined in step S21 that the termination condition has been met. If the termination condition is met (step S21: YES), the main control unit 50 terminates the process.
[0076] In the flowchart above, steps S11 and S17 correspond to the "reception process" related to the technology of this disclosure. Step S14 corresponds to the "decision process" related to the technology of this disclosure. Step S15 corresponds to the "first creation process" related to the technology of this disclosure. Step S16 corresponds to the "selection process" related to the technology of this disclosure. Step S18 corresponds to the "input process" related to the technology of this disclosure. Step S19 corresponds to the "estimation process" related to the technology of this disclosure.
[0077] As described above, according to the technology of this disclosure, when the frame rate is high, the small-scale first model M1 is selected to prioritize real-time performance, and when the frame rate is low, the large-scale second model M2 is selected to prioritize accuracy in subject tracking. When the frame rate is high, the change in shape or blur of the tracked subject between frames is small, so the accuracy of subject tracking is maintained at a constant level even with the small-scale first model M1. Also, when the frame rate is low, the frame period is long, so the real-time performance is maintained at a constant level even with the large-scale second model M2. Thus, according to the technology of this disclosure, it is possible to achieve both accuracy in subject tracking and real-time performance.
[0078] Furthermore, according to the technology disclosed herein, a first reference image T1 for the first model M1 and a second reference image T2 for the second model M2 are created, and estimation processing is performed using the reference image corresponding to the selected model. Therefore, there is no need to recreate the reference image when switching the selected model. Consequently, real-time performance can be maintained even when switching the selected model.
[0079] [Differentiation] The following shows various modifications of the above embodiment. For each modification, only the differences from the above embodiment will be explained.
[0080] In the above embodiment, the image input unit 56 inputs the entire captured image PD as a search image to the selection model, but it is also possible to input an image cropped from the captured image PD as a search image to the selection model. For example, as shown in Figure 10, the image input unit 56 sets the search range to include the region U containing the tracked subject estimated by the estimation unit 57 in the previous frame period, and crops an image within the search range from the captured image PD obtained in the current frame period and inputs it to the selection model. By limiting the search range in this way, the processing speed by the selection model is improved.
[0081] (Creating a reference image) In the above embodiment, the reference image creation unit 54 executes a creation process (first creation process) to create a first reference image T1 and a second reference image T2 from the captured image PD. The reference image creation unit 54 may be configured to execute a second creation process, which creates the first reference image T1 but does not create the second reference image T2, instead of the first creation process. For example, the reference image creation unit 54 selectively executes either the first creation process or the second creation process based on the frame rate value.
[0082] Figure 11 shows the process of creating a reference image according to a modified example. The process shown in Figure 11 is performed, for example, in step S15 of the flowchart shown in Figure 9. The reference image creation unit 54 determines whether the frame rate value is less than a certain value (step S30). If the frame rate value is less than a certain value (step S30: YES), the reference image creation unit 54 executes the first creation process (step S31). On the other hand, if the frame rate value is greater than or equal to a certain value (step S30: NO), the reference image creation unit 54 executes the second creation process (step S32).
[0083] In other words, if the model selection unit 55 selects the second model M2 as the selected model, the first creation process is executed, and if the model selection unit 55 selects the first model M1 as the selected model, the second creation process is executed. When the frame rate is high, the processing can be sped up by not creating the second reference image T2.
[0084] (Model selection) In the above embodiment, the model selection unit 55 performs selection processing using the frame rate value as factor information, but the factor information is not limited to the frame rate value. For example, the model selection unit 55 may perform selection processing using the type of subject to be tracked determined by the tracking target determination unit 53 as factor information.
[0085] The first model, M1, has fast estimation processing but low estimation accuracy, making it suitable for tracking subjects with small shape changes between frames. Subjects with small shape changes between frames are objects with high rigidity, such as vehicles and aircraft. On the other hand, the second model, M2, has slow estimation processing but high estimation accuracy, making it suitable for tracking subjects with large shape changes between frames. Subjects with large shape changes between frames are objects with low rigidity, such as humans and animals. Humans and animals are prone to shape changes due to the movement of their limbs, etc.
[0086] In the above embodiment, the selected model is not changed after the model selection unit 55 performs the selection process for the selected model. However, the selected model may be changed according to factor information that changes during the subject tracking operation. For example, as shown in the flowchart in Figure 12, if the termination condition is not met (step S21: NO), the main control unit 50 returns the process to step S16 and causes the model selection unit 55 to perform the selection process for the selected model again. In this way, the model selection unit 55 may be made to repeatedly perform the selection process until the termination condition is met.
[0087] In this case, it is preferable for the model selection unit 55 to perform the selection process using the movement speed of the tracked subject as factor information. When the movement speed of the tracked subject is high, the shape of the tracked subject changes significantly between frames. For this reason, it is preferable for the model selection unit 55 to select the second model M2 as the selected model when the movement speed of the tracked subject is above a certain value, and to select the first model M1 as the selected model when the movement speed of the tracked subject is below a certain value.
[0088] Furthermore, it is preferable for the model selection unit 55 to perform selection processing using the degree of change in the form of the tracked subject as factor information. The degree of change in the form of the tracked subject refers to, for example, the degree of change in shape or the degree of change in color. It is preferable for the model selection unit 55 to select the second model M2 as the selected model if the degree of change in the form of the tracked subject between frames is greater than or equal to a certain value, and to select the first model M1 as the selected model if the degree of change in the form of the tracked subject between frames is less than a certain value.
[0089] Furthermore, it is preferable for the model selection unit 55 to perform selection processing using the score obtained from the score map SM output from the selected model as factor information. For example, if the model selection unit 55 has selected the first model M1 as the selected model and the maximum score falls below a threshold, it may determine that the tracking accuracy has decreased and select the second model M2, which has higher tracking accuracy, as the selected model.
[0090] (Reference image updated) In the above embodiment, in order to speed up the subject tracking operation, the reference image created by the reference image creation unit 54 is not updated until the subject tracking operation is completed. This is because if the reference image is updated when the tracking subject undergoes a change in posture such as rotation or when occlusion (i.e., objects intersect) occurs during the subject tracking operation, the possibility of mistracking an object other than the tracking subject increases. Here, updating means that the reference image creation unit 54 creates a new reference image.
[0091] As described above, it is preferable not to update the reference image in principle, but the reference image creation unit 54 may update the reference image when certain conditions are met.
[0092] For example, the reference image creation unit 54 performs a first update process to update the first reference image T1 and the second reference image T2 when the selected model switches from one of the first model M1 and the second model M2 to the other. Specifically, immediately after step S16 in the flowchart shown in Figure 12, the first update process for the reference images shown in Figure 13 is performed.
[0093] In Figure 12, the reference image creation unit 54 determines whether the selected model was changed by the model selection unit 55 in step S16 (step S40). If the selected model was not changed (step S40: NO), the reference image creation unit 54 does not update the reference image. On the other hand, if the selected model was changed (step S40: YES), the reference image creation unit 54 updates the reference image (step S41). In step S41, the reference image creation unit 54 creates a first reference image T1 and a second reference image T2 based on an image cropped from the region U (see Figure 8) identified by the estimation unit 57 in the captured image PD obtained in the previous frame period.
[0094] Furthermore, if the reference image is updated when the score of region U is low, the reliability of the updated reference image will be low. Therefore, it is preferable to update the reference image only when the score is above a certain value. For example, as shown in the flowchart of Figure 14, when the selected model is changed (step S40: YES), the reference image creation unit 54 determines whether the score (e.g., maximum value) of region U identified by the estimation unit 57 is above a certain value (step S42). If the score is not above a certain value (step S42: NO), the reference image creation unit 54 does not update the reference image. On the other hand, if the score is above a certain value (step S42: YES), the reference image creation unit 54 updates the reference image (step S41).
[0095] Furthermore, the reference image creation unit 54 may perform a second update process to update the reference image based on the change in the size of the tracked subject within the field of view of the captured image PD. Changes in the size of the tracked subject occur, for example, when the tracked subject approaches the imaging device 10 or moves away from the imaging device 10. When the size of the tracked subject changes, the similarity with the reference image decreases, which reduces the accuracy of subject tracking. For this reason, it is preferable for the reference image creation unit 54 to update the reference image when the size of the tracked subject changes by a certain value or more, based on the size of the tracked subject in the reference image. The size of the tracked subject can be detected using the subject detection result from the subject detection function.
[0096] The size of the tracked subject within the field of view of the captured image PD depends on the distance from the imaging device 10 to the tracked subject. Therefore, the reference image creation unit 54 may update the reference image based on distance information detected by the phase difference pixels of the imaging sensor 20 when the distance from the imaging device 10 to the tracked subject changes by a certain value or more.
[0097] Furthermore, changes in the size of the tracked subject can also occur due to changes in the imaging magnification of the imaging device 10. For this reason, it is preferable for the reference image creation unit 54 to update the reference image after creating it if the imaging magnification changes by a certain value or more. The imaging magnification is not limited to optical zoom, but can also be changed by electronic zoom. For example, the imaging magnification can be changed by the user operating the operation unit 13.
[0098] Figure 15 is a flowchart showing an example of the second update process. The reference image creation unit 54 performs the second update process of the reference image shown in Figure 15 during the subject tracking operation. In Figure 15, the reference image creation unit 54 determines whether the imaging magnification has changed by more than a certain value (step S50). If the imaging magnification has not changed by more than a certain value (step S50: NO), the reference image creation unit 54 does not update the reference image. On the other hand, if the imaging magnification has changed by more than a certain value (step S50: YES), the reference image creation unit 54 updates the reference image (step S51).
[0099] In the second update process, it is preferable that the reference image creation unit 54 updates the reference image on the condition that the score is above a certain value.
[0100] Furthermore, the reference image creation unit 54 may periodically update the reference image during subject tracking. For example, the reference image creation unit 54 may update the reference image once every several hundred frames during subject tracking. In this case as well, it is preferable that the reference image creation unit 54 updates the reference image only if the score is above a certain value.
[0101] Furthermore, the technology disclosed herein is not limited to digital cameras, but can also be applied to electronic devices such as smartphones and tablet devices that have imaging capabilities.
[0102] In the above embodiment, the hardware structure of the control unit, with processor 40 as an example, can be any of the following types of processors. These types of processors include a CPU, which is a general-purpose processor that functions by executing software (programs), as well as processors whose circuit configuration can be changed after manufacturing, such as FPGAs. FPGAs include dedicated electrical circuits, which are processors with circuit configurations specifically designed to perform specific processing, such as PLDs or ASICs.
[0103] The control unit may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Furthermore, multiple control units may consist of a single processor.
[0104] There are several possible examples of configuring multiple control units with a single processor. A first example is a configuration where a single processor is composed of one or more CPUs and software, as exemplified by client and server computers, and this processor functions as multiple control units. A second example is a configuration where a processor that realizes the functions of the entire system, including multiple control units, on a single IC chip is used, as exemplified by System-on-a-Chip (SOC) systems. Thus, a control unit can be configured as a hardware structure using one or more of the above-mentioned types of processors.
[0105] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices.
[0106] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0107] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
Claims
1. A memory containing the first and second models that have undergone machine learning for subject tracking, A processor that receives the imaging signal from the image sensor, An estimation device comprising, The aforementioned processor, A decision process to determine the subject to be tracked, A first creation process that creates a first reference image for the first model including the tracked subject and a second reference image for the second model including the tracked subject based on the imaging signal, A selection process that selects one of the first model and the second model as the selected model based on factor information, Input processing involves inputting the captured image represented by the aforementioned imaging signal to the selected model, An estimation process that estimates the position of the tracked subject from the captured image using the selected model and the reference image for the selected model from the first reference image and the second reference image, An estimation device configured to perform the following actions.
2. The second model has more layers or larger layer sizes than the first model. The estimation device according to claim 1.
3. The second reference image has a higher resolution than the first reference image. The estimation device according to claim 2.
4. The aforementioned factor information includes the type of subject being tracked, the speed at which the subject is moving, or the degree of change in the form of the subject being tracked. The estimation device according to claim 3.
5. The aforementioned factor information is the frame rate value of the captured image input to the selected model. The estimation device according to claim 3.
6. The aforementioned processor, The system is configured to allow a second creation process, which creates the first reference image but does not create the second reference image, to be executed in place of the first creation process. Based on the frame rate value, select either the first creation process or the second creation process. The estimation device according to claim 5.
7. The aforementioned processor, In the selection process, when the selected model switches from one of the first model and the second model to the other, a first update process is executed to update the first reference image and the second reference image. An estimation device according to any one of claims 1 to 6.
8. The aforementioned processor, The system is configured to perform a second update process that updates the first and second reference images based on the change in the size of the tracked subject within the field of view of the captured image. The estimation device according to any one of claims 1 to 7.
9. The aforementioned processor, The system is configured to perform the second update process based on a change in the imaging magnification of the imaging device having the image sensor. The estimation device according to claim 8.
10. A memory containing the first and second models that have undergone machine learning for subject tracking, A method for driving an estimation device comprising, A reception process for receiving imaging signals from the image sensor, The decision process for determining the subject to be tracked, A first creation step of creating a first reference image for the first model including the tracked subject and a second reference image for the second model including the tracked subject based on the imaging signal, A selection step in which one of the first model and the second model is selected as the selected model based on factor information, An input step of inputting the captured image represented by the aforementioned imaging signal to the selected model, An estimation step of estimating the position of the tracked subject from the captured image using the selected model and the reference image for the selected model from the first reference image and the second reference image, A method for driving an estimation device, including the method described above.
11. A memory containing the first and second models that have undergone machine learning for subject tracking, A program for operating an estimation device equipped with, The reception process receives the imaging signal from the image sensor, A decision process to determine the subject to be tracked, A first creation process that creates a first reference image for the first model including the tracked subject and a second reference image for the second model including the tracked subject based on the imaging signal, A selection process that selects one of the first model and the second model as the selected model based on factor information, Input processing involves inputting the captured image represented by the aforementioned imaging signal to the selected model, An estimation process that estimates the position of the tracked subject from the captured image using the selected model and the reference image for the selected model from the first reference image and the second reference image, A program that causes the estimation device to execute the above.