Estimation device, method for driving estimation device, and program
The estimation device uses dual models with varying scales and resolutions to achieve accurate and real-time subject tracking by selecting the appropriate model based on frame rate, addressing the challenge of balancing accuracy and speed in imaging devices.
Patent Information
- Application Number
- JP2023549392
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-09-27
- Filing Date
- 2022-07-15
- Publication Date
- 2026-01-14
- Estimated Expiration
- 2042-07-15
AI Technical Summary
Existing technologies face challenges in achieving both accuracy and real-time performance in subject tracking, particularly in imaging devices.
An estimation device equipped with two machine-learned models, a smaller first model for high-speed estimation and a larger second model for high accuracy, along with reference images of varying resolutions, is used to select the appropriate model based on frame rate and other factors for optimal tracking.
This approach enables both high accuracy and real-time performance in subject tracking by dynamically selecting the most suitable model based on frame rate and other factors, maintaining performance even when models are switched.
Smart Images

Figure 0007798904000001 
Figure 0007798904000002 
Figure 0007798904000003
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to an estimation device, a driving method for the estimation device, and a program. [Background technology]
[0002] Japanese Patent Application Laid-Open Publication No. 2020-038410 discloses a solid-state imaging device that includes a deep neural network (DNN) processing unit that executes DNN on an input image based on a DNN model, and a DNN control unit that receives control information generated based on evaluation information of the DNN execution results and changes the DNN model based on the control information.
[0003] Japanese Patent Application Laid-Open No. 2019-118097 discloses an imaging device having a selection process for selecting one of a plurality of learning models that have learned standards for recording an image generated by an imaging element, a determination process for determining whether or not the image generated by the imaging element satisfies the standards using the selected learning model, and a recording process for recording the image generated by the imaging element in a memory when it is determined in the determination process that the image generated by the imaging element satisfies the standards. method The process of selecting one of the learning models is performed by the user. to The evaluation is based on at least one of the following: shooting instructions from the camera, the user's evaluation of the image, the environment when the image was generated by the image sensor, and the scores of multiple learning models for the image generated by the image sensor. Summary of the Invention [Problem to be solved by the invention]
[0004] One embodiment of the technique of the present disclosure provides an estimation device, a method for driving the estimation device, and a program that enable both accuracy and real-time performance in subject tracking. [Means for solving the problem]
[0005] In order to achieve the above object, the estimation device of the present disclosure is an estimation device that includes a memory that stores a first model and a second model that have been machine-learned for object tracking, and a processor that receives an image signal from an image sensor, and the processor is configured to execute the following processes: a determination process that determines a tracking object to be tracked; a first creation process that creates a first reference image for the first model that includes the tracking object and a second reference image for the second model that includes the tracking object based on the image signal; a selection process that selects one of the first model and the second model as a selected model based on factor information; an input process that inputs the captured image represented by the image signal into the selected model; and an estimation process that estimates the position of the tracking object from the captured image using the selected model and one of the first and second reference images for the selected model.
[0006] The second model preferably has a larger number of layers or larger layer sizes than the first model.
[0007] The second reference image preferably has a higher resolution than the first reference image.
[0008] The factor information is preferably the type of the subject being tracked, the moving speed of the subject being tracked, or the degree of change in the form of the subject being tracked.
[0009] The factor information is preferably a value of the frame rate of the captured image input to the selection model.
[0010] The processor is configured to be able to execute a second creation process, which creates a first reference image but does not create a second reference image, instead of the first creation process, and it is preferable to select the first creation process or the second creation process based on the value of the frame rate.
[0011] Preferably, the processor is configured to execute a first update process to update the first reference image and the second reference image when the selected model is switched from one of the first model and the second model to the other during the selection process.
[0012] The processor is preferably configured to execute a second update process for updating the first reference image and the second reference image based on a change in size of the tracked subject within the angle of view of the captured image.
[0013] The processor is preferably configured to execute the second update process based on a change in imaging magnification of an imaging device having an imaging element.
[0014] The method for driving the estimation device of the present disclosure includes: Li a first creation step of creating, based on the imaging signal, a first reference image for a first model including the tracking subject and a second reference image for a second model including the tracking subject; a selection step of selecting, based on factor information, one of the first model and the second model as a selected model; an input step of inputting the imaging image represented by the imaging signal into the selected model; and an estimation step of estimating the position of the tracking subject from the imaging image, using the selected model and one of the first and second reference images for the selected model.
[0015] The program of the present disclosure includes a memory storing a first model and a second model that have been machine-learned for subject tracking. Li The program causes the estimation device to execute the following processes: a reception process for receiving an image signal from an image sensor; a determination process for determining a tracking subject to be tracked; a first creation process for creating a first reference image for a first model including the tracking subject and a second reference image for a second model including the tracking subject based on the image signal; a selection process for selecting one of the first model and the second model as a selected model based on factor information; an input process for inputting the image signal represented by the image signal into the selected model; and an estimation process for estimating the position of the tracking subject from the image signal using the selected model and one of the first and second reference images for the selected model. [Brief explanation of the drawings]
[0016] [Figure 1] FIG. 1 is a diagram illustrating an example of the internal configuration of an imaging device. [Figure 2] FIG. 2 is a block diagram illustrating an example of a functional configuration of a processor. [Figure 3] 10A to 10C are diagrams conceptually illustrating an example of a process for determining a subject to be tracked and a process for creating a reference image. [Figure 4] FIG. 2 is a diagram illustrating an example of a configuration of a first model. [Figure 5] FIG. 10 is a diagram illustrating an example of the configuration of a second model. [Figure 6] FIG. 10 is a diagram illustrating an example of training data used in machine learning of a first model. [Figure 7] FIG. 10 is a diagram illustrating an example of training data used in machine learning of a second model. [Figure 8] FIG. 10 is a diagram illustrating an example of a score map. [Figure 9] 10 is a flowchart illustrating a processing procedure of a subject tracking function. [Figure 10] FIG. 10 is a diagram showing an example in which an image cut out from a captured image is used as a search image. [Figure 11] 10 is a flowchart showing a reference image creation process according to a modified example. [Figure 12] 10 is a flowchart illustrating a processing procedure of a subject tracking function according to a modified example. [Figure 13] 10 is a flowchart illustrating an example of a processing procedure for a first update process. [Figure 14] 10 is a flowchart showing another example of the processing procedure of the first update process. [Figure 15] 10 is a flowchart illustrating an example of a processing procedure for a second update process. DETAILED DESCRIPTION OF THE INVENTION
[0017] An example of an embodiment of the technology of the present disclosure will be described with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following explanation, "IC" is an abbreviation for "Integrated Circuit." "CPU" is an abbreviation for "Central Processing Unit." "ROM" is an abbreviation for "Read Only Memory." "RAM" is an abbreviation for "Random Access Memory." "CMOS" is an abbreviation for "Complementary Metal Oxide Semiconductor."
[0020] "FPGA" is an abbreviation for "Field Programmable Gate Array." "PLD" is an abbreviation for "Programmable Logic Device." "ASIC" is an abbreviation for "Application Specific Integrated Circuit." "OVF" is an abbreviation for "Optical View Finder." "EVF" is an abbreviation for "Electronic View Finder." "JPEG" is an abbreviation for "Joint Photographic Experts Group." "CNN" is an abbreviation for "Convolutional Neural Network."
[0021] The technology of the present disclosure will be described using an interchangeable lens digital camera as an example of an embodiment of an imaging device. Note that the technology of the present disclosure is not limited to interchangeable lens digital cameras, and can also be applied to digital cameras with an integrated lens.
[0022] 1 shows an example of the configuration of an imaging device 10. The imaging device 10 is a digital camera with interchangeable lenses. The imaging device 10 is composed of a main body 11 and an imaging lens 12 that is interchangeably attached to the main body 11. The imaging lens 12 is attached to the front side of the main body 11 via a camera-side mount 11A and a lens-side mount 12A.
[0023] The main body 11 is provided with an operation unit 13 including a dial, a release button, etc. Operation modes of the imaging device 10 include, for example, a still image capturing mode, a video capturing mode, and an image display mode. The operation unit 13 is operated by the user when setting the operation mode. The operation unit 13 is also operated by the user when starting to capture a still image or a video. The operation unit 13 includes a touch panel provided on a display 15, etc., which will be described later.
[0024] The imaging device 10 is also provided with a subject tracking function that, in video imaging mode, tracks a subject designated by the user as a tracking target. The subject tracking function can also be activated when a live view image is displayed before still image capture or video capture. The imaging device 10 is an example of an "estimation device" according to the technology of the present disclosure.
[0025] The main body 11 is also provided with a viewfinder 14. Here, the viewfinder 14 is a hybrid viewfinder (registered trademark). A hybrid viewfinder is a viewfinder that selectively uses, for example, an optical viewfinder (hereinafter referred to as "OVF") and an electronic viewfinder (hereinafter referred to as "EVF"). The user can observe an optical image or a live view image of the subject displayed by the viewfinder 14 through a viewfinder eyepiece (not shown).
[0026] A display 15 is provided on the rear side of the main body 11. Images based on image signals obtained by imaging, various menu screens, etc. are displayed on the display 15. The user can also observe a live view image displayed on the display 15 instead of the viewfinder 14.
[0027] The main body 11 and the imaging lens 12 are electrically connected by electrical contacts 11B provided on the camera-side mount 11A coming into contact with electrical contacts 12B provided on the lens-side mount 12A.
[0028] The imaging lens 12 includes an objective lens 30, a focus lens 31, a rear lens 32, and an aperture 33. The components are arranged along the optical axis A of the imaging lens 12 in the following order from the objective side: objective lens 30, aperture 33, focus lens 31, and rear lens 32. teeth , constitute the imaging optical system. The type, number and arrangement order of the lenses that constitute the imaging optical system are not limited to the example shown in FIG.
[0029] The imaging lens 12 also has a lens drive control unit 34. The lens drive control unit 34 is configured with, for example, a CPU, RAM, and ROM. The lens drive control unit 34 is electrically connected to a processor 40 in the main body 11 via electrical contacts 12B and 11B.
[0030] The lens drive control unit 34 drives the focus lens 31 and the diaphragm 33 based on a control signal transmitted from the processor 40. The lens drive control unit 34 controls the drive of the focus lens 31 based on a control signal for focus control transmitted from the processor 40 in order to adjust the focus position of the imaging lens 12. The processor 40 may also perform focus control based on an estimation result R that indicates the position of a subject to be tracked, which will be described later.
[0031] The diaphragm 33 has an aperture whose diameter is variable around the optical axis A. The lens drive control unit 34 controls the drive of the diaphragm 33 based on an aperture adjustment control signal sent from the processor 40 in order to adjust the amount of light incident on the light receiving surface 20A of the image sensor 20.
[0032] Also provided inside the main body 11 are an image sensor 20, a processor 40, and a memory 42. The operations of the image sensor 20, the memory 42, the operation unit 13, the viewfinder 14, and the display 15 are controlled by the processor 40.
[0033] The processor 40 is configured by, for example, a CPU, a RAM, a ROM, etc. In this case, the processor 40 executes various processes based on a program 43 stored in a memory 42. The processor 40 may be configured by a collection of multiple IC chips.
[0034] The memory 42 also stores a first model M1 and a second model M2 that have undergone machine learning for object tracking. As will be described in detail later, the first model M1 and the second model M2 are configured by neural networks, and the second model M2 is larger in scale than the first model M1. "Large scale" refers to a large number of layers (convolutional layers, pooling layers, fully connected layers, etc.) that configure the neural network, and / or a large layer size (the number of neurons that configure the layer). The first model M1 is small in scale, so it can quickly estimate the object to be tracked, but has low estimation accuracy. Conversely, the second model M2 is large in scale, so it can slowly estimate the object to be tracked, but has high object tracking accuracy.
[0035] The imaging sensor 20 is, for example, a CMOS image sensor. The imaging sensor 20 is disposed so that its optical axis A is perpendicular to the light-receiving surface 20A and is located at the center of the light-receiving surface 20A. Light (subject image) that has passed through the imaging lens 12 is incident on the light-receiving surface 20A. A plurality of pixels that generate image signals by performing photoelectric conversion are formed on the light-receiving surface 20A. The imaging sensor 20 generates and outputs image signals by photoelectrically converting the light incident on each pixel. The imaging sensor 20 is an example of an "imaging element" according to the technology of the present disclosure.
[0036] A Bayer color filter array is arranged on the light receiving surface of the image sensor 20, with a color filter of any one of R (red), G (green), and B (blue) arranged opposite each pixel. Note that some of the pixels arranged on the light receiving surface of the image sensor 20 may be phase difference pixels for performing focus control.
[0037] Fig. 2 shows an example of the functional configuration of the processor 40. The processor 40 realizes various functional units by executing processes in accordance with a program 43 stored in a memory 42. As shown in Fig. 2, for example, the processor 40 realizes a main control unit 50, an imaging control unit 51, an image processing unit 52, a tracking target determination unit 53, a reference image creation unit 54, a model selection unit 55, an image input unit 56, an estimation unit 57, and a display control unit 58.
[0038] The main control unit 50 performs overall control of the operation of the imaging device 10 based on instruction signals input from the operation unit 13. The imaging control unit 51 controls the imaging sensor 20 to execute imaging processing that causes the imaging sensor 20 to perform imaging operations. The imaging control unit 51 drives the imaging sensor 20 in a still image imaging mode or a video imaging mode. The imaging sensor 20 outputs an imaging signal RD generated by the imaging operation. The imaging signal RD is so-called RAW data.
[0039] The image processing unit 52 performs a receiving process to receive the imaging signal RD output from the imaging sensor 20. The image processing unit 52 also performs image processing, including demosaic processing, on the received imaging signal RD to generate a captured image PD. For example, the captured image PD is a color image in which each pixel is represented by the three primary colors R, G, and B. More specifically, for example, the captured image PD is a 24-bit color image in which each R, G, and B signal contained in one pixel is represented by 8 bits.
[0040] The tracking target determination unit 53 performs a determination process to determine a subject designated by the user as a tracking target. For example, the user uses the operation unit 13 to designate a subject to be tracked from within the captured image PD displayed on the display 15. The tracking target determination unit 53 determines the subject designated by the user from within the captured image PD as the tracking subject to be tracked.
[0041] If the imaging device 10 has a subject detection function for detecting a subject based on the captured image PD, the tracking target determination unit 53 may determine a specific subject detected by the subject detection function as the tracking subject.
[0042] The reference image creation unit 54 performs a creation process to create, based on the captured image PD, a first reference image T1 for a first model including the tracking subject determined by the tracking target determination unit 53, and a second reference image T2 for a second model including the tracking subject. The creation process in this embodiment corresponds to the "first creation process" according to the technology of the present disclosure.
[0043] The reference image creation unit 54 creates a first reference image T1 and a second reference image T2 by cutting out an area including the tracking subject from within the captured image PD. The second reference image T2 is a reference image for the second model M2, which is larger in scale than the first model M1, and therefore has a higher resolution than the first reference image T1. Note that high resolution refers to a large amount of information, such as a large number of pixels in the image, a large amount of data for high-frequency components, or a large number of bits for each pixel that makes up the image. Hereinafter, when there is no need to distinguish between the first reference image T1 and the second reference image T2, they will simply be referred to as reference images. The reference image is a so-called template.
[0044] The model selection unit 55 performs a selection process to select one of the first model M1 and the second model M2 stored in the memory 42 as the selected model based on the factor information. In this embodiment, the model selection unit 55 performs the selection process using the value of the frame rate as factor information. The frame rate is the reciprocal of the repetition period of the imaging operation by the image sensor 20.
[0045] The frame rate value is changed, for example, by a user performing a setting operation using the operation unit 13. In addition, the frame rate value may be reduced by selecting a synthesis mode that synthesizes images of multiple frames in order to increase the brightness of the image.
[0046] The first model M1 has high-speed estimation processing but low estimation accuracy, and is therefore suitable for tracking a subject whose shape or blurring amount is small between frames. When the frame rate is high, high-speed subject tracking processing is required, but the time difference between frames is small and the subject's shape or blurring amount is small, so the model selection unit 55 selects the first model M1 as the selected model.
[0047] On the other hand, the second model M2 has a slow estimation process but high estimation accuracy, and is therefore suitable for tracking a subject whose shape changes or the amount of blurring is large between frames. When the frame rate is low, high-speed subject tracking processing is not necessary, but the time difference between frames is large, resulting in a large change in the shape of the subject or a large amount of blurring, so the model selection unit 55 selects the second model M2 as the selected model.
[0048] The image input unit 56 performs an input process of inputting the captured image PD represented by the imaging signal RD to the selected model selected by the model selection unit 55. In this embodiment, Image input unit 56 The captured image PD input to the selection model is a search image for searching for a tracking subject included in a reference image.
[0049] Furthermore, the image input unit 56 changes the resolution of the captured image PD to be input to the selected model in accordance with the resolution of the reference image to be input to the selected model. When the selected model is the second model M2, the image input unit 56 increases the resolution of the captured image PD compared to when the selected model is the first model M1.
[0050] The estimation unit 57 performs estimation processing to estimate the position of the tracking subject from the captured image PD using the selected model selected by the model selection unit 55 and the reference image for that selected model. Specifically, when the model selection unit 55 selects the first model M1, the estimation unit 57 inputs the first reference image T1 into the selected model. On the other hand, when the model selection unit 55 selects the second model M2, the estimation unit 57 inputs the second reference image T2 into the selected model.
[0051] The selection model outputs a score map SM that indicates the similarity of each region in the captured image PD with the reference image. The estimation unit 57 outputs information on the position with the highest score (i.e., the highest similarity) in the score map SM to the display control unit 58 as an estimation result R of the position of the tracking subject.
[0052] The display control unit 58 displays the estimation result R together with the captured image PD on the display 15. Specifically, the display control unit 58 displays the position of the tracking subject in the captured image PD so that the position can be recognized, based on the estimation result R. For example, the display control unit 58 displays the tracking subject in the captured image PD so that the position can be recognized. A rectangular frame is displayed around the
[0053] Fig. 3 conceptually shows an example of the process of determining a tracking subject and the process of creating a reference image. In Fig. 3, an area S is an area designated as a tracking target from within a captured image PD by a user using the operation unit 13. The tracking target determination unit 53 determines a subject included in the designated area S as a tracking subject H.
[0054] The reference image creation unit 54 creates a first reference image T1 by cutting out an area including the tracking subject H from the captured image PD and reducing the resolution of the cut-out image (in other words, the resolution of the second reference image T2 is higher than the resolution of the first reference image T1). The reference image creation unit 54 also creates a second reference image T2 by cutting out an area including the tracking subject H from the captured image PD.
[0055] 4 shows an example of the configuration of the first model M1. The first model M1 is composed of a first convolutional network (hereinafter referred to as the first CNN) 61A, a second convolutional network (hereinafter referred to as the second CNN) 62A, and a convolution calculation unit 63A.
[0056] The first CNN 61A is composed of multiple convolution layers and multiple pooling layers. Similarly, the second CNN 62A is composed of multiple convolution layers and multiple pooling layers. The convolution operation unit 63A is composed of multiple fully connected layers.
[0057] A first reference image T1 is input to the first CNN 61A. A captured image PD is input to the second CNN 62A. The first CNN 61A converts the input first reference image T1 into a feature map FM1 and outputs it. The second CNN 62A converts the input captured image PD into a feature map FM2 and outputs it. The feature maps FM1 and FM2 are input to a convolution calculation unit 63A.
[0058] The first CNN 61A and the second CNN 62A have the same configuration, but the size of the input layer to which an image is input corresponds to the size (number of neurons) of the input image. That is, the size of the input layer differs between the first CNN 61A and the second CNN 62A.
[0059] The convolution calculation unit 63A generates a score map SM by convolving the feature map FM2 with the feature map FM1 as a kernel, and outputs the generated score map SM to the estimation unit 57. The score map SM is an image that represents the similarity between each region in the captured image PD and the first reference image T1. The higher the similarity, the higher the score.
[0060] 5 shows an example of the configuration of the second model M2. The second model M2 is composed of a first CNN 61B, a second CNN 62B, and a convolution operation unit 63B. The first CNN 61B, the second CNN 62B, and the convolution operation unit 63B each have a larger number of layers than the first CNN 61A, the second CNN 62A, and the convolution operation unit 63A. Note that the first CNN 61B, the second CNN 62B, and the convolution operation unit 63B may each have a larger layer size than the first CNN 61A, the second CNN 62A, and the convolution operation unit 63A.
[0061] The second model M2 has the same configuration as the first model M1, except that it has a larger number of layers and / or larger layer sizes. A larger number of layers refers to a larger number of convolution layers or pooling layers. A larger layer size refers to a larger number of operations or a larger amount of operations in the convolution layers or pooling layers.
[0062] The first CNN 61B receives input of a second reference image T2. The second CNN 62B receives input of a captured image PD. The first CNN 61B converts the input second reference image T2 into a feature map FM1 and outputs it. The second CNN 62B converts the input captured image PD into a feature map FM2 and outputs it. The feature maps FM1 and FM2 are input to a convolution calculation unit 63B.
[0063] The convolution operation unit 63B generates a score map SM by convolving the feature map FM2 with the feature map FM1 as a kernel, and outputs the generated score map SM to the estimation unit 57.
[0064] FIG. 6 shows an example of training data used in machine learning of the first model M1. The machine learning of the first model M1 is performed using a pair of two frames selected from a video. Specifically, machine learning is performed by inputting training data into the first model M1, the training data being a pair of a first reference image T1 generated from the first frame and a captured image PD generated from the second frame. For the machine learning of the first model M1, it is preferable to use two frames with a small time difference and small change in the shape of the subject.
[0065] 7 shows an example of training data used in machine learning of the second model M2. The machine learning of the second model M2 is performed using a pair of two frames selected from a video. Specifically, machine learning is performed by inputting training data, which is a pair of a second reference image T2 generated from the first frame and a captured image PD generated from the second frame, into the second model M2. For the machine learning of the second model M2, it is preferable to use two frames with a large time difference and large changes in the shape of the subject.
[0066] 8 shows an example of the score map SM. As shown in Fig. 8, the estimation unit 57, for example, identifies an area U in the score map SM that includes a position with the highest score, and outputs position information of the identified area U to the display control unit 58 as an estimation result R.
[0067] FIG. 9 is a flowchart illustrating the processing procedure of the subject tracking function when capturing a moving image or displaying a live view image.
[0068] The main control unit 50 determines whether or not a user has operated the operation unit 13 to issue an instruction to start video capture or live view image display (step S10). If a start instruction has been issued (step S10: YES), the main control unit 50 controls the imaging control unit 51 to cause the imaging sensor 20 to perform an imaging operation and acquires the imaging signal RD output from the imaging sensor 20 (step S11). The display control unit 58 causes the display 15 to display the captured image PD generated by the image processing unit 52 based on the imaging signal RD (step S12).
[0069] The main control unit 50 determines whether or not the user has specified an area to be tracked within the captured image PD using the operation unit 13 (step S13). If the user has not specified an area (step S13: NO), the main control unit 50 returns the process to step S11 and causes the image sensor 20 to perform an image capturing operation. The processes of steps S11 to S12 are repeatedly executed until it is determined in step S13 that the user has specified an area.
[0070] When the user designates an area (step S13: YES), the main control unit 50 causes the tracking target determination unit 53 to determine a tracking target (step S14). In step S14, the tracking target determination unit 53 determines a subject included in the designated area as the tracking subject H.
[0071] The reference image creation unit 54 cuts out an area including the tracking subject H from the captured image PD, and creates a first reference image T1 and a second reference image T2 (step S15). Here, the second reference image T2 has a higher resolution than the first reference image T1.
[0072] The model selection unit 55 selects either the first model M1 or the second model M2 as the selected model using the value of the frame rate as factor information (step S16). In step S16, the model selection unit 55 selects the first model M1 as the selected model if the value of the frame rate is equal to or greater than a certain value, and selects the second model M2 as the selected model if the value of the frame rate is less than the certain value.
[0073] The main control unit 50 controls the imaging control unit 51 to cause the imaging sensor 20 to perform an imaging operation and acquires the imaging signal RD output from the imaging sensor 20 (step S17). The image input unit 56 inputs the captured image PD generated by the image processing unit 52 based on the imaging signal RD to the selected model at a resolution corresponding to the selected model selected by the model selection unit 55 (step S18).
[0074] The estimation unit 57 inputs the reference image for the selected model, selected by the model selection unit 55 from the first reference image T1 and the second reference image T2, into the selected model, estimates the position of the tracking subject from the imaging signal RD based on the score map SM output from the selected model, and outputs the estimation result R to the display control unit 58 (step S19). The display control unit 58 displays the estimation result R together with the captured image PD on the display 15 (step S20).
[0075] The main control unit 50 determines whether a predetermined termination condition is met (step S21). The termination condition may be, for example, that the user has performed an operation to stop video capture using the operation unit 13. If the termination condition is not met (step S21: NO), the main control unit 50 returns the process to step S17 and causes the image sensor 20 to perform an image capture operation. The processes of steps S17 to S20 are repeatedly executed until it is determined in step S21 that the termination condition is met. If the termination condition is met (step S21: YES), the main control unit 50 terminates the process.
[0076] In the above flowchart, steps S11 and S17 correspond to the "receiving process" according to the technology of the present disclosure. Step S14 corresponds to the "determining process" according to the technology of the present disclosure. Step S15 corresponds to the "first creating process" according to the technology of the present disclosure. Step S16 corresponds to the "selecting process" according to the technology of the present disclosure. Step S18 corresponds to the "input process" according to the technology of the present disclosure. Step S19 corresponds to the "estimating process" according to the technology of the present disclosure.
[0077] As described above, according to the technology of the present disclosure, when the frame rate is high, a small-scale first model M1 is selected with an emphasis on real-time performance, and when the frame rate is low, a large-scale second model M2 is selected with an emphasis on object tracking accuracy. When the frame rate is high, the change in shape or amount of blur of the object being tracked between frames is small, so even the small-scale first model M1 maintains a constant level of object tracking accuracy. Furthermore, when the frame rate is low, the frame period is long, so even the large-scale second model M2 maintains a constant level of real-time performance. In this way, according to the technology of the present disclosure, it is possible to achieve both object tracking accuracy and real-time performance.
[0078] Furthermore, according to the technology disclosed herein, a first reference image T1 for the first model M1 and a second reference image T2 for the second model M2 are created, and the estimation process is performed using the reference images according to the selected model, so there is no need to recreate the reference images when switching the selected model, and real-time performance can be maintained even when the selected model is switched.
[0079] [Variations] Various modifications of the above embodiment are shown below, and in each modification, only the differences from the above embodiment will be described.
[0080] In the above embodiment, the image input unit 56 inputs the entire captured image PD as a search image to the selection model, but an image cut out from the captured image PD may also be input as a search image to the selection model. For example, as shown in Fig. 10, the image input unit 56 sets a search range to include an area U containing the tracking subject estimated by the estimation unit 57 in the previous frame period, and cuts out an image within the search range from the captured image PD obtained in the current frame period and inputs it to the selection model. By limiting the search range in this way, the processing speed by the selection model is improved.
[0081] (Creating a reference image) In the above embodiment, the reference image creation unit 54 executes a creation process (first creation process) that creates a first reference image T1 and a second reference image T2 from the captured image PD. The reference image creation unit 54 may be configured to execute a second creation process that creates the first reference image T1 but does not create the second reference image T2 instead of the first creation process. For example, the reference image creation unit 54 selectively executes the first creation process or the second creation process based on the value of the frame rate.
[0082] Fig. 11 shows a reference image creation process according to a modified example. The process shown in Fig. 11 is executed, for example, in step S15 of the flowchart shown in Fig. 9. The reference image creation unit 54 determines whether the frame rate value is less than a certain value (step S30). If the frame rate value is less than a certain value (step S30: YES), the reference image creation unit 54 executes a first creation process (step S31). On the other hand, if the frame rate value is equal to or greater than a certain value (step S30: NO), the reference image creation unit 54 executes a second creation process (step S32).
[0083] That is, the first creation process is executed when the second model M2 is selected as the selected model by the model selection unit 55, and the second creation process is executed when the first model M1 is selected as the selected model by the model selection unit 55. When the frame rate is high, the processing can be speeded up by not creating the second reference image T2.
[0084] (Model Selection) In the above embodiment, the model selection unit 55 performs the selection process using the value of the frame rate as factor information, but the factor information is not limited to the value of the frame rate. For example, the model selection unit 55 may perform the selection process using the type of tracking subject determined by the tracking target determination unit 53 as factor information.
[0085] The first model M1 has a high speed estimation process but low estimation accuracy, and is therefore suitable for tracking a subject whose shape changes little between frames. A subject whose shape changes little between frames is an object with high rigidity, such as a vehicle or an aircraft. On the other hand, the second model M2 has a low speed estimation process but high estimation accuracy, and is therefore suitable for tracking a subject whose shape changes little between frames. A subject whose shape changes little between frames is an object with low rigidity, such as a human or an animal. The shape of a human or an animal is easily changed by the movement of its limbs, etc.
[0086] In the above embodiment, the selected model is not changed after the model selection unit 55 performs the selection process for the selected model, but the selected model may be changed according to factor information that changes during the tracking operation of the subject. For example, as shown in the flowchart in Fig. 12, if the termination condition is not satisfied (step S21: NO), the main control unit 50 returns the process to step S16 and causes the model selection unit 55 to perform the selection process for the selected model again. In this way, the model selection unit 55 may be caused to repeatedly perform the selection process until the termination condition is satisfied.
[0087] In this case, it is preferable that the model selection unit 55 performs the selection process using the movement speed of the tracking subject as factor information. When the movement speed of the tracking subject is high, the shape of the tracking subject changes significantly between frames. For this reason, it is preferable that the model selection unit 55 selects the second model M2 as the selected model when the movement speed of the tracking subject is equal to or greater than a certain value, and selects the first model M1 as the selected model when the movement speed of the tracking subject is less than the certain value.
[0088] It is also preferable that the model selection unit 55 performs the selection process using the degree of change in the form of the tracked subject as factor information. The degree of change in the form of the tracked subject is, for example, the degree of change in shape or the degree of change in color. When the degree of change in the form of the tracked subject between frames is equal to or greater than a certain value, the model selection unit 55 selects a model. to It is preferable to select the second model M2 as the selected model, and select the first model M1 as the selected model when the degree of change in the form of the tracking subject between frames is less than a certain value.
[0089] It is also preferable that the model selection unit 55 performs the selection process using the score obtained from the score map SM output from the selected model as factor information. For example, when the model selection unit 55 selects the first model M1 as the selected model and the maximum value of the score becomes less than the threshold, it determines that the tracking accuracy has decreased and selects the second model M2 with higher tracking accuracy as the selected model.
[0090] (Updated reference image) In the above embodiment, in order to speed up the subject tracking operation, the reference image created by the reference image creation unit 54 is not updated until the subject tracking operation is completed. This is because if the reference image is updated when the tracking subject undergoes a change in posture, such as rotation, or when occlusion (i.e., intersection of objects) occurs during the subject tracking operation, the possibility of erroneously tracking an object other than the tracking subject increases. Here, "updating" refers to the reference image creation unit 54 creating a new reference image.
[0091] As described above, it is preferable that the reference image not be updated in principle, but the reference image creation unit 54 may update the reference image when a specific condition is met.
[0092] For example, when the selected model is switched from one of the first model M1 and the second model M2 to the other, the reference image creation unit 54 executes a first update process for updating the first reference image T1 and the second reference image T2. Specifically, immediately after step S16 of the flowchart shown in FIG. 12, the reference image creation unit 54 executes the first update process of the reference image shown in FIG. 13.
[0093] 12, the reference image creation unit 54 determines whether the selected model has been changed by the model selection unit 55 in step S16 (step S40). If the selected model has not been changed (step S40: NO), the reference image creation unit 54 does not update the reference image. On the other hand, if the selected model has been changed (step S40: YES), the reference image creation unit 54 updates the reference image (step S41). In step S41, the reference image creation unit 54 creates a first reference image T1 and a second reference image T2 based on an image cut out from a region U (see FIG. 8) identified by the estimation unit 57 in a captured image PD obtained in the immediately preceding frame period.
[0094] Note that if the reference image is updated when the score of region U is low, the reliability of the updated reference image will be low, so it is preferable to update the reference image on the condition that the score is equal to or greater than a certain value. For example, as shown in the flowchart of FIG. 14, when the selected model is changed (step S40: YES), the reference image creation unit 54 determines whether the score (e.g., the maximum value) of region U identified by the estimation unit 57 is equal to or greater than a certain value (step S42). When the score is not equal to or greater than the certain value (step S42: NO), the reference image creation unit 54 does not update the reference image. On the other hand, when the score is equal to or greater than the certain value (step S42: YES), the reference image creation unit 54 updates the reference image (step S41).
[0095] Furthermore, the reference image creation unit 54 may execute a second update process that updates the reference image based on a change in the size of the tracked subject within the angle of view of the captured image PD. A change in the size of the tracked subject occurs, for example, when the tracked subject approaches or moves away from the image capture device 10. When the size of the tracked subject changes, the similarity with the reference image decreases, thereby reducing the accuracy of subject tracking. For this reason, it is preferable that the reference image creation unit 54 update the reference image when the size of the tracked subject changes by more than a certain value, using the size of the tracked subject in the reference image as a reference. The size of the tracked subject can be detected using the subject detection result obtained by the subject detection function.
[0096] The size of the tracking subject within the angle of view of the captured image PD depends on the distance from the imaging device 10 to the tracking subject. Therefore, the reference image creation unit 54 may update the reference image based on distance information detected by the phase difference pixels of the imaging sensor 20 when the distance from the imaging device 10 to the tracking subject has changed by a certain value or more.
[0097] Furthermore, the size of the tracking subject may change due to a change in the imaging magnification of the imaging device 10. Therefore, it is preferable that the reference image creation unit 54 updates the reference image when the imaging magnification changes by a certain value or more after creating the reference image. The imaging magnification may change not only due to optical zoom but also due to electronic zoom. For example, the imaging magnification may change when the user operates the operation unit 13.
[0098] Fig. 15 is a flowchart showing an example of the second update process. The reference image creation unit 54 executes the second update process of the reference image shown in Fig. 15 during the subject tracking operation. In Fig. 15, the reference image creation unit 54 determines whether the imaging magnification has changed by a certain value or more (step S50). If the imaging magnification has not changed by a certain value or more (step S50: NO), the reference image creation unit 54 does not update the reference image. On the other hand, if the imaging magnification has changed by a certain value or more (step S50: YES), the reference image creation unit 54 updates the reference image (step S51).
[0099] In the second update process, it is also preferable that the reference image creation unit 54 updates the reference image on the condition that the score is equal to or greater than a certain value.
[0100] Furthermore, the reference image creation unit 54 may periodically update the reference image during the subject tracking operation. For example, the reference image creation unit 54 updates the reference image once every several hundred frames during the subject tracking operation. In this case, too, it is preferable that the reference image creation unit 54 updates the reference image on the condition that the score is equal to or greater than a certain value.
[0101] The technology of the present disclosure is not limited to digital cameras, but can also be applied to electronic devices such as smartphones and tablet terminals that have an imaging function.
[0102] In the above embodiment, the hardware structure of the control unit, with processor 40 being an example, can use the following various processors. The above various processors include a CPU, which is a general-purpose processor that functions by executing software (programs), as well as a processor such as an FPGA, whose circuit configuration can be changed after manufacture. FPGAs include dedicated electrical circuits, such as PLDs or ASICs, which are processors with a circuit configuration designed specifically to execute specific processes.
[0103] The control unit may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, multiple control units may be configured with a single processor.
[0104] There are several possible examples of configuring multiple control units with a single processor. A first example is a form in which one processor is configured with a combination of one or more CPUs and software, as typified by computers such as client and server, and this processor functions as multiple control units. A second example is a form in which a processor is used to realize the functions of an entire system including multiple control units on a single IC chip, as typified by system-on-chip (SOC). In this way, the control unit can be configured as a hardware structure using one or more of the various processors described above.
[0105] Furthermore, more specifically, the hardware structure of these various processors can be an electric circuit that combines circuit elements such as semiconductor elements.
[0106] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[0107] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
Claims
1. a memory that stores a first model and a second model that have been machine-learned for object tracking; a processor that receives an imaging signal from the imaging element; An estimation device comprising: The processor: a determination process for determining a tracking subject to be tracked; a first generation process for generating a first reference image for the first model including the tracking subject and a second reference image for the second model including the tracking subject based on the imaging signal; a selection process of selecting one of the first model and the second model as a selected model based on factor information; an input process of inputting a captured image represented by the imaging signal into the selection model; an estimation process of estimating a position of the tracking subject from the captured image by using the selection model and a reference image for the selection model out of the first reference image and the second reference image; The estimation device is configured to perform the following:
2. The second model has a larger number of layers or a larger layer size than the first model. The estimation device according to claim 1 .
3. The second reference image has a higher resolution than the first reference image. The estimation device according to claim 2 .
4. The factor information is the type of the tracking subject, the moving speed of the tracking subject, or the degree of change in the shape of the tracking subject. The estimation device according to claim 3 .
5. the factor information is a value of a frame rate of the captured image input to the selection model; The estimation device according to claim 3 .
6. The processor: a second generation process for generating the first reference image but not the second reference image can be executed instead of the first generation process; selecting the first creation process or the second creation process based on the value of the frame rate; The estimation device according to claim 5 .
7. The processor: a first update process for updating the first reference image and the second reference image when the selected model is switched from one of the first model and the second model to the other in the selection process; The estimation device according to claim 1 .
8. The processor: a second update process is executed to update the first reference image and the second reference image based on a change in size of the tracking subject within an angle of view of the captured image. The estimation device according to any one of claims 1 to 7.
9. The processor: The second update process is executed based on a change in the imaging magnification of an imaging device having the imaging element. The estimation device according to claim 8 .
10. A method for driving an estimation device including a memory that stores a first model and a second model that have been machine-learned for object tracking, the method comprising: a receiving step of receiving an imaging signal from the imaging element; a determination step of determining a tracking subject to be tracked; a first creation step of creating a first reference image for the first model including the tracking subject and a second reference image for the second model including the tracking subject based on the imaging signal; a selection step of selecting one of the first model and the second model as a selected model based on factor information; an input step of inputting a captured image represented by the imaging signal into the selection model; an estimation step of estimating a position of the tracking subject from the captured image by using the selection model and a reference image for the selection model out of the first reference image and the second reference image; A method for driving an estimation device, comprising:
11. A program for operating an estimation device including a memory storing a first model and a second model that have been machine-learned for object tracking, a receiving process for receiving an imaging signal from the imaging element; a determination process for determining a tracking subject to be tracked; a first generation process for generating a first reference image for the first model including the tracking subject and a second reference image for the second model including the tracking subject based on the imaging signal; a selection process of selecting one of the first model and the second model as a selected model based on factor information; an input process of inputting a captured image represented by the imaging signal into the selection model; an estimation process of estimating a position of the tracking subject from the captured image by using the selection model and a reference image for the selection model out of the first reference image and the second reference image; A program that causes the estimation device to execute the above.
Citation Information
Patent Citations
Key point detection method, device and equipment and readable medium
CN109584276A
Generating a customized machine-learning model to perform tasks using artificial intelligence
CN112088386A
System and method for adverse event detection or severity estimation from surgical data
US20200265273A1