Imaging device, method for controlling the imaging device, and program
The imaging device addresses the challenge of tracking multiple moving subjects by using motion vector detection and weight-based control to accurately track and focus on primary subjects, reducing misdetection and mistracking.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-02-16
- Publication Date
- 2026-03-16
AI Technical Summary
Existing imaging devices struggle with detecting and tracking multiple moving subjects, especially when subjects' movements are large, leading to issues like mistracking of background subjects and misdetection of stationary subjects as moving subjects.
The imaging device includes a motion vector detection unit, subject detection unit, and a control unit that calculates the centroid position of multiple moving subjects based on motion vectors and weight coefficients, using a pan-tilt drive unit for tracking control, adjusting thresholds based on shaking state and pan-tilt angular velocity to accurately track multiple subjects.
Enables effective tracking of multiple moving subjects, reducing misdetection and mistracking, and maintaining focus on primary subjects even in complex scenes.
Smart Images

Figure 0007830155000003 
Figure 0007830155000004 
Figure 0007830155000005
Abstract
Description
Technical Field
[0001] The present invention relates to an imaging device, a control method for an imaging device, and a program.
Background Art
[0002] Patent Document 1 discloses an imaging device that detects a moving subject and performs automatic tracking.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Non-Patent Documents
[0004]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] In the imaging device disclosed in Patent Document 1, detecting moving subjects is difficult when the subject's movement is large, such as in sports scenes. When imaging a moving subject, one can consider a method of detecting the moving subject on the image and centering the image on the detected moving subject. However, if automatic tracking is limited to a specific moving subject among multiple moving subjects, other moving subjects may move out of the field of view. Furthermore, it is difficult to perform appropriate control due to the mistracking of background moving subjects other than the target moving subject, or the misdetection or mistracking of stationary subjects as moving subjects.
[0006] Therefore, the present invention aims to provide an imaging device that can be controlled in a manner suitable for imaging multiple moving subjects. [Means for solving the problem]
[0007] An imaging device as one aspect of the present invention includes an imaging unit, a motion vector detection unit that detects motion vectors based on image data output from the imaging unit, a subject detection unit that detects a plurality of moving subjects based on the motion vectors, and the imaging unit To pan and tilt It has a drive unit that drives the object and a control unit that performs tracking control by controlling the drive unit, and the control unit is Based on the position of each of the multiple moving subjects in the image data and the weight coefficient of each of the multiple moving subjects, the centroid position of the multiple moving subjects in the image data is determined. Calculate and the above Center of gravity position The drive unit is controlled based on this.
[0008] Other objects and features of the present invention are described in the following embodiments. [Effects of the Invention]
[0009] According to the present invention, it is possible to provide an imaging device that can be controlled in a manner suitable for imaging multiple moving subjects. [Brief explanation of the drawing]
[0010] [Figure 1] This is an external view of the imaging device in the first embodiment. [Figure 2] This is a block diagram of the imaging device in the first embodiment. [Figure 3]This is an explanatory diagram of communication between the imaging device and external equipment in the first embodiment. [Figure 4] This is a flowchart showing the tracking operation process in the first embodiment. [Figure 5] This is an explanatory diagram of the motion object detection process in the first embodiment. [Figure 6] This is an explanatory diagram of the motion object detection process in the first embodiment. [Figure 7] This is an explanatory diagram of the frequency distribution processing in the first embodiment. [Figure 8] This is an explanatory diagram of the tracking operation process in the first embodiment. [Figure 9] This is an explanatory diagram of the operation of the imaging device in the first embodiment. [Figure 10] This is a flowchart showing the imaging process in the second embodiment. [Figure 11] This is an example of a weight calculation table in the first embodiment. [Modes for carrying out the invention]
[0011] The embodiments of the present invention will now be described in detail with reference to the drawings. The configurations described in each embodiment are merely illustrative, and the present invention is not limited to the configurations described in each embodiment.
[0012] (First Embodiment) <Configuration of imaging device 101> First, referring to FIG. 1, the external configuration of the imaging device 101 in the first embodiment will be described. FIGS. 1(a) to (d) are external views of the imaging device 101. The imaging device 101 is provided with a power switch (not shown). As shown in FIG. 1(a), the imaging device 101 has a lens barrel 102 including an imaging lens group (imaging optical system) for imaging and an imaging element. The imaging device 101 to which the lens barrel 102 is attached has a mechanism that can be rotationally driven with respect to the fixed portion 103. The tilt rotation unit 104 includes a motor drive mechanism that rotates the imaging device 101 in the "pitch direction" (see FIG. 1(b)). The pan rotation unit 105 includes a motor drive mechanism that rotates the imaging device 101 in the "yaw direction" (see FIG. 1(b)).
[0013] The tilt rotation unit 104 and the pan rotation unit 105 constitute a pan-tilt unit that performs tilt rotation and pan rotation of the imaging device 101. Note that "pitch", "yaw", and "roll" shown in FIG. 1(b) are rotations about the X-axis, Y-axis, and Z-axis, respectively. The X-axis, Y-axis, and Z-axis are axes defined at the fixed position of the fixed portion 103.
[0014] The fixed portion 103 is provided with an angular velocity meter 106 and an accelerometer 107. Based on the outputs from the angular velocity meter 106 and the accelerometer 107, the vibration of the imaging device 101 can be detected. Then, the tilt rotation unit 104 and the pan rotation unit 105 are rotationally driven based on the detected shake angle. Thereby, the shake and tilt of the imaging device 101 can be corrected.
[0015] In FIGS. 1(a), (c), and (d), reference numeral 108 indicates the optical axis direction. FIG. 1(c) shows a state in which the imaging device 101 is tilt-rotated by 90 degrees from the state of FIG. 1(a). FIG. 1(d) shows a state in which the imaging device 101 is tilt-rotated by θ degrees (0 degrees < θ < 90 degrees) from the state of FIG. 1(a).
[0016] Next, the internal configuration of the imaging device 101 will be described with reference to Figure 2. Figure 2 is a block diagram of the imaging device 101. The control unit 223 includes, for example, a CPU (MPU), memory (DRAM, SRAM), non-volatile memory (EEPROM), etc. By executing a program, the control unit 223 realizes the control of each part of the imaging device 101, the control of data transfer between each part, and various other functions. The non-volatile memory 216 is an electrically erasable and recordable memory that stores parameters, programs, etc. necessary for the operation of the control unit 223. The control unit 223 also has a subject detection unit 225 that detects multiple moving subjects based on motion vectors.
[0017] The lens barrel 120 includes a zoom unit 201, a focus unit 203, and an imaging unit 206. The zoom unit 201 has a zoom lens that performs magnification. The zoom drive unit 202 drives and controls the zoom unit 201. In other words, the zoom drive unit 202 changes the focal length of the imaging optical system. The focus unit 203 has a lens that adjusts the focus. The focus drive unit 204 drives and controls the focus unit 203. The imaging unit 206 has an image sensor (not shown) on which an image of the subject is formed. The image sensor of the imaging unit 206 is a photoelectric conversion element such as a CMOS sensor or a CCD sensor, which receives light incident through the imaging optical system of the lens barrel 102 and outputs charge information corresponding to the amount of light received as analog image data (imaging image) to the image processing unit 207.
[0018] The image processing unit 207 performs image processing such as distortion correction, white balance adjustment, and color interpolation on the digital image data obtained from the analog image data, and outputs the digital image data. The image processing unit 207 also functions as a motion vector detection unit that detects motion vectors in the image data output from the imaging unit 206. The image recording unit 208 converts the digital image data output from the image processing unit 207 into a recording format such as JPEG. The converted image data is transmitted to the memory 215 and the video output unit 217.
[0019] The pan-tilt drive unit 205 is a rotational drive unit for driving the tilt rotation unit 104 and the pan rotation unit 105. In other words, the pan-tilt drive unit 205 rotates the imaging device 101, which is equipped with a lens barrel 102, in the tilt direction and the pan direction. The vibration detection unit 209 includes, for example, an angular velocity meter (gyro sensor) 106 for detecting the angular velocity of the imaging device 101 in the three axial directions, and an accelerometer (accelerometer sensor) 107 for detecting the acceleration of the imaging device 101 in the three axial directions. Based on the detected signals, the rotation angle of the imaging device 101, the amount of shift of the imaging device 101, etc., are calculated.
[0020] The operation unit 210 is provided for performing various operations on the imaging device 101, and includes, for example, a power button and a button for triggering imaging of the imaging device 101. When the power button is operated, power is supplied to the imaging device 101 and the imaging device 101 is started up. The audio input unit 213 uses a microphone provided on the imaging device 101 to pick up audio signals from the vicinity of the imaging device 101 and transmits the analog-to-digital converted digital audio signal to the audio processing unit 214. The audio processing unit 214 performs audio processing such as optimization processing of the received digital audio signal. The audio signal that has undergone various processing in the audio processing unit 214 is then sent to the memory 215 by the control unit 223. The memory 215 temporarily stores the image signal or audio signal obtained by the image processing unit 207 and the audio processing unit 214.
[0021] The image processing unit 207 reads the image signal temporarily stored in the memory 215 and generates a compressed image signal by encoding the image signal, etc. The audio processing unit 214 reads the audio signal temporarily stored in the memory 215 and generates a compressed audio signal by encoding the audio signal, etc. The control unit 223 transmits the compressed image signal and the compressed audio signal to the recording and playback unit 220.
[0022] The recording and playback unit 220 records the compressed image signal, compressed audio signal, or other image capture-related control data, etc., generated by the image processing unit 207 and the audio processing unit 214, respectively, onto the recording medium 221. If the audio signal is not compressed, the control unit 223 transmits the audio signal generated by the audio processing unit 214 and the compressed image signal generated by the image processing unit 207 to the recording and playback unit 220 for recording on the recording medium 221.
[0023] The recording medium 221 is a recording medium (such as an HD) built into the imaging device 101, or a recording medium (such as a USB memory stick or memory card) that can be attached to or removed from the imaging device 101. The recording medium 221 can record various types of data such as compressed image signals, compressed audio signals, or audio signals generated by the imaging device 101, and generally has a larger capacity than the non-volatile memory 216. Examples of recording medium 221 include hard disks, optical discs, magneto-optical discs, CD-Rs, DVD-Rs, magnetic tapes, non-volatile semiconductor memory, and flash memory.
[0024] The recording and playback unit 220 has the function of reading and playing back compressed image signals, compressed audio signals, audio signals, various data, or programs recorded on the recording medium 221. The compressed image signals and compressed audio signals read by the recording and playback unit 220 are sent by the control unit 223 to the image processing unit 207 and the audio processing unit 214. The image processing unit 207 and the audio processing unit 214 temporarily store the compressed image signals and compressed audio signals in the memory 215, decode them according to a predetermined procedure, and transmit the decoded signals to the video output unit 217.
[0025] The audio output unit 218 has a speaker built into the imaging device 101 and, for example, outputs a pre-set audio pattern from the speaker during imaging. The LED control unit 224 has multiple LEDs and controls the lighting of multiple LEDs according to a set lighting / flashing pattern, for example, during imaging. The video output unit 217 consists of, for example, a video output terminal and outputs an image signal to display video on a connected external display or the like. Alternatively, the audio output unit 218 and the video output unit 217 may be combined into a single terminal. In other words, an HDMI (High-Definition Multimedia Interface) terminal or the like may be used.
[0026] The communication unit 222 communicates between the imaging device 101 and external devices. The communication unit 222 transmits and receives data such as audio signals, image signals, compressed audio signals, or compressed image signals. It also has a function to transmit information indicating the internal state of the imaging device 101, such as error information, to external devices when the imaging device 101 detects an abnormal condition. The communication unit 222 is a wireless communication module such as an infrared communication module, a Bluetooth® communication module, a wireless LAN communication module, a WirelessUSB, or a GPS receiver.
[0027] <Communication with external devices> Next, referring to Figure 3, the communication between the imaging device 101 and the external device 301 will be explained. Figure 3 is an explanatory diagram of the communication between the imaging device 101 and the external device 301.
[0028] The imaging device 101 is a device having an imaging function, and the external device 301 is a smart device including a Bluetooth® communication module or a wireless LAN communication module. The external device 301 is, for example, a smartphone. The imaging device 101 and the external device 301 can communicate with each other via first communication 302 and second communication 303. For example, first communication 302 is wireless LAN communication compliant with the "IEEE802.11" standard series, and second communication 303 is communication with a master-slave relationship such as a control station and a slave station, such as "Bluetooth® Low Energy (BLE)".
[0029] Note that Wi-Fi and BLE are just examples of communication methods, and the imaging device 101 and the external device 301 have communication capabilities using multiple types of communication methods. For example, other communication methods may be used if it is possible to control one communication method within the relationship between a control station and a slave station. However, without loss of generality, the first communication method 302, such as Wi-Fi, is capable of faster communication than the second communication method 303, such as BLE. Also, the second communication method 303 is either lower power-consuming or has a shorter communication range than the first communication method 302, or at least one of the other.
[0030] <Tracking action processing> Next, the imaging process (tracking operation process) in this embodiment will be described with reference to Figure 4. Figure 4 is a flowchart of the tracking operation process (subject tracking control) in this embodiment. Each step in Figure 4 is mainly performed by the image processing unit 207 or the control unit 223.
[0031] First, in step S401, the image processing unit 207 generates an image processed for subject detection using the imaging signal captured by the imaging unit 206. Subject detection, such as people or objects, is performed on the generated image. When detecting a person, the face or body of the subject is detected. In the "face detection process," a pattern for determining a person's face is pre-set, and within the captured image, areas that match the pre-set pattern can be detected as a person's face image.
[0032] The image processing unit 207 also simultaneously calculates a "confidence score" indicating the likelihood that the subject is indeed a face. The "confidence score" is calculated based on factors such as the size of the "face region" in the image and the degree of agreement with the face pattern. Similarly, for object recognition, objects can be recognized by determining whether or not they match a pre-set pattern. By calculating an "evaluation value" for each image region of the recognized subject, the image region of the subject with the highest "evaluation value" can be determined as the "main subject region (specific subject)".
[0033] Next, in step S402, the control unit 223 uses the vibration detection unit 209 to acquire angular velocity outputs for the three axes from the angular velocity meter 106 set on the fixed unit 103. The control unit 223 also acquires the current angle position of the pan / tilt from the output of the encoders that can acquire the rotation angle, which are installed on the tilt rotation unit 104 and the pan rotation unit 105, respectively. The control unit 223 also acquires the "motion vector" calculated (detected) by the image processing unit 207 (vector information acquisition). As a method for detecting the "motion vector," the image processing unit 207 first divides the image (an image based on image data output from the imaging unit 206) into multiple regions. The image processing unit 207 then compares the image from the previous frame, which is stored in advance, with the current image (two consecutive images), and calculates the amount of motion of the image from the relative displacement information of the images. Once the "motion vector" has been acquired, the process proceeds to step S403.
[0034] Next, in step S403, the control unit 223 determines the shaking state (vibration state) of the imaging device 101 from the angular velocity output in three axes from the angular velocity meter 106 set on the fixed unit 103. For example, within a predetermined period TIMEA, the number of times the output of the angular velocity meter 106 exceeds the threshold Thresh1 can be counted, and if the count value exceeds the threshold Thresh2, it can be determined that the amount of shaking (vibration) is large. Alternatively, the thresholds can be set in stages, and the shaking state can be determined in stages according to the magnitude of the shaking. Alternatively, the amount of shaking can be calculated by filtering. For example, the output of the angular velocity meter 106 can be offset by cutting the low frequency band with a high-pass filter (HPF), and the signal obtained by converting the HPF angular velocity to an absolute value and passing it through a low-pass filter (LPF) can be calculated as the amount of shaking. By any of the above methods, it is possible to determine whether the shaking state of the imaging device 101 is large or small. Once the shaking state is determined, the process proceeds to step S404.
[0035] <Moving subject determination method> Next, in step S404, the subject detection unit 225 of the control unit 223 determines whether or not there is a "moving subject detection area" on the captured image based on the information acquired in step S402 (moving subject determination). Now, let's explain the moving subject determination. First, the control unit 223 determines whether or not there is a subject with a prominent feature (prominent subject) for each image frame from the image processing unit 207. In this embodiment, "prominence" is the degree of prominence of the feature, and "prominence" is determined based on hue, saturation, and brightness.
[0036] The more pronounced the distinction from the background, the higher the "strikingness." The method for calculating "strikingness" is disclosed, for example, in Non-Patent Document 2. "Strikingness" can be calculated using the known strikingness calculation method disclosed in Non-Patent Document 2.
[0037] In this embodiment, the control unit 223 determines whether or not there is a "moving subject detection area" based on the "prominent subject" within the image frame and the "motion vector detection position" in the image (moving subject detection process). The moving subject detection process will now be explained with reference to Figures 5(a) to (d) and Figure 6. Figures 5(a) to (d) and Figure 6 are explanatory diagrams of the moving subject detection process.
[0038] In Figure 5(a), reference numeral 501 indicates a "stationary person". The area indicated by reference numeral 504 is a subject in which face detection is possible. Reference numeral 502 indicates a subject whose face is hidden and therefore cannot be detected, and which is moving, resulting in significant image changes between frames. Reference numeral 503 indicates a subject in an area where characteristics such as hue, saturation, and brightness are prominent in the image, and which is a subject with almost no movement.
[0039] Figure 5(b) shows the results of extracting the "severity calculation area" for determining "severity" using the method described above, indicated by symbols 505 to 510. Furthermore, as indicated by symbol 511 in Figure 5(c), the detection of "motion vectors" is the amount of pixel movement between image frames in a region set at a specific location in an image. The detection positions for "motion vectors" are arranged to cover the entire image so that "motion vectors" can be detected. Within the "severity calculation area" (symbols 505 to 510), the "severity calculation area" in which the number of vectors whose amount of movement is greater than or equal to "threshold 1" is detected as "threshold 2" or more is determined to be a "moving subject detection area".
[0040] Furthermore, as indicated by symbols 508 to 510, if the "sampling area calculation areas" overlap or are very close together, and a "moving subject" is detected in each "sampling area calculation area," the following occurs. That is, the "moving subject detection area" is determined to be one area (symbol 512) in Figure 5(d). Here, within the image frame, the subject (symbol 501) and the subject (symbol 503) are not moving. Therefore, the detected amount of "motion vector" in the "sampling area calculation areas" of symbols 505, 506, and 507 is a very small value, and it is not determined to be a "moving subject detection area." Within the "sampling area calculation areas" of symbols 508, 509, and 510, if the number of vectors whose magnitude of movement is greater than or equal to "threshold 1" is detected to be greater than or equal to "threshold 2," it is determined to be a "moving subject detection area." Furthermore, as indicated by symbols 508, 509, and 510, the "severity calculation areas" overlap and are therefore judged as a single "moving subject detection area." As a result, there is only one "moving subject detection area" (symbol 512).
[0041] In this embodiment, as shown in Figure 1, the imaging device 101 has a pan-tilt mechanism. While "subject tracking" is performed by pan-tilt, even if the subject is not moving, driving the pan-tilt causes movement on the imaging surface. As a result, the output value at the detection position of each "motion vector" becomes large. Also, when imaging is performed while moving the imaging device 101 by hand, image blur occurs due to hand shake. This causes movement on the imaging surface even if the actual subject is not moving, and the output value at the detection position of each "motion vector" becomes large. Therefore, in this embodiment, the amount of "motion vector" after removing the image blur caused by pan-tilt driving and hand shake is calculated using the method shown in Figure 6, and a determination is made as to whether or not there is a "moving subject detection area" (moving subject detection process). This will be explained with reference to Figure 6.
[0042] The gyro output 601, which is the output from the gyroscope, is multiplied by the conversion gain from "angular velocity" to "image plane blur pixels" to match the unit system with the motion vector output 600 (vector conversion 604). In addition, the "pan-tilt angular velocity" is calculated by applying differentiation 605 to the pan-tilt angle 602. Next, it is multiplied by the conversion gain from "pan-tilt angular velocity" to "image plane blur pixels" (vector conversion 606). At this time, the "pan-tilt angular velocity" and gyro angular velocity (gyro output 601) are axially converted to the blur component on the image plane based on the pan-tilt angle, and the "blur angular velocity" on the vertical and horizontal axes on the image plane is calculated. The adder / subtractor 607 subtracts the gyro output 601 after conversion to vectors and the pan-tilt angular velocity after conversion to vectors from each motion vector output result and inputs it to the motion area determination 608. Meanwhile, the "calculated significance" 603 calculated in the "significance calculation area" is also input to the motion area determination 608. This allows for the determination of whether or not there is a "motion subject detection area" using the method described with reference to Figures 5(a) to (d).
[0043] Furthermore, the amount of movement of each motion vector within the motion detection area and its confidence level are calculated from each motion vector within the motion detection area. The range of motion vector values is divided into several intervals from all motion vectors detected within the motion detection area. Then, a frequency distribution process is performed to list the frequency of occurrence of the motion vector values belonging to each interval.
[0044] The "frequency distribution processing" will be explained with reference to Figures 7(a) and (b). Figures 7(a) and (b) are explanatory diagrams of the frequency distribution processing. Figure 7(a) shows an example of detecting a "motion vector" in the "moving subject detection area". In Figure 7(b), the horizontal axis shows "movement amount (pixels)" and the vertical axis shows "frequency". From the histogram where the number of detected "motion vectors" (frequency) is equal to or greater than the threshold of code 701, the interval 702 where the distribution is most concentrated is set as the "moving subject detection area". Based on the average value of the movement amount of the "motion vector" within the set "moving subject detection area" (interval 702), the "representative motion vector amount" of the "moving subject" is calculated. In addition, the confidence level of the movement amount of the "representative motion vector" calculated from the variance of the movement amount of all "motion vectors" within the "moving subject detection area" is also calculated. Here, a large variance value is judged as low "confidence level", and a small variance value is judged as high "confidence level".
[0045] Furthermore, referring to Figure 6, a method for determining the motion region was explained by subtracting the shaking of the imaging device 101 (camera shake) and the movement of the imaging device 101 due to pan-tilt drive (camera motion vector) from the motion vector output 600. However, there is a problem that the output values of the camera shake and camera motion vector during pan-tilt drive differ depending on the distance from the imaging device 101 to the subject.
[0046] Figure 9(a)~( c Figure 9(a) is an explanatory diagram of the operation of the imaging device 101. Figure 9(b) shows the rotation direction of the imaging device 101. Figure 9(b) shows the motion vector calculated by the image processing unit 207 in the imaging device 101. Figure 9(c) shows the vector amount obtained by subtracting the camera motion vector from the motion vector.
[0047] The blur δ occurring on the imaging plane is calculated using the following equation (1), given the parallel runout Y at the principal point position of the imaging optical system, the runout angle θ of the imaging optical system, the focal length f of the imaging optical system, and the magnification β.
[0048] δ = (1 + β)fθ + βY ···(1) As shown in equation (1), the blur δ occurring on the image sensor varies depending on the focal length f and the magnification β. The focal length can be calculated from the information of the imaging optical system, but the magnification β differs for each subject. When comparing a stationary subject 901 at a distance in the image with a stationary subject 902 at a close distance, the vector quantity 903 in the region of subject 902 is output to be larger than the vector quantity 904 in the region of subject 901. Thus, because the output value of the vector for stationary objects changes depending on the distance between the imaging device 101 and the subject, there is a problem in that a stationary subject at a close distance is mistakenly detected as a moving subject (vector quantity 904).
[0049] Therefore, the detection threshold for a moving subject is varied based on the result of the shaking state determination in step S403 in Figure 4, or the pan-tilt angular velocity (output of derivative 605 in Figure 6). If the number of vectors whose magnitude of movement is greater than or equal to "threshold 1" is detected to be greater than or equal to "threshold 2", it is determined to be a "moving subject detection area". However, for example, if the result of the shaking state determination is determined to be "high vibration", "threshold 1" and "threshold 2" are set to be smaller, and the moving subject determination is performed. Yasu To make it less likely to detect moving subjects. Also, if the result of the vibration state judgment is determined to be "low vibration", set "threshold 1" and "threshold 2" to be set to be larger to make it less likely to detect moving subjects. Also, if the pan-tilt angular velocity is large, set "threshold 1" and "threshold 2" to be set to be smaller to make it less likely to detect moving subjects. Yasu To make this less likely, and when the pan-tilt angular velocity is small, the "threshold 1" and "threshold 2" are set to be larger to make it less likely for the system to detect moving subjects. This method prevents the problem of misidentifying stationary subjects that are close to the subject as moving subjects.
[0050] Another method involves obtaining distance information (depth information) from multiple image data from different viewpoints using a phase-difference detection method with a pupil-splitting image sensor. From the multiple images obtained in this way, it is possible to generate an image displacement map, a defocus amount map calculated by multiplying the image displacement amount by a predetermined conversion coefficient, and a distance map or distance image obtained by converting the defocus amount into distance information of the subject. Alternatively, a method can be used to calculate β at each image position from the distance map and calculate the camera motion vector.
[0051] However, calculation errors can sometimes make it difficult to accurately calculate β. In such cases, after subtracting the camera motion vector calculated considering β from the distance map, the system further varies "threshold 1" and "threshold 2" based on the shaking state determination result and pan-tilt speed to prevent malfunctions in motion subject detection.
[0052] <Method for calculating tracking evaluation value for moving subjects> Next, in step S405 of Figure 4, the control unit 223 calculates an evaluation value (evaluation value information) for tracking control for each moving subject detected in step S404. Then, in step S406, the control unit 223 calculates the center of gravity position of each moving subject based on the image position information of each moving subject detected in step S404 and the evaluation value information of each moving subject calculated in step S405. The control unit 223 then controls the pan-tilt drive unit to transition the center of gravity position to the tracking target position (for example, the center of the image). 205 Calculate the tracking amount to instruct the system.
[0053] Next, in step S407, the control unit 223, based on the tracking amount detected in step S406, controls the pan-tilt drive unit 205 The system is driven to perform tracking control of the moving subject. Subsequently, in step S408, the control unit 223 terminates this process and enters a wait state, waiting for this process to be executed in the next imaging cycle.
[0054] Here, with reference to Figures 8(a)-(d) and 11(a)-(c), the calculation of evaluation values for each moving subject, the calculation of the tracking amount, and the tracking control in steps S405, S406, and S407 of the tracking operation process will be described in detail. Figures 8(a)-8(d) are explanatory diagrams of the tracking operation process. Figures 11(a)-(c) are examples of weight calculation tables.
[0055] The following information is available for each of the detected moving subjects 801 to 808.
[0056] (1) Subject size: The size of a moving subject is calculated using the method described with reference to Figures 5(a) to (d), and this is included as additional information for each moving subject. A weight calculation table is used, as shown in Figure 11(a), in which the weight coefficient is largest when the subject size is at threshold Th1, becomes smaller when it is smaller than threshold Th1, and becomes smaller when it is greater than or equal to threshold Th1.
[0057] (2) Magnitude of the subject's velocity: The magnitude of the velocity of the moving subject is calculated using the method described with reference to Figures 7(a) and (b), and is recorded as additional information for each moving subject. The magnitude of the subject's velocity is the threshold Th 1 The system has a weight calculation table as shown in Figure 11(a), where the weight coefficient is largest when the threshold Th1 is greater than or equal to Th1, becomes smaller when the threshold is less than or equal to Th1, and becomes smaller when the threshold is greater than or equal to Th1.
[0058] (3) Direction of movement of the subject: The direction of movement of the moving subject is calculated using the method described with reference to Figures 7(a) and (b), and this is stored as additional information for each moving subject. It is also calculated whether the subject is moving away from the center of the image or towards the center of the image. The weight coefficient is calculated so that it is larger when the subject is moving away from the center of the image, and smaller when the subject is moving towards the center of the image.
[0059] (4) Number of detected vectors: The number of effective vectors (number of vectors detected) within the detected moving subject is calculated using the method described with reference to Figures 7(a) and (b), and this is stored as additional information for each moving subject. Alternatively, the ratio of effective vectors to the total number of vectors within the moving subject may be used. A weight calculation table is used, as shown in Figure 11(b), where the weight coefficient increases as the number of vectors increases.
[0060] (5) Variance of detected vectors: The variance of vectors within the detected moving subjects is calculated using the method described with reference to Figures 7(a) and (b), and this is stored as additional information for each moving subject. A weight calculation table is used, as shown in Figure 11(c), where the weight coefficient decreases as the variance of the vector increases.
[0061] (6) Detection information of specific subjects around moving subjects: If a specific subject is detected within or around a detected moving subject, this information is added to each moving subject. A weighting coefficient is calculated so that the weight is increased when a specific subject is detected and decreased when a specific subject is not detected.
[0062] From the weight coefficients calculated in (1) to (6) above, the final weight coefficient for each moving subject is calculated. This can be done by adding up the weight coefficients from (1) to (6), or by multiplying each coefficient by another for each level of importance.
[0063] In Figure 8(a), we will explain using the case where moving subjects 801-808 are detected as an example. Moving subjects 802 and 803 are large and moving away from the center of the image, so large weight coefficients are calculated. If the number of effective vectors is large, the vector variance is small, or a specific subject (person) is detected, the weight coefficients will be set to be even larger. Moving subjects 801 and 804 are large, and even if their weight coefficients are determined to be large in other cases, they are moving towards the center of the image, so their weight coefficients will be set to be small. Moving subject 808 is also moving towards the center of the image, so its weight coefficient will be set to be small. Moving subjects 805, 806, and 807 are small and have fewer vectors, so their weight coefficients will be set to be small.
[0064] The center of gravity position 809(H) of each moving subject is calculated from the weight coefficient x and the subject position y of each moving subject.
[0065]
number
[0066] Pan-tilt drive unit for transitioning the center of gravity to the tracking target position (e.g., the center of the image) 205 The tracking amount to instruct the pan-tilt drive unit is calculated. However, if the tracking amount is too large, the angle of view changes due to the abrupt changes in the image will become noticeable, and control oscillations may occur due to the effects of delays in image detection and mechanical drive. Therefore, a tracking amount is calculated that gradually shifts the center of gravity towards the center of the image. Based on the calculated tracking amount, the pan-tilt drive unit 205 It drives the motor and performs tracking control of moving subjects.
[0067] In Figure 8(b), which shows the next frame after Figure 8(a), tracking control is also performed based on the calculation of the center of gravity. Since the moving subject 808 is moving towards the center of the image, its weight is set to be small.
[0068] Furthermore, in Figures 8(c) and (d) showing the next frame, the moving subject 808 is in the image. centerAlthough it is moving away from the image, the movement is in a constant direction from Figure 8(a) to Figures 8(c) and (d). In the case of objects such as cars passing in the background, they almost always move in a constant direction. Therefore, in order to determine that the moving subject 808 is a background object, if the moving subject moving from the edge of the image towards the center of the image as shown in Figure 8(a) continues to move in a constant direction until it reaches the center of the image, the moving subject 808 will be determined to be a background object. The detection of the moving subject 808 is performed by determining the direction and speed of the movement of the moving subject 808 in the next frame. 808 The system predicts the movement of the subject and determines that any moving subject detected within the predicted area is moving subject 808, thereby determining if they are the same moving subject. If, upon transitioning to the center of the image, the system determines that it is a background moving object, it sets a smaller weight for moving subject 808 to prevent the tracking from being pulled along by the movement of moving subject 808.
[0069] Using the method described above, a weighting coefficient corresponding to the evaluation value of each detected moving subject is calculated. Then, by performing moving subject tracking based on the weighting coefficient and the center of gravity position of the moving subject, it becomes possible to perform automatic tracking control focusing on subjects that are within the field of view, even among multiple moving subjects, including falsely detected moving subjects.
[0070] In this embodiment, a method for calculating the tracking amount by determining the center of gravity of a moving subject using the position of the moving subject and a weighting coefficient has been described, but the method is not limited to this. Alternatively, a method may be used in which a target tracking amount z is calculated for each moving subject based on the position of the moving subject, and the final tracking amount C of the moving subject is calculated from the weighting coefficient x calculated in the same manner as above and the target tracking amount z.
[0071]
number
[0072] Furthermore, the pan-tilt tracking control amount is varied by multiplying the tracking control amount by a coefficient ks based on the result of the shaking state determination in step S403, or the pan-tilt angular velocity (output of derivative 605). If the shaking state determination result is determined to be "high vibration", the coefficient ks is set to be small, and the pan-tilt tracking control amount is reduced. If the shaking state determination result is determined to be "low vibration", the coefficient ks is set to be large, and the pan-tilt tracking control amount is increased. Also, if the pan-tilt angular velocity is large, the coefficient ks is set to be small, and the pan-tilt tracking control amount is reduced. Also, if the pan-tilt angular velocity is small, the coefficient ks is set to be large, and the pan-tilt tracking control amount is increased. In this way, if a stationary subject that is close to the subject is mistakenly detected as a moving subject, the angle of view fluctuation due to misdetection and mistracking of moving objects can be reduced by reducing the tracking control amount.
[0073] As described above, in this embodiment, the imaging device 101 includes an imaging unit 206, a motion vector detection unit (image processing unit 207), a subject detection unit 225, a drive unit (pan-tilt drive unit 205), and a control unit 223. The motion vector detection unit detects motion vectors based on image data output from the imaging unit. The subject detection unit detects multiple moving subjects based on the motion vectors. The drive unit drives the imaging unit, and the control unit controls the drive unit to perform tracking control (subject tracking control). The control unit calculates evaluation values for each of the multiple moving subjects based on at least one of the following: information about the multiple moving subjects, information about the shaking (vibration) of the imaging device, or information about the drive state of the drive unit, and controls the drive unit based on the evaluation values.
[0074] Preferably, the evaluation value is a weighting coefficient for the multiple moving subjects. The control unit calculates the center of gravity position of the multiple moving subjects based on the positions of the multiple moving subjects and the weighting coefficients, and controls the drive unit based on the center of gravity position. Also preferably, the subject detection unit calculates the shaking vector of the imaging unit based on information about the drive state of the drive unit and information about the shaking of the imaging device, and detects the multiple moving subjects based on the vector obtained by subtracting the shaking vector from the motion vector. Also preferably, the control unit calculates the evaluation value based on the size, speed, direction of movement of the multiple moving subjects, the number of detected vectors, the variance value of the detected vectors, or at least one of the detection information of specific subjects around the multiple moving subjects.
[0075] Preferably, the control unit sets the evaluation value to the first evaluation value when the amount of shaking of the imaging device is the first amount of shaking, and sets the evaluation value to the second evaluation value which is smaller than the first evaluation value when the amount of shaking is the second amount of shaking which is greater than the first amount of shaking. Also preferably, the subject detection unit sets the detection conditions for each of the multiple moving subjects to the first condition when the amount of shaking of the imaging device is the first amount of shaking, and sets the detection conditions to the second condition which is stricter than the first condition when the amount of shaking is the second amount of shaking which is greater than the first amount of shaking. Also preferably, the control unit sets the tracking amount in tracking control to the first tracking amount when the amount of shaking of the imaging device is the first amount of shaking, and sets the tracking amount to the second tracking amount which is smaller than the first tracking amount when the amount of shaking is the second amount of shaking which is greater than the first amount of shaking.
[0076] Preferably, the control unit sets the evaluation value to a first evaluation value when the speed of the drive unit is a first speed, and sets the evaluation value to a second evaluation value smaller than the first evaluation value when the speed is a second speed faster than the first speed. Also preferably, the subject detection unit sets the detection conditions for each of the multiple moving subjects to a first condition when the speed of the drive unit is a first speed, and sets the detection conditions to a second condition that is stricter than the first condition when the speed is a second speed faster than the first speed. Also preferably, the control unit sets the tracking amount in tracking control to a first tracking amount when the speed of the drive unit is a first speed, and sets the tracking amount to a second tracking amount smaller than the first tracking amount when the speed is a second speed faster than the first speed. Also preferably, the control unit performs tracking control during video recording.
[0077] According to this embodiment, it is possible to provide an imaging device that can be controlled in a manner suitable for imaging multiple moving subjects.
[0078] (Second Embodiment) Next, a second embodiment of the present invention will be described. In the first embodiment, a method for pan-tilt tracking control for multiple moving subjects was described. On the other hand, in this embodiment, a method for solving the problem that occurs when the shooting angle of view turns in an unintended direction due to false detection or false tracking of a moving subject, and subsequently the subject to be photographed no longer appears in the angle of view will be described.
[0079] <Imaging Processing> Referring to Figure 10, the "imaging process" and "tracking operation process" will be explained. Figure 10 is a flowchart of the imaging process in this embodiment.
[0080] First, in step S1001, the control unit 223 determines whether video recording is in progress. If video recording is not in progress (NO), the process proceeds to step S1009, where the control unit 223 determines whether a video recording instruction has been given. If it is determined in step S1009 that no video recording instruction has been given (NO), the process proceeds to step S1012, where the control unit 223 terminates this process and enters a wait state, waiting for this process to be executed in the next imaging cycle. On the other hand, if it is determined in step S1009 that a video recording instruction has been given (YES), the process proceeds to step S1010, where the control unit 223 starts video recording and proceeds to step S1011, where it performs tracking operation processing in the manner described with reference to Figure 4. After the tracking operation processing, the process proceeds to step S1012, where the control unit 223 terminates this process and enters a wait state, waiting for this process to be executed in the next imaging cycle.
[0081] If it is determined in step S1001 that video recording is in progress (YES), the process proceeds to step S1002, where the control unit 223 determines whether or not a video stop instruction has been given. If a video stop instruction has been given (YES), the process proceeds to step S1003, where the control unit 223 stops video recording. Then, in step S1012, the control unit 223 terminates this process and enters a wait state, waiting for this process to be executed in the next imaging cycle.
[0082] Here, the movement in step S1009 picture This section describes how to issue shooting instructions and how to issue video stop instructions in step S1002. Video may be started or stopped by user operations such as pressing the shutter button on the imaging device 101, lightly tapping the imaging device 101 with a finger, voice command input, or instructions from an external device 301. Alternatively, video may be started or stopped by an automatic shooting detection process that automatically determines the timing for starting or stopping video recording.
[0083] The automatic shooting determination function determines whether or not to automatically shoot based on the detected subject. For example, if a specific person is detected and their facial expression and pose meet predetermined conditions, video recording may begin, and the video may stop when the specific person is no longer present. Alternatively, video recording may begin when a moving subject is detected, and stop if no moving subjects are detected for an extended period.
[0084] If no instruction to stop video recording has been given in step S1002 and video recording is continuing (NO), the process proceeds to step S1004. In step S1004, the control unit 223 determines whether the counter, which is provided to return the pan-tilt position (the position of the pan-tilt unit, i.e., the position of the imaging unit 206) to the reference position, is greater than or equal to the threshold ThC. If the counter is greater than or equal to the threshold ThC (YES), the process proceeds to step S1005. In step S1005, the control unit 223 moves the pan-tilt position to the reference position. Subsequently, in step S1012, the control unit 223 terminates this process and enters a wait state, waiting for this process to be executed in the next imaging cycle. On the other hand, if the counter in step S1004 is less than the threshold ThC, the process proceeds to step S1006.
[0085] In step S1006, the control unit 223 determines whether the count clearing condition is met. If the count clearing condition is met (YES), the process proceeds to step S1007. In step S1007, the control unit 223 clears the counter to 0. Subsequently, in step S1011, the control unit 223 performs the tracking operation. On the other hand, if the count clearing condition is not met in step S1006 (NO), the process proceeds to step S1008. In step S1008, the control unit 223 increments the counter. Subsequently, in step S1011, the control unit 223 performs the tracking operation. After performing the tracking operation in step S1011, the process proceeds to step S1012, where the control unit 223 terminates this process and enters a wait state, waiting for this process to be executed in the next imaging cycle.
[0086] Next, we will explain how to determine the count clearing condition in step S1006, and how to count up in step S1008.
[0087] <How to set the reference position> Before the video starts, a reference position for the pan-tilt angle (pan-tilt position) is set in advance. Any of the methods described below can be used to set the reference position.
[0088] The first method described is setting the pan-tilt position at the start of video recording as the reference position. In step S1009, when video recording is started manually by the user, the user can confirm whether the lens barrel 102 of the imaging device 101 is facing the direction they want to shoot before issuing the recording command. Therefore, there is a high probability that the subject being filmed will be at the pan-tilt position at the start of the video. Also, even when a video start command is issued by the automatic shooting process, automatic video recording is performed when the subject to be filmed is detected, so there is a high probability that the subject being filmed will be at the pan-tilt position at the start of the video. Therefore, the pan-tilt position at the start of video recording is set as the reference position.
[0089] Another method involves the user setting the reference position. For example, a dedicated application installed on an external device 301 allows the user to view the image from the imaging device 101, move to the pan-tilt position to be set as the reference position, and then set the reference position. Alternatively, the pan-tilt position at which a specific command is detected via voice command can be set as the reference position, or the position at which the user manually rotates the pan-tilt rotation unit and stops it can be set as the reference position. In any of these methods, the reference position can be set by user operation.
[0090] <How to count up> Next, the method for counting up in step S1008 will be explained. The count-up value COUNT is varied according to the time elapsed since the rotational position of the pan-tilt unit (position of the imaging unit 206) deviated from the reference position, the amount of the deviation, the detection information from the motion detection unit, or the detection information of a specific subject, as shown in the following equation (4).
[0091] COUNT=K1×K2×K3×K4×Base ···(4) K1 is a coefficient that varies depending on the elapsed time since the rotational position of the pan-tilt unit deviated from the reference position. The value of K1 is set to increase as the elapsed time since the rotational position of the pan-tilt unit passed the reference position increases. K2 is the pan-tilt Department This is a coefficient that varies depending on the angle at which the rotation position of the pan-tilt unit is moved away from the reference position. The larger the angle at which the rotation position of the pan-tilt unit is moved away from the reference position, the larger the value of K2 is set to be. K3 is a coefficient that varies depending on whether or not a moving subject is detected. The period during which a moving subject is detected is K3Scope The size decreases, and the period during which no moving subject is detected is K3Scope It is set to be large. K4 is a coefficient that varies depending on whether a specific subject is detected or not. The period during which a specific subject is detected is K4 value The size decreases, and the period during which no specific subject is detected is K4 value It is set to be large. Base is a fixed parameter for the minimum count value, and coefficients K1 to K4 are variable values of 1 or greater.
[0092] As described above, the longer the time elapsed since the pan-tilt rotation position passed the reference position, and the greater the angle at which the pan-tilt rotation position deviates from the reference position, the larger the count-up value is set. This makes it easier to return to the reference position. Also, if no moving subject or specific subject is detected, the count-up value is increased to make it easier to return to the reference position.
[0093] In this embodiment, it is not necessary to base the setting on the time the imaging unit 206 is separated from the reference position, the amount of angle of separation, the detection information of a moving subject, and the detection information of a specific subject (it is not necessary for all coefficients K1 to K4 to be variable). For example, the count-up value COUNT can be set based on at least two of these factors (it is sufficient to make at least two of the coefficients K1 to K4 variable).
[0094] <Count Clear Condition Check> Next, we will explain how to clear the count in step S1007. Department The count is reset to 0 when the rotation position passes the reference position. Alternatively, the count may be reset to 0 when a specific subject is detected, when a moving subject that meets specific conditions is detected, or based on the conditions of a specific subject and a moving subject. For example, the count may be reset when a specific subject that was detected at the start of automatic video recording is detected, or the count may be reset when a specific subject that has been set in advance by the user via a dedicated application installed on an external device 301 is detected. Furthermore, the count may be reset when multiple moving subjects are detected, or when a specific subject is detected and a moving subject is detected near the specific subject.
[0095] As described above, in this embodiment, the subject detection unit 225 detects a moving subject based on a motion vector or a specific subject, and the control unit 223 sets the reference position of the imaging unit 206. The control unit also determines whether or not to move the imaging unit to the reference position based on at least two of the time the imaging unit is away from the reference position, the amount of angle of departure, the detection information of the moving subject, or the detection information of the specific subject. Preferably, the reference position is set by user instruction. Preferably, the reference position is the position at which video recording is started.
[0096] According to this embodiment, if the shooting angle of view moves significantly away from the reference position due to false detection or false tracking of a moving subject, and the subject to be photographed subsequently disappears from the angle of view, the pan-tilt function can be activated under predetermined conditions. positionThe device is returned to its reference position. This increases the probability that the subject to be photographed can be detected again after returning to the reference position. Although this embodiment describes an example in which a "moving subject" is detected based on the captured image, it can also be applied to cases where a mechanism is used to detect movement around the imaging device 101 using infrared light, ultrasound, visible light, etc.
[0097] (Other embodiments) The present invention can also be realized by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit (e.g., an ASIC) that implements one or more functions.
[0098] According to each embodiment, it is possible to provide an imaging device, a control method for the imaging device, and a program that are capable of being controlled in a manner suitable for imaging multiple moving subjects.
[0099] Although preferred embodiments of the present invention have been described above, the present invention is not limited to these embodiments, and various modifications and changes are possible within the scope of its gist. [Explanation of symbols]
[0100] 101 Imaging device 205 Pan-tilt drive unit (drive unit) 206 Imaging Department 207 Image Processing Unit (Motion Vector Detection Unit) 223 Control Unit 225 Subject detection unit
Claims
1. Imaging unit, A motion vector detection unit detects motion vectors based on image data output from the imaging unit, A subject detection unit that detects multiple moving subjects based on the aforementioned motion vector, A drive unit that drives the imaging unit to pan and tilt, It has a control unit that performs tracking control by controlling the aforementioned drive unit, The control unit, Based on the respective positions of the multiple moving subjects in the image data and the respective weight coefficients of the multiple moving subjects, the centroid position of the multiple moving subjects in the image data is calculated. An imaging device characterized by controlling the drive unit based on the position of the center of gravity.
2. The imaging apparatus according to claim 1, characterized in that the control unit calculates the weight coefficient based on at least one of the size, speed, direction of movement, number of detected vectors, variance of the detected vectors, or detection information of specific subjects in the vicinity of the plurality of moving subjects.
3. The subject detection unit is If the amount of shaking of the imaging device is the first amount of shaking, the determination threshold for each of the multiple moving subjects is set to the first threshold. The imaging apparatus according to claim 1 or 2, characterized in that, when the amount of shaking is a second amount of shaking which is greater than the first amount of shaking, the determination threshold is set to a second threshold which is greater than the first threshold.
4. The control unit, If the amount of shaking of the imaging device is the first amount of shaking, the tracking amount in the tracking control is set to the first tracking amount. The imaging device according to any one of claims 1 to 3, characterized in that, when the amount of shaking is a second amount of shaking which is greater than the first amount of shaking, the tracking amount is set to a second tracking amount which is smaller than the first tracking amount.
5. The subject detection unit is When the speed of the drive unit is a first speed, the determination threshold for each of the plurality of moving subjects is set to the first threshold. The imaging apparatus according to any one of claims 1 to 4, characterized in that, when the speed is a second speed that is faster than the first speed, the determination threshold is set to a second threshold that is greater than the first threshold.
6. The control unit, When the speed of the drive unit is the first speed, the tracking amount in the tracking control is set to the first tracking amount. The imaging apparatus according to any one of claims 1 to 5, characterized in that, when the speed is a second speed that is faster than the first speed, the tracking amount is set to a second tracking amount that is smaller than the first tracking amount.
7. The imaging device according to any one of claims 1 to 6, characterized in that the control unit performs the tracking control during video recording.
8. The imaging unit further comprises a pan-tilt unit that rotates the imaging unit in the pan and tilt directions. The imaging apparatus according to any one of claims 1 to 7, characterized in that the drive unit drives the pan-tilt unit.
9. A control method for an imaging device having an imaging unit, a motion vector detection unit that detects motion vectors based on image data output from the imaging unit, a subject detection unit that detects multiple moving subjects based on the motion vectors, a drive unit that drives the imaging unit to pan and tilt, and a control unit that performs tracking control by controlling the drive unit, A calculation step of calculating the centroid position of the multiple moving subjects in the image data based on the respective positions of the multiple moving subjects in the image data and the respective weight coefficients of the multiple moving subjects, A control method for an imaging device, characterized by comprising: a control step of performing tracking control by controlling the drive unit based on the center of gravity position calculated in the calculation step.
10. A program characterized by causing a computer to execute the control method described in claim 9.
Citation Information
Patent Citations
Imaging apparatus
JP2008177972A
Image processing apparatus, image processing method and program
JP2011139231A
Imaging apparatus
JP2012124575A
Control apparatus, optical instrument, imaging apparatus, and control method
JP2016208230A
Imaging device and control method of the same
JP2018205551A