Image processing apparatus and method, and computer-readable storage medium

By detecting the vibration of the shooting device and the motion vector of the subject, and processing the clipping positions of the subject and background respectively, the problem of background jitter in the subject tracking function is solved, and the stability and clarity of the captured image are achieved.

CN116113977BActive Publication Date: 2025-12-02JVC KENWOOD CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202180058607.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-09-17
Filing Date
2021-06-01
Publication Date
2025-12-02
Estimated Expiration
2041-06-01

AI Technical Summary

Technical Problem

In cameras equipped with subject tracking, the background is prone to shaking due to camera shake, and existing technologies struggle to suppress overall image shake.

Method used

By detecting the vibration of the shooting device and calculating the amount of movement, and combining it with the motion vector of the subject, the cropping positions of the subject and the background are determined separately, and the images are composited to eliminate shakiness.

Benefits of technology

It effectively suppresses overall shaking in the captured images, ensuring that the subject is clear and the background is natural, resulting in stable images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116113977B_ABST
    Figure CN116113977B_ABST
Patent Text Reader

Abstract

The object detection unit (23) detects objects of interest within frames of an image acquired from the imaging unit (10). The motion vector detection unit (25) detects motion vectors of objects between frames of the image. The cropping position determination unit (27) determines the cropping position for cropping an image of a predetermined size from each frame of the acquired image. The cropping position determination unit (27) determines a cropping position for the object that moves according to a composite vector and a cropping position for the background that moves according to a correction vector, wherein the composite vector is a vector synthesized by combining the correction vector and the motion vector of the object, and the correction vector is used to eliminate jitter caused by vibrations applied to the imaging unit (10). The image compositing unit (31) synthesizes the image data of the object cropped from the cropping position for the object and the image data of the background other than the object cropped from the cropping position for the background.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image processing apparatus and method for processing images captured by a camera, as well as a computer-readable storage medium. Background Technology

[0002] Cameras equipped with electronic image stabilization are becoming increasingly common. Electronic image stabilization adaptively shifts the field of view cropped from the shooting range in a way that counteracts camera shake, resulting in a shake-reduced image.

[0003] In addition, cameras equipped with subject tracking functions are becoming increasingly common. This subject tracking function tracks subjects such as people. In subject tracking, as the subject moves, the cropping range of the field of view shifts according to the subject's movement, thereby controlling the cropping to keep the subject's position within the field of view as still as possible. For example, the field of view can be cropped in a way that keeps the subject fixed in the center within the field of view.

[0004] Patent Document 1 discloses a technique for synthesizing natural candid or wide-angle images from multiple images acquired through panning photography. Specifically, the area near the camera's movement is used as the background area. In each of the acquired images, the main subject area is extracted from the background area, ensuring the positions of the main subject areas in the multiple images are matched. The background area is then filtered to generate the synthesized image. However, this technique is not one that synthesizes the extracted subject and background to generate an image with a typical field of view; rather, it is a technique used in special photography such as panoramic photography.

[0005] Existing technical documents

[0006] Patent documents

[0007] Patent document 1: Japanese Patent Application Publication No. 2017-098776. Summary of the Invention

[0008] When the subject tracking function described above moves the field of view clipping range according to the movement of the subject, it can suppress subject jitter, but jitter will occur in the background.

[0009] This embodiment was made in view of the following situation, and its purpose is to provide a technique for suppressing shaking in the captured image as a whole.

[0010] To address the aforementioned issues, an image processing apparatus according to one embodiment includes: an image acquisition unit that acquires images captured by an imaging unit; a motion determination unit that determines the motion of the imaging unit based on the vibration applied to the imaging unit, based on the output signal of a sensor used to detect vibrations applied to the imaging unit or the inter-frame difference of the acquired images; an object detection unit that detects objects of interest within frames of the acquired images; a motion vector detection unit that detects motion vectors of the objects between frames of the images; a cropping position determination unit that determines a cropping position for cropping images of a predetermined size from each frame of the acquired images, and determines a cropping position for the objects that moves according to a composite vector and a cropping position for the background that moves according to a correction vector, the composite vector being a vector synthesized by combining the correction vector and the motion vector of the objects, the correction vector being used to eliminate jitter caused by the vibration; and an image synthesis unit that synthesizes image data of the objects cropped from the cropping position for the objects and image data of the background other than the objects cropped from the cropping position for the background.

[0011] Furthermore, any combination of the above-mentioned constituent elements, and any variations of the embodiments of this implementation in the form of apparatus, method, system, recording medium, computer program, etc., are also valid as embodiments of this implementation.

[0012] According to this embodiment, shaking within the captured image can be suppressed overall. Attached Figure Description

[0013] Figure 1 This is a diagram showing the structure of the imaging device according to Embodiment 1 of the present invention.

[0014] Figure 2 Figures (a)-(c) are diagrams (one of which) illustrating a specific example of the synthetic image generation process of the image processing apparatus according to Embodiment 1.

[0015] Figure 3 Figures (a)-(c) are diagrams (second example) illustrating a specific example of the synthetic image generation process of the image processing apparatus according to Embodiment 1.

[0016] Figure 4 This is a diagram illustrating the structure of the imaging device according to Embodiment 2 of the present invention.

[0017] Figure 5 Figures (a)-(c) are diagrams (one of which) illustrating a specific example of the synthetic image generation process of the image processing apparatus according to Embodiment 2.

[0018] Figure 6Figures (a)-(c) are diagrams (second example) illustrating a specific example of the synthetic image generation process of the image processing apparatus according to Embodiment 2.

[0019] Figure 7 Figures (a)-(c) are diagrams (third example) illustrating a specific example of the synthetic image generation process of the image processing apparatus according to Embodiment 2.

[0020] Figure 8 This is a diagram illustrating the structure of the imaging device according to Embodiment 3 of the present invention.

[0021] Figure 9 Figures (a)-(c) are specific examples illustrating the limitation processing of the total cropping range of the image processing apparatus according to embodiments 1 and 3. Detailed Implementation

[0022] Figure 1 This diagram illustrates the structure of the shooting device 1 according to Embodiment 1 of the present invention. The shooting device 1 may be a single camera or a camera module mounted on an information device such as a smartphone, tablet computer, or laptop computer.

[0023] The imaging apparatus 1 according to Embodiment 1 includes an imaging unit 10, a vibration detection sensor 11, and an image processing device 20. The imaging unit 10 includes a lens, a solid-state imaging element, and a signal processing circuit. The solid-state imaging element can be, for example, a CMOS (Complementary Metal Oxide Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor. The solid-state imaging element converts light incident through the lens into an electrical image signal and outputs it to the signal processing circuit. The signal processing circuit performs signal processing such as A / D conversion and noise removal on the image signal input from the solid-state imaging element and outputs it to the image processing device 20.

[0024] The vibration detection sensor 11 detects the vibration applied to the imaging unit 10 and outputs it to the image processing device 20. The vibration detection sensor 11 can be, for example, a gyroscope sensor. The gyroscope sensor detects the vibration applied in the yaw and pitch directions of the imaging unit 10 as angular velocities, respectively.

[0025] The image processing apparatus 20 includes an image acquisition unit 21, an image recognition unit 22, an object detection unit 23, an object tracking unit 24, a motion vector detection unit 25, a cut position determination unit 27, a vibration information acquisition unit 28, a movement amount determination unit 29, a cut unit 30, an image synthesis unit 31, and a pixel supplementation unit 32. These components can be implemented through the cooperation of hardware and software resources, or solely through hardware resources. Hardware resources can utilize CPUs, ROMs, RAMs, GPUs (Graphics Processing Units), DSPs (Digital Signal Processors), ISPs (Image Signal Processors), ASICs (Application Specific Integrated Circuits), FPGAs (Field-Programmable Gate Arrays), and other LSIs. Software resources can utilize firmware and other programs.

[0026] The image acquisition unit 21 acquires the image captured by the imaging unit 10. The vibration information acquisition unit 28 acquires the output signal of the vibration detection sensor 11 as vibration component information. The movement amount determination unit 29 integrates the output signal acquired by the vibration information acquisition unit 28 to determine the movement amount of the imaging unit 10. For example, the movement amount determination unit 29 integrates the angular velocity signals in the yaw and pitch directions acquired by the vibration information acquisition unit 28 respectively, and calculates the movement angle of the imaging unit 10 in the yaw and pitch directions.

[0027] Furthermore, the movement determination unit 29 can also determine the overall movement of the background area as the movement of the imaging unit 10 when, based on the inter-frame difference detection, the entire background area, except for the object identified by the image recognition unit 22 described later, moves in the same direction. In this case, even if the vibration detection sensor 11 is omitted, the movement of the imaging unit 10 based on the vibration applied to it can be estimated. Omitting the vibration detection sensor 11 can reduce costs.

[0028] The movement amount determination unit 29 generates a correction vector to compensate for the determined movement amount of the shooting unit 10. That is, a correction vector with the same correction amount as the movement amount of the shooting unit 10 is generated in the direction of the shaking of the shooting unit 10. The movement amount determination unit 29 outputs the generated correction vector to the clipping position determination unit 27. When the shooting unit 10 is ideally stationary, the value of the correction vector becomes zero.

[0029] The image recognition unit 22 searches for objects within frames of images acquired by the image acquisition unit 21. The image recognition unit 22 has a dictionary of objects, generated by learning from multiple images of specific objects. Specific objects include, for example, a person's face, a person's whole body, an animal's face (e.g., a dog's or cat's), an animal's whole body, and vehicles (e.g., railway vehicles).

[0030] The image recognition unit 22 uses object recognizers to search within the frame of the image. Object recognition can use HOG (Histogram of Oriented Gradients) features, for example. Alternatively, Haar-like features, LBP (Local Binary Patterns) features, etc., can also be used. When an object is present within the frame, the image recognition unit 22 fills in the object with a rectangular detection box.

[0031] The object detection unit 23 decides whether to detect the object identified by the image recognition unit 22 as an object of interest. An object of interest is an object that is estimated to be a subject of interest to the user (photographer) of the shooting device 1.

[0032] The object detection unit 23 determines whether to include the identified object as a subject based on any combination of one or more of the following criteria: (a) whether the object is larger than a specified size, (b) whether the object is located in the central region of the frame, (c) whether the object is a person or an animal, (d) whether the distance to the object is less than a specified value (the method for estimating the distance to the object will be described later), (e) whether there is any part hidden within the object, and (f) whether the motion of the object is less than a specified value (the method for motion detection will be described later).

[0033] When multiple objects are identified within a frame, the object detection unit 23 can detect all objects that meet the above-mentioned judgment criteria as subjects, or it can select a subject based on any combination of one or more of the following judgment criteria: (a) whether it is the largest object in the frame, (b) whether it is the object located in the center of the frame, (d) whether it is the closest to the object, and (f) whether it is the object with the least movement among the moving objects.

[0034] These criteria are based on the following rules of thumb: the majority of the subject being photographed occupies a large area within the frame, is located in the center of the frame, is located near the front of the depth of field, and has minimal motion within the frame due to the photographer tracking the subject using the shooting device 1.

[0035] Furthermore, when the viewfinder of the shooting device 1 is a touch panel, the object detection unit 23 can also determine the object that is touched by the user of the shooting device 1 among the objects displayed in the viewfinder as the subject.

[0036] The object tracking unit 24 tracks the objects identified by the image recognition unit 22 in subsequent frames. For object tracking, for example, a particle filter or mean shift method can be used. The tracked objects can be all objects identified by the image recognition unit 22, or only those objects detected as objects of interest by the object detection unit 23. Furthermore, if the motion of objects is used as a reference for selecting the subject, it is necessary to track all objects identified by the image recognition unit 22.

[0037] The motion vector detection unit 25 detects the amount of movement of the object of interest between frames of the image as the motion vector of that object. This motion vector represents the forward vector (tracking vector) of the object's movement.

[0038] The cropping position determination unit 27 determines the cropping position for cropping an image of a predetermined size from each frame of the image. The cropping position determination unit 27 crops a portion of the image from all the shooting range captured by all pixels of the solid-state imaging element to determine the range of the image to be displayed or recorded.

[0039] In this embodiment, an electronic shake correction function is employed. In this electronic shake correction function, the position of the field of view cropped from all shooting ranges adaptively changes to counteract hand shakes in the shooting device 1. Additionally, in this embodiment, a subject tracking function is also employed. In this subject tracking function, the position of the field of view cropped from all shooting ranges changes in a way that keeps the position of the subject within the field of view relative to the subject's movement as fixed as possible. As described above, in this embodiment, real-time field of view cropping is performed based on the electronic shake correction function and the subject tracking function.

[0040] The cut position determination unit 27 obtains a correction vector from the movement amount determination unit 29 and obtains the position and motion vector of the object of interest from the motion vector detection unit 25. The cut position determination unit 27 determines the cut position of the object of interest and the cut position of the background, respectively.

[0041] The cut position determination unit 27 moves the cut position of the reference frame for the object according to the composite vector obtained by combining the correction vector and the motion vector of the object, and determines the cut position of the object for the current frame.

[0042] When the reference frame is the previous frame, the correction vector becomes a correction vector used to eliminate the motion of the camera unit 10 between the previous frame and the current frame, and the motion vector of the object becomes a motion vector representing the motion of the object between the previous frame and the current frame. When the reference frame is the frame at which tracking of the object begins, the correction vector becomes a correction vector used to cancel the motion of the camera unit 10 between the frame at which tracking begins and the current frame, and the motion vector of the object becomes a motion vector representing the motion of the object between the frame at which tracking begins and the current frame.

[0043] The clipping position determination unit 27 moves the clipping position for the background of the reference frame according to the correction vector, and determines the clipping position for the background of the current frame. When the reference frame is the previous frame, the correction vector becomes a correction vector for eliminating the motion of the imaging unit 10 between the previous frame and the current frame. When the reference frame is the frame at the start of tracking the object, the correction vector becomes a correction vector for offsetting the motion of the imaging unit 10 between the frame at the start of tracking and the current frame.

[0044] The cut position determination unit 27 can also determine the cut position for the background by moving the cut position for the object in the current frame according to the inverse vector of the object's motion vector. In this case, the object's motion vector becomes the motion vector representing the motion between the frame at the start of tracking and the current frame.

[0045] The cropping unit 30 crops only the image data of the object from the image data for the cropping position of the object. The cropping unit 30 also crops the image data of the background, excluding the object, from the image data for the cropping position of the background. The image compositing unit 31 combines the cropped image data of the object and the cropped image data of the background.

[0046] Detailed explanations will follow, but when multiple objects of interest are set within a frame, defective pixels may be generated in the synthesized image. In this case, the pixel supplementation unit 32 supplements the defective pixels in the synthesized image based on at least one valid pixel that is spatially or temporally close to the defective pixel.

[0047] Figure 2 Figures (a)-(c) are diagrams (one of which) illustrating a specific example of the synthetic image generation process of the image processing apparatus 20 according to Embodiment 1. Figure 3 Figures (a)-(c) are diagrams (second example) illustrating a specific example of the synthetic image generation process of the image processing apparatus 20 according to Embodiment 1.

[0048] like Figure 2As shown in (a), the cut position determination unit 27 sets the cut position C0 at the center of frame F0 by default. The object detection unit 23 detects object OB1 as an object of interest within frame F0. The object tracking unit 24 begins tracking object OB1. Figure 2 In the example shown in (a), the object OB1 moves to the left. Additionally, the camera unit 10 moves upwards due to the photographer's hand tremor.

[0049] like Figure 2 As shown in (b), the cut position determination unit 27 moves the cut position C0 of the reference frame F0 in the current frame F1 according to the composite vector of the hand shakiness correction vector and the motion vector of the object OB1, thereby determining the cut position C1 for the object OB1.

[0050] like Figure 2 As shown in (c), the cut position determination unit 27 moves the cut position C1 of the object OB1 in the current frame F1 according to the inverse vector of the motion vector of the object OB1, thereby determining the cut position C2 for the background.

[0051] like Figure 3 As shown in (a), the clipping unit 30 clips the image data of object OB1 from the image data of clipping position C1 used for object OB1 in the current frame F1. Figure 3 As shown in (b), the clipping unit 30 clips the background image data, excluding the image data of the object OB1, from the image data of the current frame F1 at the clipping position C2 for the background. Figure 3 As shown in (c), the image synthesis unit 31 synthesizes the image data of the cut object OB1 and the image data of the cut background to generate a new synthesized image Ic.

[0052] As explained above, according to Embodiment 1, when cropping the field of view while tracking the subject from all shooting ranges, the field of view for the subject and the field of view for the background are cropped separately, and background shake correction is applied to the field of view for the background. After correcting the background shake, the field of view for the subject and the field of view for the background are combined, thereby suppressing background shake. At this time, hand shake correction is applied to both the field of view for the subject and the field of view for the background. As a result, an image in which shake is suppressed across the entire field of view can be generated.

[0053] Therefore, it is possible to capture the subject clearly while generating an image with a natural background. For example, when shooting a cat in a living room, the cat is tracked in the subject tracking function, thus suppressing the cat's movement. Furthermore, the background does not shake with the cat's movement, resulting in an image with a natural background. Additionally, when the photographer pans, background movement can also be suppressed, preventing the image from becoming unrecognizable. Moreover, although details will be described later, even with multiple subjects, it is possible to include them within a single field of view as much as possible. Furthermore, by applying the same strength of the correction vector used to compensate for hand shake as the subject's motion vector, it is possible to capture images with a static background, such as virtual backgrounds. Additionally, it is possible to create materials that can be used with virtual backgrounds.

[0054] Figure 4 This diagram illustrates the structure of the imaging device 1 according to Embodiment 2 of the present invention. The imaging device 1 according to Embodiment 2 includes an imaging unit 10, a vibration detection sensor 11, a distance detection sensor 12, and an image processing device 20. The image processing device 20 according to Embodiment 2 includes an image acquisition unit 21, an image recognition unit 22, an object detection unit 23, a vector intensity adjustment unit 26, a cropping position determination unit 27, a vibration information acquisition unit 28, a movement amount determination unit 29, a cropping unit 30, an image compositing unit 31, a pixel supplementation unit 32, a distance information acquisition unit 33, and a distance determination unit 34.

[0055] The differences from Embodiment 1 will be explained below. The distance detection sensor 12 is a sensor used to detect the distance between an object located in the shooting direction and the shooting unit 10. For example, a TOF (Time of Flight) sensor can be used for the distance detection sensor 12. Representative TOF sensors include ultrasonic sensors (sonar) and LiDAR (Light Detection and Ranging). An ultrasonic sensor sends ultrasonic waves in the shooting direction, measures the time until its reflected wave is received, and detects the distance to an object located in the shooting direction. A LiDAR sensor illuminates a laser in the shooting direction, measures the time until its reflected light is received, and detects the distance to an object located in the shooting direction.

[0056] The distance information acquisition unit 33 acquires the output signal of the distance detection sensor 12 as distance information. The distance determination unit 34 generates a distance image corresponding to the visible light image captured by the imaging unit 10 based on the acquired distance information.

[0057] The distance determination unit 34 can also estimate the distance from the shooting unit 10 to the object displayed in the frame through image recognition. For example, the relationship between the general size of the object registered as dictionary data, the size of the object displayed in the frame, and the distance from the shooting unit 10 to the object can be specified in advance using a table or function. The distance determination unit 34 estimates the distance from the shooting unit 10 to the object based on the size of the object identified by the image recognition unit 22 in the frame and by referring to the table or function. In this case, even if the distance detection sensor 12 is omitted, the distance to the object displayed in the frame can be estimated. If the distance detection sensor 12 is omitted, the cost can be reduced.

[0058] Furthermore, the distance determination unit 34 can also use the distance obtained from the autofocus adjustment unit (not shown) if it is possible to obtain the distance from the shooting unit 10 to the object displayed in the frame. Additionally, if the shooting unit 10 is composed of two cameras, the distance determination unit 34 can also estimate the distance to the object displayed in the frame based on the parallax between the images captured by the shooting units 10 of both cameras.

[0059] In Embodiment 2, the subject tracking function is omitted. Therefore, it is basically unnecessary to calculate the motion vector of the object. In Embodiment 2, the object tracking unit 24 and the motion vector detection unit 25 are omitted from the image processing device 20.

[0060] In Embodiment 2, an electronic hand-shake correction function is also employed. That is, the position of the field of view cut off from all shooting ranges is adaptively moved so that the hand-shake of the shooting device 1 is canceled out. In Embodiment 2, a vector intensity adjustment unit 26 is added, which is used to adjust the intensity of the correction vector used to eliminate hand-shake.

[0061] The vector intensity adjustment unit 26 obtains a correction vector from the movement amount determination unit 29 and the distance to the target object from the distance determination unit 34. The vector intensity adjustment unit 26 adjusts the intensity of the correction vector obtained from the movement amount determination unit 29 based on the distance to the target object obtained from the distance determination unit 34. When the distance to the target object is closer, the vector intensity adjustment unit 26 sets the intensity of the correction vector to be stronger; when the distance to the target object is farther, the vector intensity adjustment unit 26 sets the intensity of the correction vector to be weaker.

[0062] For example, the relationship between the amount of hand shakiness, the distance from the shooting unit 10 to the object, and the amount of movement caused by hand shakiness of the object reflected in the frame can be specified in advance using a table or function. This relationship can also be derived based on the designer's experiments or simulations. The vector intensity adjustment unit 26 adjusts the intensity of the correction vector obtained from the movement amount determination unit 29 based on the distance from the shooting unit 10 to the object obtained from the distance determination unit 34 and with reference to the table or function.

[0063] In Embodiment 2, the cut position determination unit 27 obtains the position of the object of interest from the object detection unit 23 and obtains the correction vector of the object from the vector intensity adjustment unit 26. The cut position determination unit 27 moves the cut position of a reference frame (e.g., the previous frame) according to the correction vector to determine the cut position of the current frame. At this time, the cut position determination unit 27 determines the cut position for the object and the cut position for the background respectively based on the intensity of the correction vectors for the object and the background.

[0064] When multiple objects of interest are defined within a frame, the cut position determination unit 27 determines the cut position according to each object of interest. For example, if a first object and a second object located deeper than the first object are detected as objects of interest within the frame, the vector intensity adjustment unit 26 sets the intensity of the correction vector of the first object to be stronger than the correction intensity of the second object.

[0065] Furthermore, regarding the background correction vector, it can be used directly without adjusting its intensity, or its intensity can be adjusted by the vector intensity adjustment unit 26. For example, the overall background can be divided into multiple backgrounds based on distance distinctions, and the vector intensity adjustment unit 26 can adjust the intensity of the correction vector for each of the segmented backgrounds. Alternatively, the vector intensity adjustment unit 26 can also adjust the intensity of the correction vector based on representative values ​​(e.g., average, median, or most frequent values) of the distances to each object constituting the background.

[0066] Figure 5 Figures (a)-(c) are diagrams (one of which) illustrating a specific example of the synthetic image generation process of the image processing apparatus 20 according to Embodiment 2. Figure 6 Figures (a)-(c) are diagrams (second example) illustrating a specific example of the synthetic image generation process of the image processing apparatus 20 according to Embodiment 2. Figure 7 Figures (a)-(b) are diagrams (third example) illustrating a specific example of the synthetic image generation process of the image processing apparatus 20 according to Embodiment 2.

[0067] like Figure 5As shown in (a), the cut position determination unit 27 sets the cut position C0 at the center of frame F0 by default. The object detection unit 23 detects the first object OB1 and the second object OB2 as objects of interest within frame F0. From the perspective of the imaging unit 10, the second object OB2 exists inside the first object OB1, and a portion of the second object OB2 is hidden by the first object OB1 within frame F0.

[0068] Figure 5 (b) shows the camera part 10 falling from the camera due to the photographer's hand tremor. Figure 5 The state shown in (a) is a state that moves to the right. Figure 5 (b) shows the case where the electronic hand shake correction function is off. In this case, the shear position C0 remains unchanged, and the first object OB1 and the second object OB2 move to the left of the field of view.

[0069] At this time, the camera unit 10 moves to the right, and the viewpoint moves to the right, thus changing the way the relative positions of the first object OB1 and the second object OB2 are observed. Specifically, in the current frame F1, the first object OB1, which is closer to the viewpoint, moves significantly to the left than the second object OB2, which is farther from the viewpoint. That is, compared with the reference frame F0 before the hand shake, in the current frame F1, the overlapping portion of the first object OB1 and the second object OB2 becomes larger, and the portion of the second object OB2 that is hidden by the first object OB1 becomes larger.

[0070] Figure 5 (c) shows the case where the electronic shake correction function is enabled, and the clipping position C0 of the reference frame F0 is simply moved to the left according to the shake correction vector. Due to the movement of the clipping position C0, at the clipping position C0' of the current frame F1, the first object OB1 and the second object OB2 are positioned at the center of the field of view. However, compared to the reference frame F0, the composition of the first object OB1 and the second object OB2 within the current frame F1 is altered.

[0071] In contrast, Figure 6 (a)-(c) respectively determine the cut-off positions of the first object OB1, the second object OB2, and the background. For example... Figure 6 As shown in (a), the cut position determination unit 27 moves the cut position C0 of the reference frame F0 according to the correction vector of the first object OB1 to determine the cut position C1 for the first object OB1. Figure 6 As shown in (b), the cut position determination unit 27 moves the cut position C0 of the reference frame F0 according to the correction vector of the second object OB2 to determine the cut position C2 for the second object OB2. Figure 6As shown in (c), the cut position determination unit 27 moves the cut position C0 of the reference frame F0 according to the background correction vector to determine the cut position C3 for the background.

[0072] like Figure 7 As shown in (a), the cropping unit 30 crops image data of the first object OB1 from image data at cropping position C1 used for the first object OB1, crops image data of the second object OB2 from image data at cropping position C2 used for the second object OB2, and crops image data of the background other than the first object OB1 and the second object OB2 from image data at cropping position C3 used for the background. The image compositing unit 31 combines the cropped image data of the first object OB1, the cropped image data of the second object OB2, and the cropped background image data to generate a new composite image Ic.

[0073] The composition of the first object OB1 and the second object OB2 in the newly generated composite image Ic is the same as the composition of the first object OB1 and the second object OB2 in the clipping position C0 of the reference frame F0. That is, it is possible to perform a pseudo-viewpoint transformation as if the imaging unit 10 has not moved.

[0074] The overlapping portions of the first object OB1 and the second object OB2 in the newly generated composite image Ic from the current frame F1 are less than the overlapping portions of the first object OB1 and the second object OB2 actually reflected in the current frame F1. Therefore, in the newly generated composite image Ic, pixel loss occurs in the region of difference between the overlapping portions of the two. That is, a portion of the region that was previously hidden by the first object OB1 becomes the defective pixel region Rm.

[0075] The pixel supplementation unit 32 supplements the pixels of the defective pixel region Rm based on at least one valid pixel that is spatially or temporally close to the defective pixel region Rm. As a first supplementation method, the pixel supplementation unit 32 generates supplementary pixels based on the surrounding pixels adjacent to the defective pixel region Rm in the current frame F1.

[0076] For example, the pixel supplementation unit 32 assigns pixels to each pixel within the defective pixel region Rm that are identical to the effective pixels closest to each pixel. Alternatively, the pixel supplementation unit 32 may assign representative values ​​of multiple effective pixels adjacent to the defective pixel region Rm to each pixel within the defective pixel region Rm. Furthermore, for each pixel within the defective pixel region Rm, multiple nearby effective pixels may be determined, and representative values ​​of these determined multiple effective pixels may be calculated.

[0077] For example, the pixel supplementation unit 32 can also interpolate multiple pixels within the defective pixel region Rm based on the effective pixels that are adjacent to the defective pixel region Rm to the left, right, or top and bottom. For example, linear interpolation can be performed on a row-by-row basis. In this case, a gradient can also be added to interpolate multiple pixels within the defective pixel region Rm. When a gradient is added, the color difference between multiple pixels can be equal, or the color difference between each pixel can vary according to a certain rule (e.g., an exponential function variation).

[0078] As a second supplementary method, the pixel supplementation unit 32 searches for frames in which there are valid pixels in the region (hereinafter referred to as the corresponding region) corresponding to the defective pixel region Rm of the current frame F1 among multiple frames that are close to the current frame F1 in time, and supplements the pixels of the defective pixel region Rm of the current frame F1 based on the valid pixels in the corresponding region within the searched frame.

[0079] For example, the pixel supplementation unit 32 supplements the defective pixel region Rm of the current frame F1 with the effective pixels of the corresponding region in the frame that is temporally closest to the current frame F1, where there are effective pixels in the corresponding region. For example, the pixel supplementation unit 32 determines the frame in the frame where there are effective pixels in the corresponding region, and the representative value of the effective pixels in the corresponding region is closest to the representative value of the pixels in the region of the second object OB2 in the current frame F1 where there is a defect. The pixel supplementation unit 32 supplements the defective pixel region Rm of the current frame F1 with the effective pixels of the corresponding region in the determined frame.

[0080] As a third supplementary method, the pixel supplementation unit 32 estimates the original shape of the second object OB2 with a defective portion in the current frame F1, divides the defective pixel region Rm into the region of the second object OB2 and the background region, and supplements pixels for each region.

[0081] exist Figure 7 In the example shown in (b), the pixel supplementation unit 32 draws a straight line L1 between the two points where the outer periphery of the second object OB2 intersects with the outer periphery of the defective pixel region Rm. The pixel supplementation unit 32 estimates the region Rm2 to the left of the straight line L1 in the defective pixel region Rm as the region of the second object OB2, and estimates the region Rm1 to the right of the straight line L1 as the region of the background.

[0082] The pixel supplementation unit 32 supplements pixels in the left region Rm2 of the defective pixel region Rm based on the effective pixels of the second object OB2. At this time, the first supplementation method can be used based on setting the reference range within the effective pixels of the second object OB2. The pixel supplementation unit 32 supplements pixels in the right region Rm1 of the defective pixel region Rm based on the effective pixels of the background near the right region Rm1. At this time, the first supplementation method can be used based on setting the reference range within the effective pixels of the background.

[0083] As explained above, according to Embodiment 2, hand shake correction is performed separately for the subject and background within the same frame. At this time, the correction intensity for the subject and background is changed according to the distance from the shooting unit 10, thereby correcting hand shake while maintaining the composition within the frame. Furthermore, in the event of pixel loss areas due to cropping position movement, the unnaturalness of the image can be reduced by supplementing with spatially or temporally close pixels. Through these methods, even with hand shake, an image that appears as if the shooting device 1 is stationary can be generated.

[0084] Figure 8 This diagram illustrates the structure of the imaging device 1 according to Embodiment 3 of the present invention. The imaging device 1 according to Embodiment 3 includes an imaging unit 10, a vibration detection sensor 11, a distance detection sensor 12, and an image processing device 20. The image processing device 20 according to Embodiment 3 includes an image acquisition unit 21, an image recognition unit 22, an object detection unit 23, an object tracking unit 24, a motion vector detection unit 25, a vector intensity adjustment unit 26, a cropping position determination unit 27, a vibration information acquisition unit 28, a movement amount determination unit 29, a cropping unit 30, an image compositing unit 31, a pixel supplementation unit 32, a distance information acquisition unit 33, and a distance determination unit 34.

[0085] The differences from Embodiments 1 and 2 will be explained below. In Embodiment 3, similar to Embodiment 1, both subject tracking and electronic shakiness correction are used. In Embodiment 3, the vector intensity adjustment unit 26 obtains the motion vector of the object of interest from the motion vector detection unit 25 and the distance to the object from the distance determination unit 34. The vector intensity adjustment unit 26 adjusts the intensity of the motion vector of the object of interest obtained from the motion vector detection unit 25 based on the distance to the object obtained from the distance determination unit 34. The closer the distance to the object, the stronger the intensity of the motion vector of the object is set by the vector intensity adjustment unit 26; the farther the distance to the object, the weaker the intensity of the motion vector is set by the vector intensity adjustment unit 26.

[0086] For example, the relationship between the amount of movement caused by the actual motion of the object, the distance from the shooting unit 10 to the object, and the amount of movement of the object relative to the actual motion of the object within a frame can be specified in advance using a table or function. This relationship can also be derived based on the designer's experiments or simulations. The vector intensity adjustment unit 26 adjusts the intensity of the motion vector of the object obtained from the motion vector detection unit 25 based on the distance from the shooting unit 10 to the object obtained from the distance determination unit 34, referring to the table or function.

[0087] In embodiment 3, the cut position determination unit 27 moves the cut position of the reference frame according to the composite vector of the hand shakiness correction vector and the motion vector of the object adjusted by the vector intensity adjustment unit 26, thereby determining the cut position (reference frame) for the object. Figure 2 (b)). The following processing is the same as in Implementation 1.

[0088] As mentioned above, sometimes multiple objects of interest are set within a frame. In such cases, these objects may move in different directions. The composition within the image synthesized using the field of view for each object may deviate significantly from the actual composition. Therefore, it is also possible to impose a limit on the total cropping range before compositing multiple cropping positions.

[0089] Figure 9 Figures (a)-(c) are specific examples illustrating the limitation processing of the total cropping range of the image processing apparatus 20 according to embodiments 1 and 3. Figure 9 As shown in (a), the cut position determination unit 27 sets the cut position C0 at the center of frame F0 by default. The object detection unit 23 detects the first object OB1 and the second object OB2 as objects of interest within frame F0. The object tracking unit 24 begins tracking the first object OB1 and the second object OB2. Figure 9 In the example shown in (a), the first object OB1 moves to the left, and the second object OB2 moves to the right. Additionally, the camera unit 10 moves upwards due to the photographer's hand tremor.

[0090] like Figure 9 As shown in (b), the cut position determination unit 27 moves the cut position C0 of the reference frame F0 in the current frame F1 according to the composite vector of the hand shakiness correction vector and the motion vector of the first object OB1, thereby determining the cut position C1 for the first object OB1. Similarly, the cut position determination unit 27 moves the cut position C0 of the reference frame F0 in the current frame F1 according to the composite vector of the hand shakiness correction vector and the motion vector of the second object OB2, thereby determining the cut position C2 for the second object OB2.

[0091] exist Figure 9In the example shown in (b), the width W1 of the total cut range before compositing the cut position C1 used for the first object OB1 and the cut position C2 used for the second object OB2 is increased. If the image data of the first object OB1 within cut position C1, the image data of the second object OB2 within cut position C2, and the cut position used for the background (in...) Figure 9 If the background data (not shown in (b)) is synthesized, then both the first object OB1 and the second object OB2 are completely included in the field of view of the synthesized image. However, this will generate an image in which the positional relationship between the first object OB1 and the second object OB2 is closer than it actually is, and the deviation from the actual positional relationship becomes larger.

[0092] Therefore, a limit is set on the total cut range. Specifically, an upper limit Wt for width and an upper limit Ht for height are set on the total cut range. When multiple objects are detected within a frame, the vector intensity adjustment unit 26 adjusts the intensity of the motion vector of at least one of the multiple objects in such a way that the total cut range before the synthesis of the multiple objects converges to the limit of the total cut range.

[0093] exist Figure 9 In the example shown in (c), the vector intensity adjustment unit 26 weakens the intensity of the motion vector of the second object OB2 to satisfy the upper limit Wt of the total cut range width. Specifically, the vector intensity adjustment unit 26 weakens the intensity of the motion vector of the second object OB2 so that the width W1' of the total cut range before compositing the cut position C1 used for the first object OB1 and the cut position C2' used for the second object OB2 is consistent with the upper limit Wt of the total cut range width. In this case, a portion of the second object OB2 is removed from the field of view within the composite image.

[0094] Furthermore, the vector intensity adjustment unit 26 can either reduce the intensity of the motion vector of the first object OB1 without reducing the intensity of the motion vector of the second object OB2, or it can reduce the intensity of both the motion vectors of the first object OB1 and the second object OB2.

[0095] Objects whose motion vector intensity is reduced can also be low-priority objects. Priority can be set as follows: larger objects have higher priority. Alternatively, priority can be set as follows: objects closer to the center of the frame have higher priority. Alternatively, priority can be set as follows: objects closer to the camera unit 10 have higher priority. Alternatively, priority can be set as follows: objects with less motion have higher priority. Furthermore, multiple combinations of these criteria can be used.

[0096] The upper limits of the total cropping range, Wt and Ht, can also be set by the user. Setting these values ​​higher increases the probability that all objects will be included within the field of view of the composite image. Conversely, setting these values ​​lower weakens the correction and reduces deviation from the actual composition.

[0097] As explained above, Embodiment 3 achieves the effects of both Embodiment 1 and Embodiment 2. That is, it can suppress jitter across the entire field of view, generating a naturally composed image within the field of view. Furthermore, Embodiment 2 can also be understood as an example of a static subject in Embodiment 3.

[0098] The present invention has been described above based on embodiments. Those skilled in the art will understand that these embodiments are illustrative, and various modifications can be made to the combination of these constituent elements and processing procedures, and such modifications are also within the scope of the present invention.

[0099] In the above Figure 9 Examples (a)-(c) illustrate how to designate multiple objects to be tracked within a frame. Alternatively, one or more of the multiple objects identified within a frame can be designated as tracked objects, while the remaining objects can be designated as non-tracked objects. In this case, the non-tracked objects are treated as background.

[0100] In embodiments 1-3 described above, an example of generating a composite image by adjusting the cropping field of view in real time during shooting was illustrated. Alternatively, image data for the entire shooting range and sensor information for each frame can be pre-recorded, and after shooting, the recorded image data for the entire shooting range and sensor information for each frame can be read out to perform the composite image generation process described in embodiments 1-3. In this case, the composite image generation process can also be performed not within the shooting device 1, but through another image reproduction device (e.g., a PC or smartphone).

[0101] Industrial availability

[0102] This invention can be used in cameras equipped with electronic image stabilization.

[0103] Symbol Explanation

[0104] 1. Imaging device, 10. Imaging unit, 11. Vibration detection sensor, 12. Distance detection sensor, 20. Image processing device, 21. Image acquisition unit, 22. Image recognition unit, 23. Object detection unit, 24. Object tracking unit, 25. Motion vector detection unit, 26. Vector intensity adjustment unit, 27. Shearing position determination unit, 28. Vibration information acquisition unit, 29. Movement amount determination unit, 30. Shearing unit, 31. Image synthesis unit, 32. Pixel supplementation unit, 33. Distance information acquisition unit, 34. Distance determination unit.

Claims

1. An image processing apparatus, comprising: The image acquisition unit acquires images captured by the shooting unit; The movement determination unit determines the movement of the imaging unit based on the vibration applied to the imaging unit, based on the output signal of a sensor used to detect the vibration applied to the imaging unit or the difference between frames of the acquired image. The object detection unit detects objects of interest within the frames of the acquired images. The motion vector detection unit detects the motion vectors of the object between frames of the image; The cropping position determination unit determines the cropping position for cropping an image of a specified size from each frame of the acquired image. The cropping position determination unit determines, within the current frame, the cropping position for the object that has been moved according to the synthesis vector compared to the previous frame and the cropping position for the background that has been moved according to the correction vector compared to the previous frame. The synthesis vector is a vector synthesized by combining the correction vector and the motion vector of the object. The correction vector is used to eliminate jitter caused by the vibration. as well as The image compositing unit, for the current frame, combines image data of the object cut out from the cut-out position for the object and image data of the background other than the object cut out from the cut-out position for the background.

2. The image processing apparatus according to claim 1, further comprising: The distance determination unit determines the distance from the shooting unit to the object based on the output signal from the sensor used to measure distance or the size of the object within the frame; as well as The vector intensity adjustment unit adjusts the intensity of the motion vector of the object based on a determined distance to the object.

3. The image processing apparatus according to claim 2, wherein, As the distance to the object increases, the vector intensity adjustment unit sets the intensity of the motion vector to be stronger.

4. The image processing apparatus according to claim 1, further comprising: The vector intensity adjustment unit adjusts the intensity of the motion vector of the object detected within the frame. When the vector intensity adjustment unit detects multiple objects within a frame, it adjusts the intensity of the motion vector of at least one of the multiple objects so that the total cut range before the multiple objects are combined falls within the limit of the total cut range.

5. An image processing method, comprising: Acquire images captured by the filming department; The amount of movement of the camera unit based on the vibration applied to the camera unit is determined based on the output signal of the sensor used to detect the vibration applied to the camera unit, or the difference between frames of the acquired image. The objects to be focused on should be detected within the frames of the acquired images; Detect the motion vectors of the objects between frames of the image; In the process of determining the cut position for cropping an image of a specified size from each frame of the acquired image, the cut position for the object, which has been moved relative to the previous frame according to the composite vector, and the cut position for the background, which has been moved relative to the previous frame according to the correction vector, are determined within the current frame. The composite vector is a vector synthesized by combining the correction vector and the motion vector of the object. The correction vector is used to eliminate jitter caused by the vibration. as well as For the current frame, the image data of the object cut out from the cut position for the object and the image data of the background other than the object cut out from the cut position for the background are combined.

6. A computer-readable storage medium storing an image processing program that causes a computer to perform the following processes: Acquire images captured by the filming department; The amount of movement of the camera unit based on the vibration applied to the camera unit is determined based on the output signal of the sensor used to detect the vibration applied to the camera unit, or the difference between frames of the acquired image. The objects to be focused on should be detected within the frames of the acquired images; Detect the motion vectors of the objects between frames of the image; In the process of determining the cut position for cropping an image of a specified size from each frame of the acquired image, within the current frame, the cut position for the object, which has been moved relative to the previous frame according to a composite vector, and the cut position for the background, which has been moved relative to the previous frame according to a correction vector, are determined. The composite vector is a vector synthesized by combining the correction vector and the motion vector of the object. The correction vector is used to eliminate jitter caused by the vibration. For the current frame, the image data of the object cut out from the cut position for the object and the image data of the background other than the object cut out from the cut position for the background are combined.

Citation Information

Patent Citations

  • Imaging apparatus, control method of imaging apparatus, and program

    JP2017098776A

  • Imaging device and image reproduction device

    CN101897174A

  • Image processor, image processing program and image processing method

    JP2009245159A