Image processing device, method, and storage medium

By detecting vibration and distance through the image processing device, adjusting the correction vector strength, cutting and synthesizing the image, the problem of composition changes in electronic hand shake correction is solved, a stable image is generated, background blur is suppressed, and the subject is clear.

CN115804101BActive Publication Date: 2025-09-19JVC KENWOOD CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202180048380.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-09-17
Filing Date
2021-06-01
Publication Date
2025-09-19
Estimated Expiration
2041-06-01

AI Technical Summary

Technical Problem

In electronic hand-shake correction, changes in the composition within the field of view cause changes in the way the positional relationship between the object and the background is observed. Existing optical systems are complex, making it difficult to generate stable images.

Method used

Through the image processing device, the vibration detection sensor and the distance detection sensor are used to determine the movement of the shooting device and the distance to the object, adjust the correction vector intensity, the cutting position determination unit determines the cutting position, and performs image synthesis, and the cutting unit and the pixel supplementation unit process defective pixels.

Benefits of technology

This generates a stable image including the composition even when the shooting viewpoint is shaken, suppresses background blur, ensures clear subjects and natural backgrounds, and is suitable for shooting multiple subjects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115804101B_ABST
    Figure CN115804101B_ABST
Patent Text Reader

Abstract

The object detection unit (23) detects an object of interest within a frame of an image acquired by the imaging unit (10). The vector intensity adjustment unit (26) adjusts the intensity of a correction vector for eliminating jitter caused by vibration applied to the imaging unit (10). The cutting position determination unit (27) determines a cutting position for cutting an image of a specified size from each frame of the acquired image. The cutting position determination unit (27) moves the cutting position according to the correction vector, and determines a cutting position for the object and a cutting position for the background respectively according to the intensity of the correction vector. The image synthesis unit (31) synthesizes image data of the object cut from the cutting position for the object and image data of the background with the object cut out from the cutting position for the background removed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image processing device and an image processing program for processing each frame of an image captured by an imaging unit. Background Art

[0002] Cameras equipped with electronic image stabilization functions are becoming increasingly common. Electronic image stabilization adaptively shifts the field of view (FOV) cut from the captured image to compensate for camera shake, producing images with reduced image shakiness. Electronic image stabilization controls the subject's position within the FOV as closely as possible. However, when the viewpoint shifts due to camera movement, the composition within the FOV changes, altering the way the positional relationship between the subject and the background is perceived.

[0003] Patent Document 1 discloses a technology that combines parallax images captured from multiple different viewpoints of a subject's space to create an image close to the image before the viewpoint change. However, this technology requires an imaging element with multiple pixels per microlens, complicating the optical system.

[0004] Prior art literature

[0005] Patent Literature

[0006] Patent Document 1: Japanese Patent Application Laid-Open No. 2016-42662. Summary of the Invention

[0007] The present embodiment has been made in view of such circumstances, and an object of the present embodiment is to provide a technology for electronically generating a stable image including the composition even if the shooting viewpoint fluctuates.

[0008] To solve the above-mentioned problem, an image processing device according to one aspect of the present embodiment includes: an image acquisition unit that acquires an image captured by an imaging unit; a movement amount determination unit that determines the movement amount of the imaging unit due to the vibration applied to the imaging unit based on an output signal from a sensor for detecting vibration applied to the imaging unit or a difference between frames of the acquired image; an object detection unit that detects an object of interest within a frame of the acquired image; a distance determination unit that determines the distance from the imaging unit to the object based on an output signal from a sensor for measuring distance or the size of the object within the frame; a vector intensity adjustment unit that adjusts the intensity of a correction vector for eliminating blur caused by the vibration; a cropping position determination unit that determines a cropping position for cropping an image of a predetermined size from each frame of the acquired image, the cropping position determination unit shifting the cropping position according to the correction vector and determining a cropping position for the object and a cropping position for the background, respectively, based on the intensity of the correction vector; and an image synthesis unit that synthesizes image data of the object cropped at the cropping position for the object and image data of the background excluding the object cropped at the cropping position for the background.

[0009] Optional combinations of the above-described components and modes in which the expressions of the present embodiment are converted between devices, methods, systems, recording media, computer programs, and the like are also applicable as modes of the present embodiment.

[0010] According to this embodiment, even if the shooting viewpoint fluctuates, a stable image including the composition can be generated electronically. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 This is a diagram showing the configuration of an imaging device according to Embodiment 1 of the present invention.

[0012] Figure 2 (a) to (c) are diagrams (part 1) for explaining a specific example of the synthetic image generation process performed by the image processing apparatus according to the first embodiment.

[0013] Figure 3 (a) to (c) are diagrams (part 2) for explaining a specific example of the synthetic image generation process performed by the image processing apparatus according to the first embodiment.

[0014] Figure 4 This is a diagram showing the configuration of an imaging device according to Embodiment 2 of the present invention.

[0015] Figure 5 (a) to (c) are diagrams (part one) for explaining a specific example of the synthetic image generation process performed by the image processing apparatus according to the second embodiment.

[0016] Figure 6 (a) to (c) are diagrams (part 2) for explaining a specific example of the synthetic image generation process performed by the image processing apparatus according to the second embodiment.

[0017] Figure 7 (a) and (b) are figures (part 3) for explaining a specific example of the synthetic image generation process performed by the image processing device according to the second embodiment.

[0018] Figure 8 This is a diagram showing the configuration of an imaging device according to Embodiment 3 of the present invention.

[0019] Figure 9 (a) to (c) are diagrams for explaining a specific example of the total cropping range limitation process performed by the image processing apparatuses according to the first and third embodiments. DETAILED DESCRIPTION

[0020] Figure 1 1 is a diagram showing the configuration of an imaging device 1 according to Embodiment 1 of the present invention. The imaging device 1 may be a single-unit video camera or a camera module mounted on information equipment such as a smartphone, tablet computer, or notebook PC.

[0021] The imaging device 1 according to the first embodiment includes an imaging unit 10, a vibration detection sensor 11, and an image processing device 20. The imaging unit 10 includes a lens, a solid-state imaging element, and a signal processing circuit. The solid-state imaging element can use, for example, a CMOS (Complementary Metal Oxide Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor. The solid-state imaging element converts light incident through the lens into a movie image signal and outputs it to the signal processing circuit. The signal processing circuit performs signal processing such as A / D conversion and noise removal on the image signal input from the solid-state imaging element and outputs it to the image processing device 20.

[0022] The vibration detection sensor 11 detects vibration applied to the imaging unit 10 and outputs the detected vibration to the image processing device 20. For example, a gyro sensor can be used as the vibration detection sensor 11. The gyro sensor detects vibration applied to the imaging unit 10 in the yaw and pitch directions as angular velocities.

[0023] The image processing device 20 includes an image acquisition unit 21, an image recognition unit 22, an object detection unit 23, an object tracking unit 24, a motion vector detection unit 25, a cropping position determination unit 27, a vibration information acquisition unit 28, a movement amount determination unit 29, a cropping unit 30, an image synthesis unit 31, and a pixel supplementation unit 32. These components can be implemented through the collaboration of hardware and software resources, or solely through hardware resources. Hardware resources include a CPU, ROM, RAM, a GPU (Graphics Processing Unit), a DSP (Digital Signal Processor), an ISP (Image Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), and other LSIs. Software resources include firmware and other programs.

[0024] The image acquisition unit 21 acquires the image captured by the imaging unit 10. The vibration information acquisition unit 28 acquires the output signal of the vibration detection sensor 11 as vibration component information. The movement amount determination unit 29 integrates the output signal acquired by the vibration information acquisition unit 28 to determine the movement amount of the imaging unit 10. For example, the movement amount determination unit 29 integrates the angular velocity signals in the yaw and pitch directions acquired by the vibration information acquisition unit 28 to calculate the movement angles of the imaging unit 10 in the yaw and pitch directions.

[0025] Furthermore, if the movement amount determination unit 29 detects, based on inter-frame differences, that the entire background region, excluding an object recognized by the image recognition unit 22 (described later), is moving uniformly in the same direction, it may determine the movement amount of the entire background region as the movement amount of the imaging unit 10. In this case, even if the vibration detection sensor 11 is omitted, the movement amount of the imaging unit 10 based on the vibration applied to the imaging unit 10 can be estimated. Omitting the vibration detection sensor 11 can reduce costs.

[0026] The movement amount determination unit 29 generates a correction vector for offsetting the determined movement amount of the imaging unit 10. Specifically, the correction vector is generated with a correction amount equal to the movement amount of the imaging unit 10 in the direction in which the imaging unit 10 is shaken. The movement amount determination unit 29 outputs the generated correction vector to the cropping position determination unit 27. When the imaging unit 10 is perfectly stationary, the value of the correction vector is zero.

[0027] The image recognition unit 22 searches for an object within the frame of the image acquired by the image acquisition unit 21. The image recognition unit 22 includes dictionary data containing identifiers for specific objects generated by learning from multiple images of specific objects. Examples of specific objects include human faces, human bodies, faces, animals (e.g., dogs and cats), vehicles (e.g., railway cars), and the like.

[0028] The image recognition unit 22 uses object identifiers to search within the image frame. For object recognition, HOG (Histograms of Oriented Gradients) features can be used, for example. Haar-like features, LBP (Local Binary Patterns) features, and other features can also be used. If an object exists within the frame, the image recognition unit 22 adds a rectangular detection frame to the object.

[0029] The object detection unit 23 determines whether to detect the object recognized by the image recognition unit 22 as an object requiring attention. The object requiring attention is an object estimated to be a subject that the user (photographer) of the imaging device 1 is paying attention to.

[0030] The object detection unit 23 determines whether to treat the identified object as a subject based on any one or more combinations of the following criteria: (a) whether the object is larger than a specified size, (b) whether the object is located in the center of the frame, (c) whether the object is a person or an animal, (d) whether the distance to the object is less than a specified value (a method for estimating the distance to the object will be described later), (e) whether there is any portion obscured by the object, and (f) whether the object's motion is less than a specified value (a method for detecting motion will be described later).

[0031] When multiple objects are identified within a frame, the object detection unit 23 may detect all objects that meet the aforementioned criteria as the subject, or may select a single subject based on any one or more combinations of the following criteria: (a) whether it is the largest object in the frame, (b) whether it is the object located most centrally within the frame, (d) whether it is the closest to the subject, and (f) whether it is the object with the least movement among moving objects.

[0032] These judgment criteria are based on the following empirical rules: most of the subject the photographer is interested in occupies a large area in the frame, is located in the center of the frame, is located near the front of the depth of field, and the photographer tracks the subject using the camera 1 to minimize movement within the frame.

[0033] Furthermore, when the viewfinder screen of the imaging device 1 is a touch panel type, the object detection unit 23 may determine as the subject an object touched by the user of the imaging device 1 among the objects reflected in the viewfinder screen.

[0034] The object tracking unit 24 tracks the object recognized by the image recognition unit 22 within subsequent frames. Object tracking can use, for example, a particle filter or a mean shift method. The tracked objects may be all objects recognized by the image recognition unit 22 or only those detected as objects of interest by the object detection unit 23. Furthermore, if the motion of the object is used as a criterion for selecting the subject, all objects recognized by the image recognition unit 22 must be tracked.

[0035] The motion vector detection unit 25 detects the amount of movement of the object of interest between frames of the video as the motion vector of the object. The motion vector represents a positive vector (tracking vector) of the movement of the object.

[0036] The cropping position determination unit 27 determines a cropping position for cropping an image of a predetermined size from each frame of the video. The cropping position determination unit 27 determines the range of the image to be displayed or recorded by cropping a portion of the entire imaging range captured by all pixels of the solid-state imaging device.

[0037] In this embodiment, an electronic hand-shake correction function is employed. This function adaptively changes the position of the field of view (FOV) clipped from the entire shooting range to compensate for hand-shake in the camera 1. Furthermore, this embodiment also employs a subject tracking function. This function adjusts the position of the FOV clipped from the entire shooting range in response to the subject's movement, ensuring that the subject's position within the FOV is as stable as possible. Thus, in this embodiment, real-time FOV clipping is performed using both the electronic hand-shake correction function and the subject tracking function.

[0038] The cropping position determination unit 27 obtains the correction vector from the movement amount determination unit 29 and obtains the position and motion vector of the attention object from the motion vector detection unit 25. The cropping position determination unit 27 determines the cropping position of the attention object and the cropping position of the background, respectively.

[0039] The cropping position determination unit 27 moves the cropping position for the object in the reference frame according to a composite vector obtained by combining the correction vector and the motion vector of the object, and determines the cropping position for the object in the current frame.

[0040] When the reference frame is the previous frame, the correction vector is used to cancel the motion of the imaging unit 10 between the previous frame and the current frame, and the object motion vector is a motion vector representing the object's motion between the previous frame and the current frame. When the reference frame is the frame at which tracking of the object is started, the correction vector is used to cancel the motion of the imaging unit 10 between the frame at which tracking is started and the current frame, and the object motion vector is a motion vector representing the object's motion between the frame at which tracking is started and the current frame.

[0041] The cropping position determination unit 27 moves the background cropping position of the reference frame according to the correction vector and determines the background cropping position of the current frame. When the reference frame is the previous frame, the correction vector is used to eliminate the motion of the imaging unit 10 between the previous frame and the current frame. When the reference frame is the frame at the start of tracking the object, the correction vector is used to eliminate the motion of the imaging unit 10 between the frame at the start of tracking and the current frame.

[0042] The cropping position determination unit 27 may determine the background cropping position by moving the object cropping position in the current frame according to the object's motion vector. In this case, the object's motion vector represents the motion between the frame at the start of tracking and the current frame.

[0043] The cutting unit 30 cuts out only the image data of the object from the image data at the cutting position for the object. The cutting unit 30 cuts out the image data of the background excluding the object from the image data at the cutting position for the background. The image combining unit 31 combines the cut image data of the object with the cut image data of the background.

[0044] As will be described in detail later, if multiple objects of interest are set within a frame, defective pixels may be generated in the synthesized image. In this case, the pixel supplementation unit 32 supplements the defective pixels in the synthesized image based on at least one valid pixel that is spatially or temporally close to the missing pixel.

[0045] Figure 2 (a) to (c) are diagrams (part 1) for explaining a specific example of the synthetic image generation process by the image processing device 20 according to the first embodiment. Figure 3 (a) to (c) are diagrams (part 2) for explaining a specific example of the synthetic image generation process performed by the image processing device 20 according to the first embodiment.

[0046] like Figure 2As shown in (a), the cropping position determination unit 27 sets the cropping position C0 at the center of the frame F0 by default. The object detection unit 23 detects the object OB1 as the object of attention in the frame F0. The object tracking unit 24 starts tracking the object OB1. Figure 2 In the example shown in (a), the object OB1 moves to the left, and the imaging unit 10 moves upward due to the photographer's hand shaking.

[0047] like Figure 2 As shown in (b), the cut position determination unit 27 moves the cut position C0 of the reference frame F0 in the current frame F1 according to the composite vector of the hand shake correction vector and the motion vector of the object OB1, and determines the cut position C1 for the object OB1.

[0048] like Figure 2 As shown in (c) of FIG. 1 , the cutout position determination unit 27 moves the cutout position C1 for the object OB1 in the current frame F1 according to the inverse vector of the motion vector of the object OB1 , and determines a cutout position C2 for the background.

[0049] like Figure 3 As shown in (a) of FIG. 1 , the clipping unit 30 clips the image data of the object OB1 from the image data of the clipping position C1 for the object OB1 in the current frame F1. Figure 3 As shown in (b), the cutting unit 30 cuts out the image data of the background excluding the image data of the object OB1 from the image data of the cutting position C2 for the background of the current frame F1. Figure 3 As shown in (c) of FIG. 8 , the image synthesis unit 31 synthesizes the image data of the clipped object OB1 and the image data of the clipped background to generate a new synthesized image Ic.

[0050] As described above, according to Embodiment 1, when the field of view is clipped while tracking the subject across the entire shooting range, the field of view for the subject and the field of view for the background are clipped separately, and background blur correction is applied to the background field of view. After background blur correction, the field of view for the subject and the background are combined to suppress background blur. In this case, camera shake correction is applied to both the field of view for the subject and the background. This allows the generation of an image with reduced shake across the entire field of view.

[0051] Therefore, it is possible to clearly capture the subject and generate an image of a natural background. For example, when shooting a cat in the living room, the cat's shaking is suppressed because it is tracked by the subject tracking function. Furthermore, the background will not move suddenly according to the movement of the cat, and it is possible to shoot an image of a natural background. In addition, even when the photographer pans, the movement of the background is suppressed, so it is possible to avoid an image that is difficult to identify what is being shot. In addition, although the details will be described later, in the case of multiple subjects, multiple subjects can be converged into one field of view as much as possible. In addition, by multiplying the intensity of the correction vector used to offset hand shake by the same intensity as the motion vector of the subject, it is possible to shoot an image with a still background such as a virtual background. In addition, it is also possible to create materials that can be used in virtual backgrounds.

[0052] Figure 4 This figure shows the configuration of an imaging device 1 according to a second embodiment of the present invention. The imaging device 1 according to the second embodiment includes an imaging unit 10, a vibration detection sensor 11, a distance detection sensor 12, and an image processing device 20. The image processing device 20 according to the second embodiment includes an image acquisition unit 21, an image recognition unit 22, an object detection unit 23, a vector intensity adjustment unit 26, a cropping position determination unit 27, a vibration information acquisition unit 28, a movement amount determination unit 29, a cropping unit 30, an image synthesis unit 31, a pixel supplementation unit 32, a distance information acquisition unit 33, and a distance determination unit 34.

[0053] The following describes the differences from embodiment 1. The distance detection sensor 12 is a sensor for detecting the distance between an object located in the shooting direction and the shooting unit 10. The distance detection sensor 12 can use a TOF (Time of Flight) sensor, for example. Representative examples of TOF sensors include ultrasonic sensors (sonar) and LiDAR (Light Detection and Ranging). The ultrasonic sensor sends ultrasonic waves in the shooting direction, measures the time until the reflected wave is received, and detects the distance to the object located in the shooting direction. The LiDAR irradiates laser light in the shooting direction, measures the time until the reflected light is received, and detects the distance to the object located in the shooting direction.

[0054] The distance information acquisition unit 33 acquires, as distance information, an output signal from the distance detection sensor 12. The distance determination unit 34 generates a distance image corresponding to the visible light image captured by the imaging unit 10 based on the acquired distance information.

[0055] The distance determination unit 34 can also estimate the distance from the imaging unit 10 to the object captured within the frame through image recognition. For example, the relationship between the general size of the object registered as dictionary data, the size of the object captured within the frame, and the distance from the imaging unit 10 to the object is predefined using a table or function. Based on the size of the object recognized within the frame by the image recognition unit 22, the distance determination unit 34 refers to this table or function to estimate the distance from the imaging unit 10 to the object. In this case, even if the distance detection sensor 12 is omitted, the distance to the object captured within the frame can still be estimated. Omitting the distance detection sensor 12 can reduce costs.

[0056] Furthermore, if the distance from the imaging unit 10 to the object captured in the frame can be obtained from the autofocus adjustment unit (not shown), the distance determination unit 34 may use the distance obtained from the autofocus adjustment unit. Furthermore, if the imaging unit 10 is configured with two eyes, the distance determination unit 34 may estimate the distance to the object captured in the frame based on the parallax between the images captured by the imaging unit 10 for both eyes.

[0057] The subject tracking function is omitted in Embodiment 2. Therefore, it is basically unnecessary to calculate the motion vector of the object, and in Embodiment 2, the object tracking unit 24 and the motion vector detection unit 25 are omitted from the image processing device 20.

[0058] The second embodiment also employs an electronic shake correction function. Specifically, the position of the field of view, which is cut from the entire imaging range, is adaptively shifted to compensate for shake in the camera 1. The second embodiment also incorporates a vector intensity adjustment unit 26 for adjusting the intensity of the correction vector used to eliminate shake.

[0059] The vector intensity adjustment unit 26 receives the correction vector from the movement amount determination unit 29 and receives the distance to the object from the distance determination unit 34. The vector intensity adjustment unit 26 adjusts the intensity of the correction vector received from the movement amount determination unit 29 based on the distance to the object received from the distance determination unit 34. The vector intensity adjustment unit 26 sets the intensity of the correction vector to be stronger as the distance to the object is closer, and sets the intensity of the correction vector to be weaker as the distance to the object is farther.

[0060] For example, the relationship between the amount of hand shake, the distance from the imaging unit 10 to the object, and the amount of hand shake-induced movement of the object captured within the frame may be defined in advance using a table or function. This relationship may also be derived based on the designer's experiments or simulations. The vector intensity adjustment unit 26 refers to this table or function based on the distance from the imaging unit 10 to the object obtained by the distance determination unit 34, and adjusts the intensity of the correction vector received from the movement determination unit 29.

[0061] In the second embodiment, the cropping position determination unit 27 obtains the position of the object of interest from the object detection unit 23 and obtains the correction vector for the object from the vector intensity adjustment unit 26. The cropping position determination unit 27 determines the cropping position of the current frame by shifting the cropping position of a reference frame (e.g., the previous frame) according to the correction vector. In this case, the cropping position determination unit 27 determines the cropping position for the object and the cropping position for the background, respectively, based on the intensities of the correction vectors for the object and background.

[0062] When multiple objects of interest are set within a frame, the cropping position determination unit 27 determines a cropping position for each object of interest. For example, when a first object and a second object located further inward than the first object are detected as objects of interest within a frame, the vector intensity adjustment unit 26 sets the intensity of the correction vector for the first object to be stronger than the intensity of the correction vector for the second object.

[0063] Furthermore, the correction vector for the background may be used directly without adjusting its intensity, or its intensity may be adjusted by the vector intensity adjustment unit 26. For example, the entire background may be divided into multiple backgrounds based on distance, and the vector intensity adjustment unit 26 may adjust the intensity of the correction vector for each of the divided backgrounds. Alternatively, the vector intensity adjustment unit 26 may adjust the intensity of the correction vector based on a representative value (e.g., the average value, median value, or mode) of the distances to the objects constituting the background.

[0064] Figure 5 (a) to (c) are diagrams (part 1) for explaining a specific example of the synthetic image generation process performed by the image processing device 20 according to the second embodiment. Figure 6 (a) to (c) are diagrams (part 2) for explaining a specific example of the synthetic image generation process performed by the image processing device 20 according to the second embodiment. Figure 7 (a) to (b) are diagrams (part 3) for explaining a specific example of the synthetic image generation process performed by the image processing device 20 according to the second embodiment.

[0065] like Figure 5 As shown in (a) of FIG. , the cropping position determination unit 27 sets the cropping position C0 at the center of frame F0 by default. The object detection unit 23 detects the first object OB1 and the second object OB2 as objects of interest within frame F0. As viewed from the imaging unit 10, the second object OB2 is located behind the first object OB1. Within frame F0, a portion of the second object OB2 is obscured by the first object OB1.

[0066] Figure 5 (b) shows the Figure 5 The state shown in (a) is a state in which the imaging unit 10 moves to the right due to hand shaking of the photographer. Figure 5 (b) shows a case where the electronic image stabilization function is turned off. In this case, the cropping position C0 remains unchanged, and the first object OB1 and the second object OB2 move to the left side of the field of view.

[0067] At this time, since the camera unit 10 moves rightward, the viewpoint shifts rightward, changing the way the relative positional relationship between the first object OB1 and the second object OB2 is observed. Specifically, in the current frame F1, the first object OB1, which is closer to the viewpoint, has moved significantly to the left compared to the second object OB2, which is farther away. In other words, compared to the reference frame F0 before the camera shake, in the current frame F1, the overlap between the first object OB1 and the second object OB2 increases, and the portion of the second object OB2 obscured by the first object OB1 increases.

[0068] Figure 5 (c) shows the case where the electronic shake correction function is enabled, and the cut position C0 of reference frame F0 is simply shifted leftward according to the shake correction vector. This shift of cut position C0 positions first object OB1 and second object OB2 at cut position C0' in current frame F1 at the center of the field of view. However, the composition of first object OB1 and second object OB2 in current frame F1 has changed compared to reference frame F0.

[0069] In contrast, Figure 6 (a) to (c) respectively determine the cutting positions of the first object OB1, the second object OB2, and the background. Figure 6 As shown in (a) of FIG. 1 , the cropping position determination unit 27 moves the cropping position C0 of the reference frame F0 according to the correction vector of the first object OB1, and determines the cropping position C1 for the first object OB1. Figure 6 As shown in (b) of FIG. 1 , the cropping position determination unit 27 moves the cropping position C0 of the reference frame F0 according to the correction vector of the second object OB2, and determines the cropping position C2 for the second object OB2. Figure 6 As shown in (c) of FIG. 8 , the cropping position determination unit 27 moves the cropping position C0 of the reference frame F0 according to the correction vector of the background, and determines a cropping position C3 for the background.

[0070] like Figure 7As shown in (a), the cutting unit 30 cuts out image data of the first object OB1 from the image data at the cutting position C1 for the first object OB1, cuts out image data of the second object OB2 from the image data at the cutting position C2 for the second object OB2, and cuts out image data of the background excluding the first object OB1 and the second object OB2 from the image data at the cutting position C3 for the background. The image combining unit 31 combines the cut-out image data of the first object OB1, the cut-out image data of the second object OB2, and the cut-out image data of the background to generate a new combined image Ic.

[0071] The composition of the first object OB1 and the second object OB2 in the newly generated synthetic image Ic is the same as the composition of the first object OB1 and the second object OB2 in the clipping position C0 of the reference frame F0. That is, a virtual viewpoint change can be performed as if the imaging unit 10 has not moved.

[0072] The overlap between first object OB1 and second object OB2 in the newly generated composite image Ic based on current frame F1 is smaller than the overlap between first object OB1 and second object OB2 actually captured in current frame F1. Therefore, in the newly generated composite image Ic, a pixel is missing in the area corresponding to the difference between the overlap. In other words, the portion of the area obscured by first object OB1 becomes defective pixel region Rm.

[0073] The pixel supplementation unit 32 supplements the pixels of the defective pixel region Rm based on at least one valid pixel that is spatially or temporally close to the defective pixel region Rm. As a first supplementation method, the pixel supplementation unit 32 generates supplementary pixels from surrounding pixels adjacent to the defective pixel region Rm in the current frame F1.

[0074] For example, the pixel supplementation unit 32 assigns the same pixel as the nearest valid pixel to each pixel in the defective pixel region Rm. Alternatively, the pixel supplementation unit 32 may assign the representative value of multiple valid pixels adjacent to the defective pixel region Rm to each pixel in the defective pixel region Rm. Alternatively, the pixel supplementation unit 32 may identify multiple valid pixels adjacent to each pixel in the defective pixel region Rm and calculate the representative value of these specific valid pixels.

[0075] For example, the pixel supplementation unit 32 may interpolate multiple pixels within the defective pixel region Rm from the valid pixels adjacent to the left, right, or above and below the defective pixel region Rm. For example, linear interpolation may be performed on a line-by-line basis. In this case, a gradient may be applied to interpolate multiple pixels within the defective pixel region Rm. When applying a gradient, the color difference between the multiple pixels may be uniform, or the color difference between the pixels may vary according to a certain rule (e.g., an exponential function).

[0076] As a second supplementation method, the pixel supplementation unit 32 searches for a frame in which there are valid pixels in an area corresponding to the defective pixel area Rm of the current frame F1 (hereinafter referred to as the corresponding area) among multiple frames that are close to the current frame F1 in time, and supplements the pixels of the defective pixel area Rm of the current frame F1 based on the valid pixels of the corresponding area within the searched frame.

[0077] For example, the pixel supplementing unit 32 supplements the defective pixel region Rm of the current frame F1 with valid pixels in the corresponding region within the frame that is temporally closest to the current frame F1. For example, the pixel supplementing unit 32 identifies a frame in which the representative value of the valid pixels in the corresponding region is closest to the representative value of the pixels in the region of the second object OB2 having the defective portion in the current frame F1. The pixel supplementing unit 32 supplements the valid pixels in the corresponding region within the identified frame to the defective pixel region Rm of the current frame F1.

[0078] As a third supplementation method, the pixel supplementation unit 32 estimates the original shape of the second object OB2 having a defective portion in the current frame F1, divides the defective pixel region Rm into a region of the second object OB2 and a region of the background, and supplements pixels for each region.

[0079] exist Figure 7 In the example shown in (b), the pixel supplementation unit 32 draws a straight line L1 between two points where the periphery of the second object OB2 intersects the periphery of the defective pixel region Rm. The pixel supplementation unit 32 estimates the area Rm2 to the left of the straight line L1 in the defective pixel region Rm as the area of ​​the second object OB2 and the area Rm1 to the right of the straight line L1 as the background area.

[0080] The pixel supplementation unit 32 supplements the pixels in the left region Rm2 of the defective pixel region Rm based on the valid pixels of the second object OB2. In this case, the first supplementation method can be used, with the reference range set to the valid pixels of the second object OB2. The pixel supplementation unit 32 supplements the pixels in the right region Rm1 of the defective pixel region Rm based on the valid pixels of the background adjacent to the right region Rm1. In this case, the first supplementation method can be used, with the reference range set to the valid pixels of the background.

[0081] As described above, according to Embodiment 2, hand shake correction is performed separately for the subject and background within the same frame. By varying the correction intensity for the subject and background based on their distance from the imaging unit 10, hand shake correction can be performed while maintaining the composition within the frame. Furthermore, if pixel loss occurs due to shifting cropping positions, image artifacts can be reduced by supplementing pixels that are spatially or temporally close. As described above, even with hand shake, an image can be generated that appears as if the imaging device 1 were stationary.

[0082] Figure 8 This figure shows the configuration of an imaging device 1 according to a third embodiment of the present invention. The imaging device 1 according to the third embodiment includes an imaging unit 10, a vibration detection sensor 11, a distance detection sensor 12, and an image processing device 20. The image processing device 20 according to the third embodiment includes an image acquisition unit 21, an image recognition unit 22, an object detection unit 23, an object tracking unit 24, a motion vector detection unit 25, a vector intensity adjustment unit 26, a cropping position determination unit 27, a vibration information acquisition unit 28, a movement amount determination unit 29, a cropping unit 30, an image synthesis unit 31, a pixel supplementation unit 32, a distance information acquisition unit 33, and a distance determination unit 34.

[0083] The following describes the differences from Embodiments 1 and 2. In Embodiment 3, similarly to Embodiment 1, both the subject tracking function and the electronic hand-shake correction function are employed. In Embodiment 3, the vector intensity adjustment unit 26 obtains the motion vector of the object of interest from the motion vector detection unit 25 and obtains the distance to the object from the distance determination unit 34. The vector intensity adjustment unit 26 adjusts the intensity of the motion vector of the object of interest obtained from the motion vector detection unit 25 based on the distance to the object obtained from the distance determination unit 34. The closer the distance to the object, the stronger the intensity of the motion vector of the object, and the farther the distance to the object, the weaker the intensity of the motion vector of the object.

[0084] For example, a table or function may be predefined to define the relationship between the amount of movement caused by the actual motion of the object, the distance from the imaging unit 10 to the object, and the amount of movement of the object within the frame relative to the actual motion of the object. This relationship may also be derived based on the designer's experiments or simulations. The vector intensity adjustment unit 26 refers to this table or function based on the distance from the imaging unit 10 to the object obtained from the distance determination unit 34, and adjusts the intensity of the motion vector of the object obtained from the motion vector detection unit 25.

[0085] In the third embodiment, the cropping position determination unit 27 moves the cropping position of the reference frame according to the composite vector of the hand shake correction vector and the motion vector of the object adjusted by the vector intensity adjustment unit 26, and determines the cropping position for the object (reference Figure 2 (b)). The following processing is the same as that of Implementation 1.

[0086] As mentioned above, multiple objects of interest may be set within a frame. In this case, multiple objects may move in different directions. In this case, the composition within the image synthesized with the field of view angles for each object may deviate significantly from the actual composition. Therefore, it is possible to impose a limit on the total cropping range before combining multiple cropping positions.

[0087] Figure 9 (a) to (c) are diagrams for explaining a specific example of the total cropping range limitation process performed by the image processing apparatus 20 in the first and third embodiments. Figure 9 As shown in (a), the cropping position determination unit 27 sets the cropping position C0 at the center of the frame F0 by default. The object detection unit 23 detects the first object OB1 and the second object OB2 as the objects to be noticed in the frame F0. The object tracking unit 24 starts tracking the first object OB1 and the second object OB2. Figure 9 In the example shown in (a), the first object OB1 moves to the left, and the second object OB2 moves to the right. In addition, the imaging unit 10 moves upward due to hand shaking of the photographer.

[0088] like Figure 9 As shown in (b) of FIG. 1 , the cutting position determination unit 27 moves the cutting position C0 of the reference frame F0 in the current frame F1 according to the composite vector of the hand shake correction vector and the motion vector of the first object OB1, thereby determining the cutting position C1 for the first object OB1. Similarly, the cutting position determination unit 27 moves the cutting position C0 of the reference frame F0 in the current frame F1 according to the composite vector of the hand shake correction vector and the motion vector of the second object OB2, thereby determining the cutting position C2 for the second object OB2.

[0089] exist Figure 9 In the example shown in (b), the width W1 of the total clipping range before the combination of the clipping position C1 for the first object OB1 and the clipping position C2 for the second object OB2 becomes wider. If the image data of the first object OB1 in the clipping position C1, the image data of the second object OB2 in the clipping position C2, and the clipping position for the background (in Figure 9If the background data within (b) is synthesized (not shown), both first object OB1 and second object OB2 are completely within the field of view within the synthesized image. However, the resulting image will be closer in position to the actual positional relationship between first object OB1 and second object OB2, significantly deviating from the actual positional relationship.

[0090] Therefore, the total cropping range is limited. Specifically, an upper width limit Wt and an upper height limit Ht are set within the total cropping range. If multiple objects are detected within a frame, the vector intensity adjustment unit 26 adjusts the intensity of the motion vector of at least one of the multiple objects so that it falls within the limits of the total cropping range.

[0091] exist Figure 9 In the example shown in (c), the vector intensity adjustment unit 26 weakens the intensity of the motion vector of the second object OB2 so as to satisfy the upper limit Wt of the width of the total cropping range. Specifically, the vector intensity adjustment unit 26 weakens the intensity of the motion vector of the second object OB2 so that the width W1' of the total cropping range before combining the cropping position C1 for the first object OB1 and the cropping position C2' for the second object OB2 coincides with the upper limit Wt of the width of the total cropping range. In this case, a portion of the second object OB2 is offset from the angle of view within the composite image.

[0092] Alternatively, the vector intensity adjustment unit 26 may weaken the intensity of the motion vector of the first object OB1 without weakening the intensity of the motion vector of the second object OB2 , or may weaken the intensities of both the motion vectors of the first object OB1 and the second object OB2 .

[0093] The object whose motion vector intensity is weakened may be a low-priority object. Alternatively, the larger the object, the higher the priority. Alternatively, the closer the object is to the center of the frame, the higher the priority. Alternatively, the closer the object is to the imaging unit 10, the higher the priority. Alternatively, the smaller the object's motion, the higher the priority. Furthermore, a combination of these criteria may be used.

[0094] The user can also set and modify the maximum width Wt and maximum height Ht values ​​for the total cropping range. Larger values ​​increase the probability that all objects will fit within the composite image's field of view. On the other hand, smaller values ​​reduce the correction intensity, minimizing deviations from the actual composition.

[0095] As described above, according to Embodiment 3, the effects of both Embodiments 1 and 2 are achieved. Specifically, it is possible to suppress jitter across the entire viewing angle and produce a naturally composed image within the viewing angle. Furthermore, Embodiment 2 can also be understood as an example of the case where the subject in Embodiment 3 is stationary.

[0096] While the present invention has been described above based on the embodiments, those skilled in the art will appreciate that the embodiments are merely examples and that various modifications are possible in the combination of components and processing steps, and that such modifications are also within the scope of the present invention.

[0097] In the above Figure 9 (a) to (c) illustrate an example of setting multiple objects to be tracked within a frame. Alternatively, one or more objects among the multiple objects identified within a frame may be set as objects to be tracked, while the remaining objects may be set as objects not to be tracked. In this case, the objects not to be tracked are treated as background.

[0098] In the above-mentioned embodiments 1-3, examples of generating a composite image by adjusting the cropped field angle in real time during shooting were described. In this regard, it is also possible to pre-record the image data of the entire shooting range and the sensor information at each frame time. After shooting is completed, the recorded image data of the entire shooting range and the sensor information at each frame time are read out to perform the composite image generation process described in the above-mentioned embodiments 1-3. In this case, the composite image generation process can also be performed not within the camera 1, but on another image reproduction device (e.g., a PC or smartphone).

[0099] Industrial Applicability

[0100] The present invention can be used in cameras equipped with an electronic hand-shake correction function.

[0101] Explanation of symbols:

[0102] 1. Camera

[0103] 10 Filming Department

[0104] 11 Vibration detection sensor

[0105] 12 distance detection sensors

[0106] 20 Image processing device

[0107] 21 Image Acquisition Department

[0108] 22 Image Recognition Unit

[0109] 23 Object Detection Unit

[0110] 24 Object Tracking Unit

[0111] 25 Motion vector detection unit

[0112] 26 Vector strength adjustment unit

[0113] 27 Cutting position determination unit

[0114] 28. Vibration information acquisition unit

[0115] 29 Movement amount determination unit

[0116] 30 Shearing section

[0117] 31 Image synthesis unit

[0118] 32-pixel supplement

[0119] 33 Distance Information Acquisition Unit

[0120] 34 Distance determination unit

Claims

1. An image processing device, comprising: An image acquisition unit, which acquires the image captured by the capturing unit; a movement amount determination unit that determines a movement amount of the imaging unit due to the vibration applied to the imaging unit based on an output signal of a sensor for detecting vibration applied to the imaging unit or a difference between frames of acquired images, and generates a correction vector based on the movement amount; An object detection unit detects an object of interest within a frame of the acquired image; a distance determining unit that determines the distance from the imaging unit to the object based on an output signal from a sensor for measuring the distance or a size of the object within a frame; a vector strength adjustment unit for adjusting the strength of a correction vector for eliminating a shake caused by the vibration; a cropping position determining unit that determines a cropping position for cropping an image of a predetermined size from each frame of the acquired image, wherein, when the object detecting unit detects a plurality of objects, the cropping position determining unit moves the cropping position along the correction vector and determines a cropping position for each of the plurality of objects and a cropping position for the background, respectively, based on the strength of the correction vector; as well as an image synthesis unit that synthesizes image data of the object cut out from a cutout position for the object and image data of the background excluding the object cut out from a cutout position for the background, The vector strength adjustment unit sets the strength of the correction vector to be stronger as the distance becomes shorter.

2. The image processing apparatus according to claim 1, wherein: When a first object and a second object located deeper than the first object are detected within a frame, the vector intensity adjustment unit sets the intensity of the correction vector used to move the cutout position for the first object to be stronger than the intensity of the correction vector used to move the cutout position for the second object.

3. The image processing apparatus according to claim 2, wherein: The image synthesizing unit further includes a pixel supplementing unit that, when a defective pixel is generated in the image synthesized by the image synthesizing unit, supplements the defective pixel based on at least one valid pixel that is spatially or temporally close to the defective pixel.

4. An image processing method comprising the following steps: acquiring images taken by a photographing unit; determining an amount of movement of the imaging unit due to the vibration applied to the imaging unit based on an output signal of a sensor for detecting vibration applied to the imaging unit or a difference between frames of acquired images, and generating a correction vector based on the amount of movement; Detecting an object of interest within a frame of the acquired image; determining the distance from the imaging unit to the object based on an output signal from a sensor for measuring distance or a size of the object within a frame; adjusting the strength of the correction vector for eliminating the shake caused by the vibration, setting the strength of the correction vector to be stronger as the distance becomes closer; In a process of determining a cutout position for cutting out an image of a predetermined size from each frame of the acquired image, when a plurality of objects are detected, the cutout position is moved along with the correction vector, and a cutout position for each of the plurality of objects and a cutout position for a background are determined separately based on the strength of the correction vector; and Image data of the object cut out from the cutout position for the object and image data of the background excluding the object cut out from the cutout position for the background are synthesized.

5. A computer-readable storage medium storing an image processing program, wherein the image processing program, when executed, causes a computer to perform the following processing: acquiring images taken by a photographing unit; determining an amount of movement of the imaging unit due to the vibration applied to the imaging unit based on an output signal of a sensor for detecting vibration applied to the imaging unit or a difference between frames of acquired images, and generating a correction vector based on the amount of movement; Detecting an object of interest within a frame of the acquired image; determining the distance from the imaging unit to the object based on an output signal from a sensor for measuring distance or a size of the object within a frame; adjusting the strength of the correction vector for eliminating the shake caused by the vibration, setting the strength of the correction vector to be stronger as the distance becomes closer; In a process of determining a cutout position for cutting out an image of a predetermined size from each frame of the acquired image, when a plurality of objects are detected, the cutout position is moved along with the correction vector, and a cutout position for each of the plurality of objects and a cutout position for a background are determined separately based on the strength of the correction vector; and Image data of the object cut out from the cutout position for the object and image data of the background excluding the object cut out from the cutout position for the background are synthesized.

Citation Information

Patent Citations

  • Image processing system, imaging apparatus, image processing method, program, and storage medium

    JP2016042662A

  • Imaging device and image reproduction device

    CN101897174A

  • Device for assisting automobile driver

    CN1373720A

  • Image processing apparatus, control method thereof, and image capture apparatus

    US20200036895A1