Image processing device, image processing method, and program
The image processing apparatus detects and corrects differences between left and right images in stereo imaging, enhancing VR video quality by notifying users of problematic areas and providing correction methods.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- CANON KK
- Filing Date
- 2024-10-30
- Publication Date
- 2026-05-15
AI Technical Summary
Existing image processing systems for stereo imaging fail to accurately detect and notify users of differences between left and right images, leading to potential flickering, double images, and eye strain during VR video viewing.
An image processing apparatus that calculates feature quantities in sub-regions of left and right images, determines differences exceeding a threshold, and notifies users of problematic areas, providing methods to correct these differences.
Enables users to identify and address image discrepancies in real-time, reducing the likelihood of flickering and double images during VR video shooting.
Smart Images

Figure 2026079403000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to image processing of an image captured by an imaging device capable of stereo imaging.
Background Art
[0002] By viewing a video captured with a stereo lens on a head-mounted display (also referred to as a head-mounted display, HMD), it is possible to view a so-called Virtual Reality (VR) video. The user views the left video presented in the VR video with the left eye and the right video with the right eye, and fuses each video to view a stereoscopic video.
[0003] However, if there is a difference between the left and right videos, the video may flicker or appear as a double image, and may not appear as an appropriate three-dimensional image. Such a video not only impairs the sense of presence of the VR video, but may also cause eye strain and motion sickness when viewed for a long time.
[0004] Therefore, it is required not to generate a difference between the left and right videos when shooting a VR video with a stereo lens. On the other hand, even if the focus of each of the left and right lenses is adjusted visually, a difference in adjustment may occur between the left and right videos, or even if the camera settings at the time of shooting are appropriately implemented, a difference in sharpness or a shift in position may occur depending on the difference in image height between the left and right videos. When shooting a VR video, the user does not know whether such a difference between the left and right videos has occurred, or if it has occurred, in which area of the captured video it has occurred. Therefore, even if there is a difference between the left and right videos, shooting is continued.
[0005] Patent Document 1 discloses a technique for detecting a human face area from multi-eye images and determining whether stereoscopic image display is possible based on the amount of displacement and size difference of the same person. With this technique, it is possible to grasp whether there is a difference between the left and right videos of the VR video in the human face area.
Prior Art Documents
Patent Documents
[0006] [Patent Document 1] Japanese Patent Publication No. 2011-221905 [Overview of the project] [Problems that the invention aims to solve]
[0007] In the technology described in Patent Document 1, the determination of whether a stereoscopic image is possible is made only for the human face region between the left and right images. However, when capturing stereo images, differences occur between the left and right images regardless of the subject, so it is not possible to fully grasp the differences that occur between the left and right images. [Means for solving the problem]
[0008] The image processing apparatus according to the present invention is characterized by comprising: an acquisition unit that acquires a left image for viewing with the left eye and a right image for viewing with the right eye; a feature quantity calculation unit that calculates feature quantities in respective sub-regions of the left image and the right image; a left-right difference determination unit that determines whether the difference in feature quantities between corresponding sub-regions of the left image and the right image is greater than or equal to a threshold; and a notification unit that notifies the region in which the difference is greater than or equal to a threshold as a result of the determination by the left-right difference determination unit. [Effects of the Invention]
[0009] This invention makes it possible to fully understand the difference that occurs between the left and right images when shooting stereo images. [Brief explanation of the drawing]
[0010] [Figure 1] A schematic diagram showing the hardware configuration of an image processing device. [Figure 2] A block diagram showing the functional configuration of the image processing apparatus in Embodiment 1. [Figure 3] A schematic diagram showing an example of input video in Embodiment 1. [Figure 4] A diagram illustrating the small region used in feature calculation in Embodiment 1. [Figure 5]A diagram illustrating the spatial frequency components calculated by feature calculation in Embodiment 1. [Figure 6] A schematic diagram showing an example of the UI notified by the notification unit in Embodiment 1. [Figure 7] A flowchart illustrating the process in Embodiment 1. [Figure 8] A block diagram showing the functional configuration of the image processing apparatus in Embodiment 2. [Figure 9] A schematic diagram showing an example of the UI notified by the notification unit in Embodiment 2. [Figure 10] A flowchart illustrating the process in Embodiment 2. [Figure 11] A block diagram showing the functional configuration of the image processing apparatus in Embodiment 3. [Figure 12] A flowchart illustrating the process in Embodiment 3. [Modes for carrying out the invention]
[0011] Embodiments of the present invention will be described below with reference to the drawings. Note that the following embodiments are not limiting to the present invention, and not all combinations of features described in these embodiments are essential to the solution of the present invention. Identical components will be denoted by the same reference numerals.
[0012] [Embodiment 1] This embodiment describes a process for determining whether there is a difference between the left and right images that may cause viewing problems when shooting VR video, and for notifying the user of the region where the difference between the left and right images occurs. First, Figure 1 shows an example of the configuration of the image processing device 101 in this embodiment.
[0013] The image processing device 101 is configured as an example of an imaging device capable of stereo imaging. For example, it may be an imaging device with two fisheye lenses arranged side by side, capable of stereo 180-degree imaging, but it can also be operated on a personal computer (PC) connected to the imaging device.
[0014] In FIG. 1, the CPU (Central Processing Unit) 102 uses the RAM (Random Access Memory) 103 as a work memory and cooperates with other components based on the operating system (OS) and various programs. That is, it controls the operation of the entire image processing apparatus 101. Also, the CPU 102 controls each component via the system bus 112. In this embodiment, it is described assuming that there is one CPU, but it is not limited thereto, and a configuration with a plurality of CPUs may be adopted.
[0015] The image signal processor (ISP) 104 is a dedicated processor for performing image processing. The stereo lens optical system 106 refers to, for example, a fisheye lens with a 180-degree angle of view having two left and right lenses, and is arranged so that incident light is imaged on the image sensor 105. The incident light obtained from the stereo lens optical system 106 is converted into an electronic image by the image sensor 105, and the ISP 104 performs corrections specific to the image sensor 105. The captured image data is temporarily stored in the RAM 103, and after the CPU 102 and the graphic processor 107 perform high-quality processing and encoding processing, it is recorded in the external storage 111.
[0016] Hereinafter, when acquiring and expressing video data, it is assumed that it is acquired from either the ISP 104 or the external storage 111. In this embodiment, the internal format is described as an RGB image, but it is not limited thereto, and it may be a YUV image or a monochrome luminance image. Also, the graphic processor 107 can encode and decode video in real time and performs the calculation processing required when displaying video on the display 109.
[0017] Display 109 is a display device that shows commands input from the user interface (I / F) 108, response outputs from an externally connected PC (not shown), etc. User interface (UI) screens and image processing results can be displayed on display 109 via the graphics processor 107. The graphics processor 107 is capable of performing geometric transformations on input images and can input / output images from RAM 103 or output directly to display 109. User I / F 108 is an interface in which a touch panel, switches, buttons, etc. are integrated and accept user operations such as starting and stopping recording.
[0018] The external storage 111 is intended to function as a so-called memory card. The external data input / output I / F 110 performs data communication via the network.
[0019] Next, the functional configuration of the image processing device 101 in this embodiment will be explained with reference to Figure 2. The image processing device 101 consists of a left / right image acquisition unit 201 (also called the acquisition unit 201), a feature quantity calculation unit 202, a left / right difference determination unit 203, and a left / right difference region notification unit 204 (also called the notification unit 204).
[0020] The acquisition unit 201 acquires video data (hereinafter referred to as input video 301) based on instructions from the CPU 102. In this embodiment, the video data displayed during the imaging process, which is displayed as a so-called live view, is treated as input video 301.
[0021] Figure 3(a) shows an example of the input video 301. The input video 301 consists of a left video 302 and a right video 303. The video acquired from the left lens of the stereo lens optical system 106 is represented as the left video 302, and the video acquired from the right lens of the stereo lens optical system 106 is represented as the right video 303. In this embodiment, the left video 302 and the right video 303 are assumed to be 180-degree images in both the vertical and horizontal directions. The input video 301 is displayed on the display 109, and the user can take pictures while checking the input video 301 displayed on the display 109. In addition, the left video 302 and the right video 303 are acquired as images with parallax. By using an HMD (not shown) or the like, the left video can be projected onto the user's left eye and the right video onto the user's right eye, it is possible to view a 180-degree image with a sense of depth.
[0022] The feature calculation unit 202 calculates and stores feature quantities for all regions of the left video 302 and right video 303 acquired by the acquisition unit 201, based on instructions from the CPU 102. Specifically, it divides the left video 302 and right video 303 into sub-regions, and then calculates feature quantities for each sub-region. All divided sub-regions are assumed to contain all regions of the left and right videos.
[0023] Furthermore, the feature quantities calculated in this embodiment refer to spatial frequency components and edge positions.
[0024] Here, we will explain why spatial frequency components and edge positions are used as features. First, regarding spatial frequency components, differences can occur between the left and right images when the degree of focus adjustment of the left and right lenses in the stereo lens optical system 106 is misaligned, or when the image height of the same subject differs between the left and right images. This indicates that the sharpness of the image differs between the left and right images, and when viewing such an image, the highly sharp and less sharp areas appear as if they are a double image. Therefore, spatial frequency components are treated as features.
[0025] Edge position can differ between left and right images when the image height of the same subject differs between the two images. Normally, users fuse the images projected onto the left and right retinas in their brains and perceive them as a single image. In normal vision, the images projected onto the left and right eyes are of the same object, so vertical positional shifts do not occur except due to the tilt of the eyes. However, in VR shooting, the position of the subject may shift vertically due to differences in image height in the input image, and in this case, the subject may not fuse properly and may appear as a double image. Therefore, edge position is treated as a feature. In particular, if a pattern such as a checkerboard is placed at the edges of the image, the vertical difference is likely to be more noticeable.
[0026] Regarding such differences between the left and right images, it is necessary to adjust while visually checking the display during shooting, but the display size is often small, making it difficult to check whether a difference is occurring and where the difference is occurring.
[0027] The following sections will explain the domain for calculating features and the method for calculating features (spatial frequency components and edge positions).
[0028] <Explanation of the area for calculating features> The following describes the regions in which the feature calculation unit 202 calculates features. As mentioned above, the input video 301 is divided into a group of sub-regions, and the sub-regions are set so that all sub-regions include the entire area of the input video. Several examples of how to set the sub-regions are shown below.
[0029] First, subject detection is performed on the video, and sub-regions are set for each subject in the left and right video. Known object detection methods such as R-CNN can be used for subject detection. Figure 3(b) shows an example of subject detection on the input video 301. Figures 3(b) 302a and 302b show that a car, one of the subjects, is detected in the left video 302 and the right video 303, respectively, and this is used as subject information. The subject information is maintained by grouping the subject and the position of the pixels within the detected image region for each subject. Grouping can be achieved by assigning a flag to each pixel value indicating the group to which it belongs. Of the subjects indicated by the subject information, the region 302a detected in the left video and 303a detected in the right video are set as sub-regions. Alternatively, a certain range of region centered on the position coordinates of regions 302a and 303a can be set as a sub-region. Furthermore, the amount of parallax between the left and right videos of the detected subject can be calculated, and regions with similar parallax amounts can be set as sub-regions. This subject detection process is repeated to set sub-regions that include the entire image area.
[0030] Alternatively, instead of performing subject detection, corresponding points may be calculated between the left and right images using luminance, chromaticity, and edge information in the video, and the area around these corresponding points may be defined as a sub-region. Specifically, feature points like 302b and 303b shown in Figure 3(b) may be extracted from the image's luminance information, and the position coordinates of 302b and 303b may be set as corresponding points in the left and right images. The pixels surrounding these corresponding points are then defined as a sub-region. This process is repeated until the sub-region includes the entire image area.
[0031] Furthermore, if there is little parallax between the left and right images, the input image can be divided into smaller regions by block division. Figure 4 shows a projection diagram of an image captured with a fisheye lens as an example of block division. The thick line in Figure 4 indicates the shooting range of the image captured with the fisheye lens, showing a shooting range of 180° in all directions (up, down, left, and right).
[0032] The fisheye lens projection method is equidistant projection, where distances and angles are equally spaced in concentric circles from the center of the image. The region 401 shown in the pattern in Figure 4 may be a sub-region in the left and right images. Note that calculations may be performed using other projection methods, such as orthogonal projection or stereoscopic projection, rather than just equidistant projection.
[0033] Using the method described above, small regions are calculated in the left image 302 and the right image 303, and this process is performed in all regions within the images.
[0034] <An example of a method for calculating spatial frequency components> The method for calculating spatial frequency components, one of the features, is explained below using Figure 5. Figure 5 is a graph showing the spatial frequency components in small regions of the left image 302 and the right image 303. The horizontal axis of the graph in Figure 5 is spatial frequency, which indicates the number of light and dark fringes per unit length of the image. The units include cycles / px and cycles / degree, which are the number of fringes per unit of pixels or angle, respectively. The vertical axis is the power spectrum of the image, which indicates the magnitude of the spatial frequency components of the image corresponding to the spatial frequency on the horizontal axis. 502 and 503 in the graph indicate the spatial frequency components in small regions of the left image 302 and the right image 303, respectively.
[0035] The spatial frequency component 502 can be calculated by performing a two-dimensional Fourier transform on a small region of the left image 302 to convert it into a two-dimensional spatial frequency component, and then converting it to one dimension by taking an equidistant average in the circumferential direction. Similarly, the same process is performed on a small region of the right image 303 for the spatial frequency component 503. Note that this is not the only method for calculating the spatial frequency component; a two-dimensional spatial frequency component may also be used, and when converting to one dimension, an equidistant average may be taken in the vertical or horizontal direction instead of the circumferential direction.
[0036] As shown in Figure 5, the power spectrum of spatial frequency component 502 is higher than that of spatial frequency component 503, except for the low-frequency components. In other words, the left image 302 is a sharper image than the right image 303, except for the low-frequency band.
[0037] <Example of edge position calculation method> Next, an example of how to calculate edge positions will be explained. For example, edge detection is performed in small regions of the left image 302 and the right image 303. Edge detection can be performed by applying edge enhancement filters such as a Prewitt filter or a Laplacian filter. The position coordinates of the edges detected in each of the small regions of the left image 302 and the right image 303 are stored, and this process is repeated for all small regions of the input image. Note that the position coordinates stored here may be limited to regions where edges were strongly detected, or they may be stored for all detected edges.
[0038] The left-right difference determination unit 203, based on instructions from the CPU 102, compares the feature quantities of corresponding sub-regions in the left image 302 and right image 303, which were calculated by the feature quantity calculation unit 202. It then determines whether the difference in feature quantities between the sub-regions (partial regions) is greater than or equal to a threshold.
[0039] The corresponding sub-regions in the left image 302 and the right image 303 refer to the areas that the user perceives as a single image when viewing the video, such as the same subject in both the left and right images. Furthermore, as explained in the feature calculation unit 202, 302a and 303a in Figure 3(b), 302b and 303b, and sub-region 401 in Figure 4 are all corresponding sub-regions in the left and right images.
[0040] Furthermore, the threshold is set to a value that may cause flickering or double images when viewing the input video 301, indicating that if the difference in feature quantities is greater than or equal to the threshold, there is a risk of problems occurring when viewing the video. Below, an example of how to determine left-right differences with respect to the spatial frequency component and edge position, which are features calculated by the feature calculation unit 202, is explained.
[0041] First, in terms of spatial frequency, the ratio of the spatial frequency component 502 of the left image 302 and the spatial frequency component 503 of the right image 303, as shown in Figure 5, is taken. If the maximum value of this ratio is greater than or equal to a threshold, it is determined that a difference between the left and right images has occurred. For example, if the threshold is set to 2, a difference between the left and right images can be considered to have occurred if the maximum value of the ratio of the spatial frequency components is twice the maximum value. Alternatively, the threshold can be set using the maximum value of the difference between the spatial frequency components instead of the ratio.
[0042] Next, at the edge position, the position coordinates of the edges at corresponding positions in the left image 302 and the right image 303, calculated by the feature calculation unit 202, are compared. If the difference in the vertical position coordinates is greater than or equal to a threshold, it is determined that a left-right difference has occurred. For example, the threshold can be set to 1.5% of the image height.
[0043] If the vertical pixel count of input video 301 is 4096px, then a difference of 61px or more in the vertical position of the edges can be considered as a left-right difference. Alternatively, instead of using 1.5% of the image height, the setting can be adjusted considering the display field of view of the HMD used by the user during viewing. If the HMD's display field of view is 90 degrees vertically, a difference of 1.35 degrees or more in pixels can be considered as a left-right difference.
[0044] Furthermore, the feature calculation unit 202 may calculate the amount of parallax between the left and right images, and these thresholds may be changed according to the amount of parallax. For example, if the amount of parallax is large, the image may appear to be at a relatively close distance when viewed, so the thresholds may be set to be smaller and stricter.
[0045] In this way, the left-right difference determination unit 203 determines whether a left-right difference occurs in each corresponding sub-region of the left image 302 and the right image 303, and performs left-right difference determination for all regions in the left image 302 and the right image 303. Note that the threshold values described above are just examples in this embodiment, and other values may be used, or the user may specify them separately through the user I / F 108.
[0046] The notification unit 204, upon receiving instructions from the CPU 102, displays a UI 601 on the display 109 that shows the areas of the left image 302 and right image 303 where the left / right difference determination unit 203 has determined that the difference between the left and right images is above a threshold, thereby notifying the user.
[0047] Figure 6 shows an example of the UI notified by the notification unit 204. In Figure 6, the shaded areas 602a and 602b indicate areas with left-right differences in the left image 302, and the shaded areas 603a and 603b indicate areas with left-right differences in the right image 303. The shaded areas 602a and 603a, and 602b and 603b are small regions corresponding to the left image 302 and right image 303, calculated by the feature calculation unit 202. The notification unit 204 allows the user to easily understand which areas on the input image 301 have left-right differences that cause viewing problems by looking at the display 109.
[0048] In Figure 6, the areas where left-right differences occur are indicated by a diagonal pattern, but these differences can also be represented by changing the color or brightness in the video, or by making the areas with left-right differences blink. Furthermore, the left-right difference determination unit 203 may increase the saturation or brightness as the value exceeds a threshold, making areas with larger left-right differences more prominent. Additionally, the color and brightness may be changed according to the differences in feature quantities calculated by the feature quantity calculation unit 202, for example, using yellow for large differences in spatial frequency and red for large differences in edge position.
[0049] Next, the processing flow of the image processing device 101 in this embodiment will be explained using the flowchart in Figure 7. The CPU 102 reads the program that implements the flowchart shown in Figure 7 and executes it using the RAM 103 as the work area. In this way, the CPU 102 fulfills the role of each functional configuration shown in Figure 2. In the following flowchart, each process (step) will be denoted as "S".
[0050] In S701, the acquisition unit 201 acquires video data of the left video 302 and the right video 303 in the input video 301.
[0051] In S702, the feature calculation unit 202 calculates features for all sub-regions in the left video 302 and right video 303 acquired in S701. Once features have been calculated for all sub-regions in the video, the process proceeds to S703.
[0052] In S703, the left-right difference determination unit 203 determines whether the difference in feature quantities in each corresponding sub-region between the left image 302 and the right image 303, calculated in S702, is greater than or equal to a threshold. The difference in feature quantities is determined for all sub-regions in the image, and if there is a region where the difference in feature quantities is greater than or equal to the threshold, the process proceeds to S704; otherwise, the process shown in Figure 7 is terminated.
[0053] In S704, the notification unit 204 notifies the user via the display 109 of the area where a left-right difference was determined in S703. After completing the processing in S704, the process shown in Figure 7 is terminated. The processing from S701 to S704 is always performed, and if the left and right images change due to a change in the imaging position, the processing is restarted from the beginning, making it possible to notify the user of the areas where a left-right difference occurs in real time.
[0054] As explained above, according to this embodiment, feature quantities are calculated for all areas of the left and right images of the input video, and the user is notified of the video areas where the difference in feature quantities exceeds a threshold. This allows the user to understand where differences that cause viewing problems occur during VR video shooting, and to take measures to reduce the left-right difference at the time of VR video shooting.
[0055] [Embodiment 2] Embodiment 1 describes a method in which feature quantities are calculated for all regions of the left and right images of the input video, and the user is notified of the video regions where the difference in feature quantities exceeds a threshold. However, even if the user can identify where the left-right difference is problematic during viewing, they may not know how to resolve the left-right difference.
[0056] Therefore, in this embodiment, we will describe a method for notifying the user of a correction method to eliminate the left-right difference when a left-right difference occurs. Note that the hardware configuration of the image processing device 101 in this embodiment is the same as in Embodiment 1, so its description will be omitted.
[0057] The functional configuration of the image processing apparatus 101 in this embodiment will be explained with reference to Figure 8. Similar to Embodiment 1, the image processing apparatus 101 in this embodiment has an acquisition unit 201, a feature quantity calculation unit 202, a left-right difference determination unit 203, and a notification unit 204, and further in this embodiment it has a left-right difference correction method notification unit 805. The acquisition unit 201, the feature quantity calculation unit 202, the left-right difference determination unit 203, and the notification unit 204 perform the same processing as in Embodiment 1, so their explanation will be omitted.
[0058] The left-right difference correction method notification unit 805, based on instructions from the CPU 102, notifies the user of a method to correct the left-right difference in the area where the notification unit 204 has notified that there is a left-right difference. The user can correct the difference between the left and right images by implementing the correction method notified by the left-right difference correction method notification unit 805. Figures 9(a) to 9(c) show examples of UIs notified by the left-right difference correction method notification unit 805. Figures 9(a) to 9(c) are displayed by selecting the appropriate UI according to the judgment result determined by the left-right difference determination unit 203. The specific notification content will be explained below using Figure 9.
[0059] First, in Figure 9(a), the notification unit 204 displays a message box 901a overlaid on the UI 601 displayed on the display 109. Figure 9(a) can be applied, for example, when the degree of focus adjustment of the left and right lenses of the stereo lens optical system 106 differs, resulting in a left-right difference in spatial frequency components across almost the entire image. The message box 901a displays a message prompting the user to readjust the focus on the image with the lower spatial frequency components. The user can suppress the occurrence of left-right differences by readjusting the focus of the stereo lens optical system 106.
[0060] In Figure 9(b), the UI 601 displayed by the notification unit 204 is overlaid with guide lines 901b to show the range in which differences in spatial frequency components and edge positions due to differences in image height are less likely to occur. Figure 9(b) can be applied when there are subjects with different image heights in the left and right images, resulting in left-right differences in spatial frequency components and edge positions. The user can suppress the occurrence of left-right differences, which are mainly caused by differences in image height, by changing the composition so that the subject fits within the range of guide lines 901b.
[0061] In Figure 9(c), the user's direction of movement is displayed as an arrow object 901c superimposed on the UI 601 displayed by the notification unit 204, so as to ensure that the composition does not include areas with left-right differences in the shooting range. Figure 9(c) can be applied when there are left-right differences in spatial frequency components, edge positions, etc., for subjects with different image heights on the left and right sides. By moving in the direction indicated by the arrow object 901c, the user can suppress the occurrence of left-right differences, which are mainly caused by differences in image height.
[0062] The processing flow of the image processing apparatus 101 in this embodiment will be explained using the flowchart in Figure 10. The processing from S701 to S704 is the same as in Embodiment 1, so its explanation will be omitted. After the processing in S704 is completed, the process moves to S1005.
[0063] In S1005, the left-right difference correction method notification unit 805 notifies the user via the display 109 of a method to correct the left-right difference in the area where a left-right difference was notified in S704. After completing the process in S1005, the process shown in Figure 10 is terminated. The processes from S701 to S1005 are always performed, and if the user implements the correction method notified in S1005 and the left-right difference is resolved, the process from S701 is restarted.
[0064] As explained above, according to this embodiment, if a left-right difference occurs between the left and right images, the user is notified of a correction method to eliminate the left-right difference. This allows the user to determine what actions to take to eliminate the left-right difference when it occurs, making it easier to reduce the left-right difference in VR images.
[0065] [Embodiment 3] Embodiment 2 describes a method for notifying the user of a correction method to eliminate left-right differences when a left-right difference occurs between the left and right images. In this embodiment, the input image is an image taken with a fisheye lens, and some users may not be familiar with images taken with a fisheye lens, and may find it difficult to grasp the location of the left-right difference even if notified.
[0066] Therefore, this embodiment will describe a method for converting the input video into an image format for viewing on an HMD. The hardware configuration of the image processing device 101 in this embodiment is the same as in Embodiment 1, and its description will be omitted.
[0067] The functional configuration of the image processing device 101 in this embodiment will be explained with reference to Figure 11. Similar to Embodiment 1, the image processing device 101 in this embodiment has an acquisition unit 201, a feature quantity calculation unit 202, a left / right difference determination unit 203, and a notification unit 204, and further has a conversion unit 1106 in this embodiment. The acquisition unit 201, the feature quantity calculation unit 202, the left / right difference determination unit 203, and the notification unit 204 perform the same processing as in Embodiment 1, so their explanation will be omitted.
[0068] The conversion unit 1106, based on instructions from the CPU 102, converts the input video acquired by the acquisition unit 201 into a video format viewable on the HMD. For example, if the video is displayed in equirectangular projection, a known method can be used to convert the input video from a fisheye lens to equirectangular projection. For example, in VR180 format equirectangular projection, the field of view is set so that the field of view per pixel is constant regardless of its position within the image, with a field of view of 180° in all directions. Therefore, when the notification unit 204 notifies the user of areas with left-right differences, it is easier to intuitively understand where the left-right differences are compared to video captured with a fisheye lens. Subsequently, the feature calculation unit 202, the left-right difference determination unit 203, and the notification unit 204 perform processing on the input video 301 converted by the conversion unit 1106.
[0069] The processing flow of the image processing device 101 in this embodiment will be explained using the flowchart in Figure 12. The processing in S701 is the same as in Embodiment 1, and its explanation will be omitted.
[0070] In S1206, the conversion unit 1106 converts the input video acquired in S701 into a video format for viewing on the HMD and stores it.
[0071] In steps S702 to S704, the same processing as in Embodiment 1 is performed on the input video converted in S1206. After the processing in S704 is completed, the process shown in Figure 12 is terminated.
[0072] As explained above, according to this embodiment, after converting the image captured with a fisheye lens into a video format for viewing on an HMD, a process is performed to determine the difference between the left and right sides. This allows the user to more intuitively understand the areas where the difference between the left and right sides occurs, and to more easily take measures to reduce the difference between the left and right sides of the VR image.
[0073] Furthermore, this disclosure includes the following components:
[0074] [Configuration 1] An acquisition unit that acquires a left image for viewing with the left eye and a right image for viewing with the right eye, A feature calculation unit that calculates feature quantities in the respective sub-regions of the left image and the right image, A left-right difference determination unit that determines whether the difference in feature quantities between corresponding subregions in the left image and the right image is greater than or equal to a threshold, Based on the results of the determination by the left-right difference determination unit, a notification unit notifies the region in which the difference is greater than or equal to a threshold. An image processing apparatus characterized by having
[0075] [Configuration 2] The feature calculation unit divides the left image and the right image into small region groups, The image processing apparatus according to configuration 1, characterized in that the group of small regions includes all regions in the left image and the right image.
[0076] [Configuration 3] The image processing apparatus according to configuration 1 or configuration 2, characterized in that the feature quantity is either the spatial frequency component in the left image and the right image, or the edge position in the left image and the right image.
[0077] [Structure 4] The image processing apparatus according to any one of configurations 1 to 3, characterized in that the corresponding partial region is obtained by any of the subject information indicating the subject area, edge information indicating the edge position, brightness information indicating brightness, or chromaticity information indicating chromaticity in the left image and the right image.
[0078] [Composition 5] The image processing apparatus according to any one of configurations 1 to 4, characterized in that the corresponding partial region is obtained by block division, which divides the left image and the right image into small regions called blocks.
[0079] [Composition 6] The image processing apparatus according to any one of configurations 1 to 5, characterized in that the left-right difference determination unit changes the threshold according to the amount of parallax between the left image and the right image.
[0080] [Composition 7] The image processing apparatus according to any one of configurations 1 to 6, further comprising a left-right difference correction method notification unit that notifies a method for correcting the image difference between the left image and the right image.
[0081] [Structure 8] The aforementioned image processing device is The image processing apparatus according to any one of configurations 1 to 7, further comprising a conversion unit that converts the left image and the right image into image formats for display on a head display device.
[0082] [Composition 9] A program for causing a computer to function as one of the means of an image processing device described in any one of configurations 1 through 8. [Explanation of Symbols]
[0083] 201 Left and Right Video Acquisition Unit 202 Feature Calculation Unit 203 Left and right difference determination section 204 Left-right difference area notification section
Claims
1. An acquisition unit that acquires a left image for viewing with the left eye and a right image for viewing with the right eye, A feature calculation unit that calculates feature quantities in the respective sub-regions of the left image and the right image, A left-right difference determination unit that determines whether the difference in feature quantities between corresponding subregions in the left image and the right image is greater than or equal to a threshold, Based on the results of the determination by the left-right difference determination unit, a notification unit notifies the region in which the difference is greater than or equal to a threshold. An image processing apparatus characterized by having
2. The feature calculation unit divides the left image and the right image into small region groups, The image processing apparatus according to claim 1, characterized in that the group of small regions includes all regions in the left image and the right image.
3. The image processing apparatus according to claim 1 or 2, characterized in that the feature quantity is either the spatial frequency component in the left image and the right image, or the edge position in the left image and the right image.
4. The image processing apparatus according to claim 1, characterized in that the corresponding partial region is obtained by any of the subject information indicating the subject area, edge information indicating the edge position, brightness information indicating brightness, or chromaticity information indicating chromaticity in the left image and the right image.
5. The image processing apparatus according to claim 1, characterized in that the corresponding sub-region is obtained by block division, which divides the left image and the right image into small regions called blocks.
6. The image processing apparatus according to claim 1, characterized in that the left-right difference determination unit changes the threshold according to the amount of parallax between the left image and the right image.
7. The image processing apparatus according to claim 1, further comprising a left-right difference correction method notification unit that notifies a method for correcting the difference between the left image and the right image.
8. The aforementioned image processing device is The image processing apparatus according to claim 1, further comprising a conversion unit that converts the left image and the right image into image formats for display on a head display device.
9. A program for causing a computer to function as each of the means of the image processing apparatus described in claim 1.
10. The acquisition process involves obtaining a left image for viewing with the left eye and a right image for viewing with the right eye. A feature calculation step in which feature quantities are calculated in each sub-region of the left image and the right image, A left-right difference determination step that determines whether the difference in feature quantities between corresponding subregions in the left image and the right image is greater than or equal to a threshold, Based on the results of the left-right difference determination step, a notification step is made to notify the region in which the difference is greater than or equal to a threshold. An image processing method characterized by having the following features.