Fixation Correction of Multi-View Images

By adjusting multi-view images to correct perceived gaze direction through image block transformation using feature vectors and reference data, the method addresses the misalignment issue in stereoscopic image systems, enhancing user interaction in teleconferencing.

CN114143495BActive Publication Date: 2025-07-15REALD SPARK LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111303849.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2016-01-05
Filing Date
2017-01-04
Publication Date
2025-07-15
Estimated Expiration
2037-01-04

AI Technical Summary

Technical Problem

During the multi-view image display process, the observer's perceived gaze direction is inconsistent with the actual gaze direction, resulting in an unnatural interactive experience, especially in systems such as conference calls.

Method used

By identifying the left and right eye image blocks of the head in a multi-view image, the feature vectors are derived and the reference displacement vector field is found, and the gaze direction is corrected using machine learning technology to ensure the transformation consistency of the image blocks and reduce the effect of wrong gaze.

Benefits of technology

It effectively corrects the perceived gaze direction, improves the naturalness and interactive experience of multi-view image display, and reduces the discomfort caused by gaze errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114143495B_ABST
    Figure CN114143495B_ABST
Patent Text Reader

Abstract

The present invention discloses gaze correction for multi-view images. The gaze is corrected by adjusting multi-view images of a head. Image patches including the left and right eyes of the head are identified, and feature vectors are derived from a plurality of local image descriptors of the image patches in at least one image of the multi-view images. The derived feature vectors are used to look up reference data generated by machine learning including reference displacement vector fields associated with possible values of the feature vectors, thereby deriving a displacement vector field representing the transformation of the image patches. The multi-view images are adjusted by transforming the image patches including the left and right eyes of the head according to the derived displacement vector field.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the invention patent application with the Chinese patent application number 201780006239.1 and the invention title "Fixation Correction of Multi-Perspective Images", which is the Chinese national phase entry of the PCT international application PCT / US2017 / 012203 filed on January 4, 2017. Technical Field

[0002] This application relates to image processing of multi-perspective images (e.g., stereo image pairs of a head) of a head based on the perceived fixation of the eyes of the head. Background Art

[0003] In many systems, a stereo image pair of a head or more generally multi-perspective images can be captured in one device and displayed on different devices for viewing by an observer. A non-limiting example is a system for conducting a teleconference between two telecommunications devices. In this case, each device can capture a stereo image pair of the head of the observer of that device or more generally multi-perspective images, and transmit it via a telecommunications network to the other device for display and viewing by the observer of the other device.

[0004] When a stereo image pair of a head or more generally multi-perspective images are captured and displayed, the fixation of the head in the displayed stereo image pair or more generally multi-perspective images may not be directed at the observer. This can be caused, for example, by the fixation of the head not being directed at the camera system for capturing the stereo image pair, e.g., because the user whose head is being imaged is observing a display in the same device as the camera system and the camera system is offset upward (or downward) from the display. In this case, the fixation in the displayed image will be perceived as downward (or upward). The human visual system has evolved to be highly sensitive to perceiving fixation using cues obtained from the relative positions of the iris and the white sclera of other observers during social interactions. Therefore, an error in the perceived fixation can be disturbing. For example, in a system for conducting a teleconference, an error in the perceived fixation can cause unnatural interactions between users. Summary of the Invention

[0005] This disclosure relates to image processing techniques for adjusting a stereo image pair of a head or more generally multi-perspective images to correct the perceived fixation.

[0006] According to a first aspect of the present disclosure, there is provided a method for adjusting multi-view images of a head to correct fixation, the method comprising: in each image of the multi-view images, respectively identifying image patches containing the left and right eyes of the head; for the image patches containing the left eye of the head in each image of the multi-view images, and also for the image patches containing the right eye of the head in each image of the multi-view images, performing the following steps: deriving a feature vector from a plurality of local image descriptors of the image patches in at least one image of the multi-view images; and using the derived feature vector to look up reference data including a reference displacement vector field associated with possible values of the feature vector, thereby deriving a displacement vector field representing the transformation of the image patch; and adjusting each image of the multi-view images by transforming the image patches containing the left and right eyes of the head according to the derived displacement vector field.

[0007] In this method, the image patches containing the left and right eyes of the head are identified and transformed. To derive the displacement vector field representing the transformation, a feature vector is derived from a plurality of local image descriptors of the image patches in at least one image of the multi-view images, and the feature vector is used to look up reference data including a reference displacement vector field associated with possible values of the feature vector. The form of the feature vector can be derived from the reference data in advance using machine learning. This method allows fixation to be corrected, thereby reducing the disturbing effect of incorrect fixation when the multi-view images are subsequently displayed.

[0008] Various methods for deriving and using the displacement vector field are possible as follows.

[0009] In a first method, the displacement vector field can be derived independently for the image patches in each image of the multi-view images. This allows fixation to be corrected, but there is a risk that the displacement vector fields for each image may be inconsistent with each other, and as a result, conflicting transformations are performed, which can distort the stereo effect and / or degrade the image quality.

[0010] However, the following alternative method overcomes this problem.

[0011] A second possible method is as follows. In the second method, the plurality of local image descriptors used in the method are the plurality of local image descriptors in two images of the multi-view images. In this case, the reference data includes reference displacement vector fields for each image of the multi-view images, which are associated with possible values of the feature vector. This allows the displacement vector field to be derived from the reference data for each image of the multi-view images. Therefore, the derived displacement vector fields for each image of the multi-view images are inherently consistent.

[0012] A potential disadvantage of this second method is that it may require reference data to be derived from a stereo image or more generally a multi-view image, which may not be convenient to derive. However, the following method allows reference data to be derived from a single field of view image.

[0013] A third possible method is as follows. In the third method, multiple local image descriptors are multiple local image descriptors in one image of a multi-view image, and a displacement vector field is derived as follows. The derived feature vectors are used to look up reference data including a reference displacement vector field associated with possible values of the feature vectors, thereby deriving a displacement vector field representing the transformation of an image patch in the said one image of the multi-view image. Then, the derived displacement vector field representing the transformation of the image patch in the said one image of the multi-view image is transformed according to an estimated value of the optical flow between the image patch in the said one image and image patches in one or more other multi-view images, thereby deriving a displacement vector field representing the transformation of image patches in one or more other multi-view images.

[0014] Thus, in the third method, the displacement vector fields derived for each image are consistent because only one displacement vector field is derived from the reference data and another displacement vector field is derived from it using a transformation based on an estimated value of the optical flow between image patches in the images of the multi-view image.

[0015] A fourth possible method is as follows. In the fourth method, multiple local image descriptors are multiple local image descriptors in two images of a multi-view image, and a displacement vector field is derived as follows. The derived feature vectors are used to look up reference data including a reference displacement vector field associated with possible values of the feature vectors, thereby deriving an initial displacement vector field that represents a conceptual transformation of a conceptual image patch in a conceptual image that has a conceptual camera position relative to the camera positions of the images of the multi-view image. Then, the initial displacement vector field is transformed according to an estimated value of the optical flow between the conceptual image patch in the conceptual image and the image patches in the images of the multi-view image, thereby deriving a displacement vector field representing the transformation of image patches in each image of the multi-view image.

[0016] Thus, in the fourth method, the displacement vector fields derived for each image are consistent because only one displacement vector field is derived from the reference data, which represents a conceptual transformation of a conceptual image patch in a conceptual image that has a conceptual camera position relative to the camera positions of the images of the multi-view image. According to an estimated value of the optical flow between the conceptual image patch in the conceptual image and the images of the multi-view image, corresponding displacement vector fields for transforming the two images of the multi-view image are derived from it using a transformation.

[0017] The fifth possible method is as follows. In the fifth method, a displacement vector field for image patches in each image of a multi-view image is derived, but a combined displacement vector field is then derived therefrom and used to transform the image patches of the left and right eyes that include the head. In this case, the displacement vector fields for each image are consistent because they are the same.

[0018] This combination can be performed in any suitable manner. For example, the combination can be a simple average, or can be an average weighted by confidence values associated with each of the derived displacement vector fields. These confidence values can be derived during machine learning.

[0019] According to a second aspect of the present disclosure, there is provided an apparatus configured to perform a method similar to that of the first aspect of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Non-limiting embodiments are shown by way of example in the drawings, where like reference numerals represent like components, and where:

[0021] Figure 1 is a schematic perspective view of a device for capturing a stereoscopic image pair;

[0022] Figure 2 is a schematic perspective view of a device for displaying a stereoscopic image pair;

[0023] Figure 3 is a flowchart of a method for adjusting a stereoscopic image pair;

[0024] Figure 4 is a schematic illustration showing the processing of a stereoscopic image pair in the method of Figure 3 ;

[0025] Figure 5 is a flowchart of the step of extracting image patches;

[0026] Figure 6 and Figure 7 are flowcharts of the steps of deriving a displacement vector field according to two alternative methods;

[0027] Figure 8 and Figure 9 are flowcharts of two alternative schemes of the step of adjusting an image;

[0028] Figure 10 is a flowchart of a transformation step within the step of adjusting an image in the methods shown in Figure 8 and Figure 9 ; and

[0029] Figure 11 is a schematic illustration of a telecommunications system in which the method can be implemented. DETAILED DESCRIPTION

[0030] Figure 1 and Figure 2 shows how misaligned gaze is perceived when a stereoscopic image pair of a head is captured by the device 10 (which will be referred to as the source device 10) shown in Figure 1 and displayed on a different device 20 (which will be referred to as the target device 20) shown in Figure 2 How misaligned gaze is perceived when a stereoscopic image pair of a head is captured by the device 10 (which will be referred to as the source device 10) shown in

[0031] The capture device 10 includes a display 11, and the camera system 12 includes two cameras 13 for capturing a stereoscopic image pair of the head of the source observer 14. The source observer 14 views the display 11 along line 15. The cameras 13 of the camera system 12 are offset from the display 11, in this case above the display 11. Thus, the cameras 13 actually look down on the source observer 14 along line 16.

[0032] The display device 20 includes a display 21, which is any known type of stereoscopic display, such as any known type of autostereoscopic display. The display 21 displays the stereoscopic image pair captured by the capture device 10. The target observer 24 views the display 21. If the target observer 24 is in the normal viewing position perpendicular to the center of the display 21 (as shown by the solid outline of the target observer 24), the gaze of the source observer 14 is perceived by the target observer 24 as being downward, rather than looking at the target observer 24, because the cameras 13 of the source device 10 look down on the source observer 14.

[0033] Although in this example the cameras 13 are above the display 11, the cameras 13 can generally be in any position adjacent to the display 11, and the gaze of the source observer 14 perceived by the target observer 24 will accordingly be incorrect.

[0034] If the target observer 24 is in an offset viewing position (as shown by the dashed outline of the target observer 24) such that the target observer 24 views the display 21 along line 26, the offset of the target observer 24 creates an additional error in the gaze of the source observer 14 perceived by the target observer 24. If the target observer 24 is in the normal viewing position along line 25, but the stereoscopic image pair is displayed at a position offset from the center of the display 25 on the display 25, a similar additional error in the perceived gaze of the source observer 14 will occur.

[0035] The stereoscopic image pair is an example of a multi-viewpoint image with two images. Although Figure 1 shows an example where the camera system 12 includes two cameras 13 for capturing a stereoscopic image pair, alternatively, the camera system can include more than two cameras 13 for capturing more than two multi-viewpoint images, in which case there is a similar problem of misperceived gaze on the display.

[0036] Figure 3 A method for adjusting a multi-view image to correct such an error in the perceived gaze is shown. For simplicity, the method will be described with respect to the adjustment of a multi-view image including a stereo image pair. Similar processing can be performed on a larger number of images in a straightforward manner, and the method can be generalized to multi-view images including more than two images.

[0037] This method can be executed in the image processor 30. The image processor 30 can be implemented by a processor executing a suitable computer program, or by dedicated hardware, or by some combination of software and hardware. In the case of using a computer program, the computer program can include instructions in any suitable language and can be stored on a computer-readable storage medium, which can be of any type, such as: a recording medium that can be inserted into a drive of a computing system and can store information magnetically, optically, or magneto-optically; a fixed recording medium of a computer system, such as a hard disk drive; or computer memory.

[0038] The image processor 30 can be provided in the source device 10, the target device 10, or any other device, such as a server on a telecommunications network, which can be suitable for the case where the source device 10 and the target device 10 communicate through such a telecommunications network.

[0039] The stereo image pair 31 is captured by the camera system 12. Although the camera system 12 is shown in Figure 1 as including two cameras 13, this is not restrictive, and more generally, the camera system 13 can have the following characteristics.

[0040] The camera system includes a set of cameras 13 having at least two cameras 13. These cameras are typically spaced apart at a distance less than the average inter-pupillary distance of a human. In an alternative where the method is applied to more than two multi-view images, there are then more than two cameras 13, i.e., one camera 13 for each image.

[0041] The cameras 13 are spatially related to each other and to the display 11. The spatial relationships between the cameras 13 themselves and between the cameras 13 and the display 11 are known in advance. Methods known for finding spatial relationships can be applied, such as calibration methods using reference images or prior specifications.

[0042] The cameras 13 face in the same direction as the display 11. Thus, when the source observer 14 is viewing the display 11, the cameras 13 face the source observer 14 and the captured stereo image pair is an image of the head of the source observer 14. The cameras in the camera system can have different fields of view.

[0043] The camera system 12 can include cameras 13 having different sensing modalities (including visible light and infrared).

[0044] The main output of the camera system 13 is a stereo image pair 31, which is typically a video image output at video rate. The output of the camera system 13 may also include data representing the spatial relationship between the camera 13 and the display 11, the nature of the sensing modality, and the internal parameters of the camera 13 (e.g., focal length, optical axis) that can be used for angular positioning.

[0045] The method performed on the stereo image pair 31 is as follows. To illustrate the method, reference is also made to Figure 4 , which shows an example of the stereo image pair 31 at various stages of the method.

[0046] In step S1, the stereo image pair 31 is analyzed to detect the position of the head, and specifically the position of the eyes of the source observer 14 within the stereo image pair 31. This is done by detecting the presence of the head, tracking the head, and locating the eyes of the head. Step S1 can be performed using a variety of techniques known in the art.

[0047] One possible technique for detecting the presence of the head is to use a Haar feature cascade, such as that disclosed in Viola and Jones, “Rapid Object Detection using a Boosted Cascade of Simple Features”, CVPR 2001, pp 1-9 (Viola and Jones, “Rapid Object Detection using a Boosted Cascade of Simple Features”, CVPR 2001, pages 1-9, which is incorporated herein by reference).

[0048] One possible technique for tracking the head is to use the method of active appearance models to provide the position of the head of the object and the position of the eyes, as disclosed, for example, in Cootes et al., "Active shape models - their training and application", Computer Vision and Image Understanding, 61(1):38 - 59, Jan. 1995 (Cootes et al., "Active shape models - their training and application", Computer Vision and Image Understanding, Volume 61, Issue 1, Pages 38 - 59, January 1995) and Cootes et al. "Active appearance models", IEEE Trans. Pattern Analysis and Machine Intelligence, 23(6):681 - 685, 2001 (Cootes et al., "Active appearance models", IEEE Transactions on Pattern Analysis and Machine Intelligence, Volume 23, Issue 6, Pages 681 - 685, 2001, which is incorporated herein by reference).

[0049] In step S1, typically, a set of individual points ("landmarks") are set to the facial region (usually the eyes, such as the corners of the eyes, the positions of the upper and lower eyelids, etc.) to locate the eyes.

[0050] In step S2, image patches containing the left and right eyes of the head are identified in each image 31 of the stereo pair. Figure 4 The identified image patch 32 of the right eye in each image 31 is shown (for clarity, Figure 4 the image patch of the left eye is omitted).

[0051] Step S2 can be performed as follows as Figure 5 shown.

[0052] In step S2 - 1, image patches 32 containing the left and right eyes of the head are identified in each image 31 of the stereo pair. This is done by identifying, in each image 31, the image patch 39 located around the identification points ("landmarks") corresponding to the eye features, as shown, for example, in Figure 4 the figure.

[0053] In step S2-2, the image patches 32 identified in step S2-1 are transformed into a normalized coordinate system, which is the same normalized coordinate system as used in the machine learning process described further below. This transformation is chosen to align the points of the eyes ("landmarks") within the image patches identified in step S1 with predetermined positions in the normalized coordinate system. The transformation may include translation, rotation, and scaling to an appropriate extent to achieve this alignment. The output of step S2-2 is the identified image patches 33 of the right eye in each image in the normalized coordinate system, as for example Figure 4 shown.

[0054] (a) The following steps are performed separately for (i) the image patches containing the left eye of the head in each image 31 of the stereo pair and (ii) the image patches containing the right eye of the head in each image 31 of the stereo pair. For the sake of brevity, the following description will only refer to the image patches and eyes without specifying the left or right eye, but it should be noted that the same steps are performed for both the left and right eyes.

[0055] In step S3, a feature vector 34 is derived from a plurality of local image descriptors of the image patches 33 in at least one image 31 of the stereo pair. According to the method and as described further below, this may be for an image patch in a single image 31 of the stereo pair, or for both images 31 of the stereo pair. Thus, these local image descriptors are local image descriptors derived in the normalized coordinate system.

[0056] The feature vector 34 is a representation of the image patches 33 suitable for finding reference data 35 including a reference displacement vector field, which represents the transformation of the image patches and is associated with the possible values of the feature vector.

[0057] The reference data 35 is obtained and analyzed in advance using machine learning techniques, which derive the form of the feature vector 34 and associate the reference displacement vector field with the possible values of the feature vector. Thus, before returning to Figure 3 the method now the machine learning techniques will be described.

[0058] The training input for the machine learning techniques is two sets of images, which may be stereo image pairs or monoscopic images, as discussed further below. Each set includes images of the heads of the same set of individuals, but captured from the camera at different positions relative to the gaze, such that the perceived gaze is different between them.

[0059] The first set is the input images, which are images of each individual with an incorrect gaze, where the incorrectness is known a priori. Specifically, the images in the first set may be captured by at least one camera at known camera positions, where the gaze of the individual is in different known directions. For example Figure 1For the source device, the camera position can be the position of camera 13, and the gaze of the individual being imaged is directed towards the center of display 11.

[0060] The second group is the output images, which are images of each individual with a correct gaze for a predetermined observer position relative to the display position of the image to be displayed. In the simplest case, the observer position is the normal viewing position perpendicular to the center of the display position, such as shown by the solid outline of the target observer 24 for the Figure 2 target device 20.

[0061] For each image in these two groups, the image is analyzed using the same technique as used in step S1 above to detect the position of the head, particularly the position of the eyes, and then the same technique as used in step S2 above is used to separately identify the image patches containing the left and right eyes of the head. Thereafter, the following steps are performed separately for (a) the image patch containing the left eye of the head in each image and (b) the image patch containing the right eye of the head in each image. For the sake of brevity, the following description will only refer to the image patches and the eyes without specifying the left or right eye, but it should be noted that the same steps are performed for both the left and right eyes.

[0062] Each image patch is transformed into the same normalized coordinate system as used in step S2 above. As described above, this transformation is chosen to align the points of the eyes ("landmarks") with a predetermined position in the normalized coordinate system. The transformation can include translation, rotation, and scaling to an appropriate extent to achieve this alignment.

[0063] Thus, the image patches of each individual's input and output images are aligned in the normalized coordinate system.

[0064] A displacement vector field is derived from the input and output images of each individual, which represents the transformation of the image patches in the input image required to obtain the image patches of the output image, as described below. In the case where the position in the image patch is defined by (x, y), the displacement vector field F is given by:

[0065] F = {u(x, y), v(x, y)}

[0066] where u and v define the horizontal and vertical components of the vector at each position (x, y).

[0067] The displacement vector field F is chosen such that the image patch of the output image O(x, y) is derived from the image patch of the input image I(x, y) as follows:

[0068] O(x, y) = I(x + u(x, y), y + v(x, y))

[0069] For image data from more than one camera, the system derives a displacement vector field of the input images from each camera.

[0070] The process in which the trial eigenvector F' = {u', v'} can be modified to minimize the error, optionally in an iterative process, for example, derives the displacement vector field F of the individual input and output images according to the following formula:

[0071] ∑|O(x,y)-I(x+u'(x,y),y+v'(x,y))| = min!

[0072] As a non-limiting example, the displacement vector field F can be derived as disclosed in Kononenko et al., "Learning To Look Up: Realtime Monocular Gaze Correction Using Machine Learning", Computer Vision and Pattern Recognition, 2015, pp. 4667-4675 (Kononenko et al., "Learning to Look Up: Using Machine Learning for Real-Time Monocular Gaze Correction", Computer Vision and Pattern Recognition, 2015, pp. 4667-4675, which is incorporated herein by reference), where the displacement vector field F is referred to as the "flow field".

[0073] Use machine learning techniques to obtain a mapping from the displacement vector field F of each individual to the corresponding eigenvector derived from multiple local image descriptors of the image patches of the input image.

[0074] The local descriptor captures the relevant information of the local part of the image patch of the input image, and this set of descriptors generally forms a continuous vector output.

[0075] The local image descriptors input into the machine learning process are of a type expected to distinguish different individuals, but the specific local image descriptors are selected and optimized by the machine learning process itself. Generally speaking, the local image descriptors can be of any suitable type, and some non-limiting examples that can be applied in any combination are as follows.

[0076] The local image descriptor can include the value of a single pixel or its linear combination. Such a linear combination can be, for example, the difference between pixels at two points, the core derived within a mask at any position, or the difference between two cores at different positions.

[0077] The local image descriptor can include the distance of the pixel position from the position of the eye point ("landmark").

[0078] Local image descriptors may include SIFT features (Scale-Invariant Feature Transform features), such as disclosed in Lowe, "Distinctive Image Features from Scale-Invariant Keypoints", International Journal of Computer Vision 60(2), pp 91-110 (which is incorporated herein by reference).

[0079] The local image descriptor may include HOG features (Histogram of Oriented Gradients features), for example as disclosed in Dalal et al. "Histograms of Oriented Gradients for Human Detection", Computer Vision and Pattern Recognition, 2005, pp. 886-893 (Dalal et al., "Histograms of Oriented Gradients for Human Detection", Computer Vision and Pattern Recognition, 2005, pp. 886-893, which is incorporated herein by reference).

[0080] The derivation of feature vectors from multiple local image descriptors depends on the type of machine learning applied.

[0081] In a first type of machine learning technique, the feature vector may include features that are values derived from local image descriptors in a discrete space that are binary values or values discretized into more than two possible values. In this case, the machine learning technique associates a reference displacement vector field F derived from the training input with each possible value of the feature vector in the discrete space, so that the reference data 35 is essentially a lookup table. This allows the reference displacement vector field F to be simply selected from the reference data 35 based on the feature vector 34 derived in step S3, as described below.

[0082] In the case where the feature vector includes features that are binary values derived from a local image descriptor, the feature vector has a binary representation. Such binary values can be derived from the value of the descriptor in a variety of ways, such as by comparing the value of the descriptor to a threshold, comparing the values of two descriptors, or comparing the distance of the pixel location from the location of the eye point ("landmark").

[0083] Alternatively, the feature vector may include features as discretized values of local image descriptors. In this case, more than two discrete values for each feature are possible.

[0084] Any suitable machine learning techniques can be applied, such as using decision trees, decision forests, decision ferns, or their sets or combinations.

[0085] As an example, a suitable machine learning technique using a feature vector including features that are binary values derived by comparing a set of individual pixels or their linear combinations with a threshold is disclosed in Ozuysal et al., “Fast Keypoint Recognition in Ten Lines of Code”, Computer Vision and Pattern Recognition, 2007, pp. 1-8 (Ozuysal et al., “Fast Keypoint Recognition in Ten Lines of Code”, Computer Vision and Pattern Recognition, 2007, pp. 1-8, which is incorporated herein by reference).

[0086] As another example, a suitable machine learning technique using the distance between the pixel position and the eye landmark position is disclosed in Kononenko et al., “Learning To Look Up: Realtime Monocular Gaze Correction Using Machine Learning”, Computer Vision and Pattern Recognition, 2015, pp. 4667-4675 (Kononenko et al., “Learning To Look Up: Realtime Monocular Gaze Correction Using Machine Learning”, Computer Vision and Pattern Recognition, 2015, pp. 4667-4675, which is incorporated herein by reference).

[0087] As another example, a suitable machine learning technique using random decision forests is disclosed in Ho, “Random Decision Forests”, Proceedings of the 3rd International Conference on Document Analysis and Recognition, Montreal, QC, 14-16 August 1995, pp. 278-282 (Ho, “Random Decision Forests”, Proceedings of the 3rd International Conference on Document Analysis and Recognition, Montreal, QC, 14-16 August 1995, pp. 278-282, which is incorporated herein by reference).

[0088] In a second type of machine learning technique, the feature vectors can include features that are discrete values of local image descriptors in a continuous space. In this case, the machine learning technique associates a reference displacement vector field F derived from the training input with the possible discrete values of the feature vectors in the continuous space. This allows the displacement vector field F to be derived from the reference data 35 by interpolation from the reference displacement vector field based on the relationship between the feature vector 34 derived in step S3 and the values of the feature vectors associated with the reference displacement vector field.

[0089] Any suitable machine learning technique can be applied, such as using support vector regression.

[0090] As an example, a suitable machine learning technique using support vector regression is disclosed in Drucker et al. “Support Vector Regression Machines”, Advances in Neural Information Processing Systems 9, NIPS 1996, 155 - 161 (Drucker et al., “Support Vector Regression Machines”, Proceedings of the 9th Conference on Advances in Neural Information Processing Systems, NIPS, 1996, pages 155 - 161, which is incorporated herein by reference). The output of this technique is a continuous set of variations in the interpolation direction that forms part of the reference data 35 and is used in the interpolation.

[0091] Machine learning techniques, regardless of their type, also inherently derive the form of the feature vectors 34 that are used to derive the reference displacement vector field F. This is the form of the feature vectors 34 derived in step S3.

[0092] Optionally, the output of the machine learning technique can be increased to provide a confidence value associated with the derivation of the displacement vector field from the reference data 35.

[0093] In the case where the feature vectors include features that are values in a discrete space, a confidence value is derived for each reference displacement vector field.

[0094] An example of deriving a confidence value is to maintain the distribution of the corresponding parts of the input images in the training data for each resulting index (value of the feature vector) in the resulting lookup table. In this case, the confidence value can be the amount of training data resulting in the same index divided by the total number of training data samples.

[0095] Another example of deriving a confidence value is to fit a Gaussian to the distribution of the input images in the training data in each indexed binary, and use the trace of the covariance matrix near the mean as the confidence value.

[0096] In the case where the feature vector includes features that are discrete values of local image descriptors in a continuous space, a confidence value can be derived according to the machine learning method used. For example, when using support vector regression, the confidence value can be the reciprocal of the maximum distance from the support vector.

[0097] When in use, the confidence value is stored as part of the reference data.

[0098] The description now returns to Figure 3 the method of

[0099] In step S4, the feature vector 34 derived in step S3 is used to search for the reference data 35, thereby deriving at least one displacement vector field 37 representing the transformation of the image patch. Since the displacement vector field 37 is derived from the reference data 35, the transformation represented thereby corrects the fixation that will be perceived when displaying the stereoscopic image pair 31.

[0100] In the case where the feature vector 34 includes features that are values in a discrete space and the reference displacement vector field of the reference data 35 includes a reference displacement vector field associated with each possible value of the feature vector in the discrete space, the displacement vector field for the image patch is derived by selecting the reference displacement field associated with the actual value of the derived feature vector 34.

[0101] In the case where the feature vector 34 includes features that are discrete values of local image descriptors in a continuous space, the displacement vector field for the image patch is derived by interpolating the displacement vector field from the reference displacement vector field based on the relationship between the actual value of the derived feature vector 34 and the value of the feature vector associated with the reference displacement vector field. In the case where the machine learning technique is support vector regression, this can be performed using the interpolation direction that forms part of the reference data 35.

[0102] Some different methods for deriving the displacement vector field 37 in step S4 will now be described.

[0103] In the first method, in step S4, the displacement vector field 37 is derived independently for each image patch in each image 31 of the stereo pair. This first method can be applied when deriving the reference data 35 from a single field of view image. This method provides correction of the fixation, but there is a risk that the displacement vector fields 37 for each image may be inconsistent with each other, and as a result, conflicting transformations are subsequently performed, which can distort the stereoscopic effect and / or degrade the image quality.

[0104] Other methods for overcoming this problem are as follows.

[0105] In a second possible method, the plurality of local image descriptors used to derive the feature vector 34 in step S3 are the plurality of local image descriptors in the two images of the stereo pair. In this case, the reference data 35 similarly includes a plurality of pairs of reference displacement vector fields for each image 31 of the stereo image pair, which are the plurality of pairs of reference displacement vector fields associated with the possible values of the feature vector 34.

[0106] This second method allows a pair of displacement vector fields 35 to be derived from the reference data 35, i.e., one displacement vector field for each image 31 of the stereo pair. Thus, the derived displacement vector fields for each image 31 of the stereo pair are inherently consistent, since they are derived together from a consistent pair of reference displacement vector fields in the reference data 35.

[0107] A disadvantage of this second method is that it requires the reference data 35 to be derived from the training input of a machine learning technique for stereo image pairs. This does not pose any technical difficulties, but can create some practical inconveniences, since monocular images are more commonly available. Thus, the following method can be applied when deriving the reference data 35 from the training input of a machine learning technique for monocular images.

[0108] In a third possible method, the feature vector 34 is derived from a plurality of local image descriptors, which are the plurality of local image descriptors derived from one image of the stereo pair. In this case, as Figure 6 shown, the displacement vector field 37 is derived as follows.

[0109] In step S4-1, a first displacement vector field 37 is derived, which represents the transformation of image patches in the one image 31 (which can be either image 31) of the stereo pair. This is done by using the derived feature vector 34 to look up the reference data 35.

[0110] In step S4-2, a displacement vector field 37 is derived, which represents the transformation of image patches in the other image 31 of the stereo pair. This is done by transforming the displacement vector field derived in step S4-1 according to an estimated value of the optical flow between the image patches in the images 31 of the stereo pair.

[0111] The optical flow represents the effect of different camera positions between the images 31 of the stereo pair. Such optical flow is itself known and can be estimated using known techniques, such as those disclosed in Zach et al., "A Duality Based Approach for Realtime TV-L1 Optical Flow", Pattern Recognition (Proc. DAGM), 2007, pp. 214-223 (Zach et al., "A Duality Based Approach for Realtime TV-L1 Optical Flow", Pattern Recognition (Proc. DAGM), 2007, pp. 214-223, which is incorporated herein by reference).

[0112] As an example, if the first displacement vector field 37 derived in step S4-1 is used for the left image L o, L i (where the subscripts o and i represent the output image and the input image, respectively), and the optical flow of the right image R o is represented by a displacement vector field G given by:

[0113] G = {s(x,y), t(x,y)}

[0114] then the second displacement vector field 37 can be derived according to the following formula:

[0115] R o (x,y) = L o (x + s(x,y), y + t(x,y)) = L i (x + s + u(x + s, y + t, y + t + v(x + s, y + t)

[0116] Therefore, in the third method, the displacement vector fields 37 derived for each image 31 of the stereo pair are consistent, because only one displacement vector field is derived from the reference data 35 and the other displacement vector field is derived from it using a transformation, and the other displacement vector field remains consistent because it is derived based on the estimated optical flow between the image patches in the images 31 of the stereo pair.

[0117] In a fourth possible method, a feature vector 34 is derived from a plurality of local image descriptors, which are a plurality of local image descriptors derived from the two images of the stereo pair. In this case, as Figure 7 shown, the displacement vector field 37 is derived as follows.

[0118] In step S4-3, an initial displacement vector field is derived, which represents the conceptual transformation of conceptual image patches in the conceptual image. The conceptual image has a conceptual camera position at a predetermined position relative to the camera position of image 31 (in the middle between the camera positions of image 31 in this example). This can be regarded as a central eye. This is done by using the derived feature vector 34 to look up reference data 35, which includes reference displacement vector fields associated with possible values of the feature vector. This means that the reference data 35 is structured accordingly, but can still be derived from the training input including the single field-of-view images.

[0119] In step S4-4, a displacement vector field 37 is derived, which represents the transformation of image patches in each of the images 31 of the stereo pair. This is done by transforming the initial displacement vector field derived in step S4-3 according to the estimated value of the optical flow between the conceptual image patches in the conceptual image and the image patches in the images 31 of the stereo pair.

[0120] The optical flow represents the effect of different camera positions between the conceptual image and the images 31 of the stereo pair. This optical flow itself is known and can be estimated using known techniques, such as those disclosed in Zach et al., "A Duality Based Approach for Realtime TV-L1 Optical Flow", Pattern Recognition (Proc. DAGM), 2007, pp. 214-223 (Zach et al., "A Duality Based Approach for Realtime TV-L1 Optical Flow", Pattern Recognition (Proc. DAGM), 2007, pp. 214-223, which is cited above and incorporated herein by reference).

[0121] As an example, if the optical flow from the left image L to the right image R is represented by a displacement vector field G given by:

[0122] G = {s(x,y), t(x,y)}

[0123] then the transformation of the conceptual image C is given by:

[0124]

[0125] Thus, in this example, in step S4-4, the initial displacement vector field F derived for the conceptual image C in step S4-3 is transformed to derive the flow fields F rc and F lc :

[0126]

[0127]

[0128] Thus, in the fourth method, the displacement vector fields 37 derived for each image 31 of the stereo pair are consistent because only one displacement vector field is derived from the reference data 35, which represents the conceptual transformation of the conceptual image blocks in the conceptual image, and the displacement vector fields for the left and right images are derived therefrom using the transformation, and the displacement vector fields are consistent because they are derived based on the estimated value of the optical flow between the image blocks in the conceptual image and the images 31 of the stereo pair.

[0129] In step S5, each image 31 of the stereo pair is adjusted by transforming the image blocks including the left and right eyes with the head according to the derived displacement vector field 37. This results in Figure 4 the adjusted stereo image pair 38 as shown, in which the fixation has been corrected. Specifically, this adjustment can be performed in two alternative ways as follows.

[0130] The first method for performing step S5 is shown in Figure 8 and is performed as follows.

[0131] In step S5-1, the image blocks are transformed in the normalized coordinate system according to the corresponding displacement vector field 37 for the same image, thereby correcting the fixation. As described above, for the displacement vector field F, the transformation of the image block of the input image I(x, y) provides the output image O(x, y) according to the following formula:

[0132] O(x, y) = I(x + u(x, y), y + v(x, y))

[0133] In step S5-2, the transformed image blocks output from step S5-1 are transformed from the normalized coordinate system back to the original coordinate system of the corresponding image 31. This is performed using the inverse transformation of the transformation applied in step S2-2.

[0134] In step S5-3, the transformed image blocks output from step S5-2 are superimposed on the corresponding image 31. This can be performed using complete replacement within the eye region corresponding to the eyes themselves and smooth transition within the boundary region around the eye region between the transformed image blocks and the original image 31. The width of the boundary region can be a fixed size or a certain percentage of the size of the image block in the original image 31.

[0135] The second method for performing step S5 is shown in Figure 9 and is performed as follows.

[0136] In this second alternative method, the transformation back to the coordinate system of the corresponding image 31 is performed before transforming the image blocks according to the transformed displacement vector field F.

[0137] In step S5-4, the displacement vector field F is transformed from the normalized coordinate system back to the original coordinate system corresponding to the image 31. This is done using the inverse transformation of the transformation applied in step S2-2.

[0138] In step S5-5, the image patch 32 in the coordinate system of the image 31 is transformed according to the displacement vector field F that has been transformed to the same coordinate system in step S5-4. As described above, for the displacement vector field F, the transformation of the image patch of the input image I(x,y) provides the output image O(x,y) according to the following formula:

[0139] O(x,y) = I(x + u(x,y), y + v(x,y))

[0140] However, at this time, this transformation is performed in the coordinate system of the original image 31.

[0141] Step S5-6 is the same as S5-3. Therefore, in step S5-6, the transformed image patch output from step S5-5 is superimposed on the corresponding image 31. This can be done using a complete replacement within the eye region corresponding to the eye itself and a smooth transition within the boundary region around the eye region between the transformed image patch and the original image 31. The width of the boundary region can be a fixed size or a certain percentage of the size of the image patch in the original image 31.

[0142] Now, the displacement vector field 37 used in step S5 will be discussed.

[0143] One option is to directly use the displacement vector field 37 derived for the left and right images in step S4 in step S5. That is, the image patch of each image 31 of the stereo patch is transformed according to the displacement vector field 37 for that image 31. This is appropriate when the displacement vector field 37 is accurate enough, for example, because they have been derived from the reference data 35, and the reference data itself is derived from the stereo images according to the second method described above.

[0144] An alternative according to the fifth method is to derive and use a combined displacement vector field 39. This can be applied in combination with any of the first to fourth methods described above. In this case, step S5 additionally includes step S5-a as shown Figure 10 which is performed before step S5-1 in the first method as shown in Figure 8 or before step S5-4 in the second method as shown in Figure 9 In step S5-a, the combined displacement vector field 39 can be derived from the displacement vector fields 37 derived for the image patches in each image 31 of the stereo pair in step S4.

[0145] Then, the remainder of step S5 is performed using the combined displacement vector field 39 for each image 31. That is, in Figure 8 the first method of Figure 8 , in step S5-1, the image blocks 33 for each image 31 of the stereo pair are transformed according to the combined displacement vector field 39. Similarly, in Figure 9 the second method of Figure 9 , in step S5-4, the combined displacement vector field 39 is transformed, and in step S5-5, the image blocks 33 for each image 31 of the stereo pair are transformed according to the combined displacement vector field 39.

[0146] In this case, the displacement vector fields for each image are consistent because they are the same.

[0147] The combination in step S5-1a can be performed in any suitable manner.

[0148] In one example, the combination in step S5-1a can be a simple average of the displacement vector fields 37 derived in step S4

[0149] In another example, the combination in step S5-1a can be an average weighted by the confidence values associated with each of the derived displacement vector fields 37. In this case, the confidence values form part of the reference data 35 in the above-described manner, and the confidence values are derived from the reference data 35 and the derived displacement vector fields 37 in step S4.

[0150] As an example, if the derived displacement vector field 37 is represented as F i , the combined displacement vector field 39 is represented as F avg , and the confidence value is represented as a i , then the combined displacement vector field 39 can be derived as follows:

[0151]

[0152] In the above example, the fixation of the target observer 24 in the observer position, which is the normal viewing position perpendicular to the center of the display position, is corrected, as shown by the solid line contour of the target observer 24 for the target device 20 of Figure 2 . This is sufficient in many cases. However, optional modifications will now be described that allow the fixation of the target observer 24 in different observer positions to be corrected, as shown by the dashed line contour of the target observer 24 for the target device 20 of Figure 2 .

[0153] In this case, the method further includes using position data 40 representing the observer position relative to the display position of the stereoscopic image pair 31. The position data 40 can be derived in the target device 20, for example, as described below. In this case, if the method is not executed in the target device 20, the position data 40 is transmitted to the device that executes the method.

[0154] The relative observer position can take into account the position of the observer relative to the display 21. This can be determined by using the camera system in the target device 20 and an appropriate head tracking module to detect the position of the target observer 24.

[0155] The relative observer position can assume that the image is displayed at the center of the display 21. Alternatively, the relative observer position can take into account the position of the observer relative to the display 21 and the position of the image displayed on the display 21. In this case, the position of the image displayed on the display 21 can be derived from the display geometry (e.g., the position and area of the display window and the size of the display 21).

[0156] To account for different observer positions, the reference data 34 includes multiple sets of reference displacement vector fields, each set associated with a different observer position. This is achieved through the training input of a machine learning technique that includes multiple second sets of output images, each second set being an image of each individual with the correct fixation of a corresponding predetermined observer position relative to the display position of the image to be displayed. Thus, in step S4, the displacement vector field 37 is derived by looking up the set of reference displacement vector fields associated with the observer position represented by the position data.

[0157] As described above, the method can be implemented in the image processor 30 provided in various different devices. As a non-limiting example, a specific implementation in a telecommunications system will now be described, which is shown in Figure 11 and arranged as follows.

[0158] In this implementation, the source device 10 and the target device 10 communicate through such a telecommunications network 50. To communicate through the telecommunications network 50, the source device 10 includes a telecommunications interface 17 and the target device 20 includes a telecommunications interface 27.

[0159] In this implementation, an image processor 30 is provided in the source device 10 and a stereoscopic image pair directly from the camera system 12 is provided for the image processor. The telecommunications interface 17 is arranged to transmit the adjusted stereoscopic image pair 38 to the target device 20 through the telecommunications network 50 for display thereon.

[0160] The target device 20 includes an image display module 28 that controls the display 26. The adjusted stereoscopic image pair 38 is received into the target device 20 via the telecommunication interface 27 and provided to the image display module 28, which causes the adjusted stereoscopic image pair to be displayed on the display 26.

[0161] In the case where the method corrects the gaze of a target observer 24 at an observer position other than the normal viewing position perpendicular to the center of the display position, the following elements of the target device 20 are optionally included. In this case, the target device 20 includes a camera system 23 and an observer position module 29. The camera system 23 captures an image of the target observer 24. The observer position module 29 derives position data 40. The observer position module 29 includes a head tracking module that uses the output of the camera system 23 to detect the position of the target observer 24. The observer position module 29. In the case where the position of the image displayed on the display 21 is also considered in the relative observer position, the observer position module 29 obtains the position of the image displayed on the display 21 from the image display module 28. The telecommunication interface 17 is arranged to transmit the position data 40 via a telecommunication network 50 to the source device 10 for its use.

[0162] Although the above description relates to a method applied to images provided from the source device 10 to the target device 20, the method is equally applicable to images provided from the target device 20 to the source device 10 in the opposite direction, in which case the target device 20 effectively becomes the "source device" and the source device 10 effectively becomes the "target device". In the case of providing images bidirectionally, the labels "source" and "target" can be applied to both devices, depending on the communication direction to be considered.

[0163] Although various embodiments in accordance with the principles disclosed herein have been described above, it should be understood that these embodiments are shown by way of example and not limitation. Accordingly, the breadth and scope of the present disclosure should not be limited by any of the above-described exemplary embodiments, but should be defined only in accordance with any issued claims of the present disclosure and their equivalents. Additionally, the above-described advantages and feature structures are provided in the described embodiments, but the application of the issued claims should not be limited to methods and structures that achieve any or all of the above-described advantages.

[0164] Additionally, the section headings of this document are provided to conform to the recommendations of 37 CFR 1.77 or for organizational cues. These headings should not limit or characterize the embodiments listed in any claims that may issue from this disclosure. Specifically and by way of example, although the heading refers to the "Technical Field," the claims should not be limited by the language chosen under that heading to describe the so-called field. Additionally, the description of the technology in the "Background Art" should not be construed as an admission that certain technology is prior art to any embodiment in this disclosure. The "Summary of the Invention" is also not to be regarded as a characterization of the embodiments described in the issued claims. Furthermore, any reference to the singular form of "the invention" in this disclosure should not be used to argue that there is only one novel point in this disclosure. Multiple embodiments may be described in terms of the limitations of multiple claims arising from this disclosure, and such claims thus define the embodiments protected by them and their equivalents. In all cases, the scope should be considered based on the characteristics of the claims themselves of this disclosure, and should not be constrained by the headings given herein.

Claims

1. A method for adjusting multi-view images of a head to correct fixation, the method comprising: Capturing multi-view images with an electronic device having a camera system, the camera system being located at a position offset from a display of the electronic device; In each image of the multi-view images, identifying image patches that include the left and right eyes of the head; For the image patches that include the left eye of the head in each image of the multi-view images, and also for the image patches that include the right eye of the head in each image of the multi-view images, perform the following steps: Deriving a feature vector from a plurality of local image descriptors of the image patches in at least one image of the multi-view images; And Using the derived feature vector to look up reference data including a reference displacement vector field associated with possible values of the feature vector, thereby deriving a displacement vector field representing the transformation of the image patch; Adjusting each image of the multi-view images by transforming the image patches that include the left and right eyes of the head according to the derived displacement vector field; And Presenting the adjusted images on the display of the electronic device.

2. The method according to claim 1, wherein The plurality of local image descriptors are a plurality of local image descriptors in each image of the multi-view images, and The reference data includes a pair of reference displacement vector fields for each image of the multi-view images, the pair of reference displacement vector fields being associated with possible values of the feature vector.

3. The method according to claim 1, wherein The plurality of local image descriptors are a plurality of local image descriptors in one image of the multi-view images, The step of deriving the displacement vector field includes: Using the derived feature vector to look up reference data including a reference displacement vector field associated with possible values of the feature vector, thereby deriving a displacement vector field representing the transformation of the image patch in the one image of the multi-view images; and And Deriving a displacement vector field representing the transformation of the image patch in one or more other multi-view images by transforming the derived displacement vector field representing the transformation of the image patch in the one image of the multi-view images according to an estimated value of the optical flow between the image patch in the one image of the multi-view images and the image patch in one or more other multi-view images.

4. The method according to claim 1, wherein the multi-view images are a stereo image pair, The plurality of local image descriptors are a plurality of local image descriptors in two images of the stereo image pair, The step of deriving the displacement vector field includes: Using the derived feature vector to look up reference data including a reference displacement vector field associated with possible values of the feature vector, thereby deriving an initial displacement vector field, the initial displacement vector field representing a conceptual transformation of a conceptual image patch in a conceptual image, the conceptual image having a conceptual camera position relative to the camera positions of the multi-view images; and A displacement vector field representing the transformation of the image blocks in each image of the multi-view image is derived by transforming the initial displacement vector field according to an estimated value of the optical flow between the conceptual image blocks in the conceptual image and the image blocks in the images of the multi-view image.

5. The method according to claim 4, wherein the multi-view image is a stereo image pair, and the conceptual camera position is between the camera positions of the images of the stereo image pair.

6. The method according to claim 1, wherein the step of deriving the displacement vector field includes deriving a displacement vector field for the image blocks in each image of the multi-view image, and the step of transforming the image blocks of the left eye and the right eye including the head is performed according to the displacement vector field derived for the image blocks in each image of the multi-view image.

7. The method according to claim 1, wherein the step of deriving the displacement vector field includes deriving a displacement vector field for the image blocks in each image of the multi-view image, and further deriving a combined displacement vector field from the displacement vector fields derived for the image blocks in each image of the multi-view image, and the step of transforming the image blocks of the left eye and the right eye including the head is performed according to the combined displacement vector field.

8. The method according to claim 7, wherein the reference displacement vector field is further associated with a confidence value, the step of deriving a displacement vector field for the image blocks in each image of the multi-view image further includes deriving a confidence value associated with each derived displacement vector field, and the combined displacement vector field is an average of the displacement vector fields derived for the image blocks in each image of the multi-view image weighted by the derived confidence values.

9. The method according to claim 1, wherein the method uses position data representing the observer position relative to the display position of the multi-view image, the reference data includes multiple sets of reference displacement vector fields associated with possible values of the feature vector, the sets being associated with different observer positions, and the step of deriving a displacement vector field representing the transformation of the image blocks is performed by using the derived feature vector to look up the set of reference displacement vector fields associated with the observer position represented by the position data.

10. The method according to claim 1, wherein the local image descriptor is a local image descriptor derived in a normalized coordinate system, and the reference displacement vector field and the derived displacement vector field are displacement vector fields in the same normalized coordinate system.

11. The method according to claim 1, wherein the feature vector includes features that are values derived from the local image descriptor in a discrete space, the reference displacement vector field includes reference displacement vector fields associated with each possible value of the feature vector in the discrete space, and The step of deriving the displacement vector field for the image patch includes selecting the reference displacement vector field associated with the actual value of the derived feature vector.

12. The method according to claim 1, wherein the feature vector includes features that are discrete values of the local image descriptor in a continuous space, and the step of deriving the displacement vector field for the image patch includes interpolating the displacement vector field from the reference displacement vector field based on the relationship between the actual value of the derived feature vector and the value of the feature vector associated with the reference displacement vector field.

13. The method according to claim 1, wherein the multi-view image is a stereo image pair.

14. A computer program product, which can be executed by a processor and is arranged to execute, when executed, or cause the processor to execute the method according to any one of the preceding claims.

15. A computer-readable storage medium storing the computer program product according to claim 14.

16. An apparatus for adjusting multi-viewpoint images of a head to correct fixation, the apparatus comprising: A display; A camera system configured to capture multi-view images of a head, the camera system being located at a position offset from the display; An image processor arranged to process the multi-view images of the head by: In each image of the multi-view images, respectively identifying image patches containing the left and right eyes of the head; For the image patch containing the left eye of the head in each image of the multi-view images, and also for the image patch containing the right eye of the head in each image of the multi-view images, performing the following steps: Deriving a feature vector from a plurality of local image descriptors of the image patch in at least one image of the multi-view images; And Using the derived feature vector to look up reference data including a reference displacement vector field associated with possible values of the feature vector, thereby deriving a displacement vector field representing the transformation of the image patch; Adjusting each image of the multi-view images by transforming the image patches containing the left and right eyes of the head according to the derived displacement vector field; And Presenting the adjusted images on the display.

17. The apparatus according to claim 16, further comprising a telecommunications interface arranged to transmit the adjusted images to a target device via a telecommunications network for display thereon.

18. The apparatus according to claim 16, wherein the multi-view image is a stereo image pair.

Citation Information

Patent Citations

  • Display device, image processing device, and image processing method, as well as computer program

    CN103534748A

  • Eye tracking system

    GB0119859D0