Image synthesis method, apparatus, computer program and computer-readable storage medium
Patent Information
- Application Number
- EP2023801529
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-10-31
- Publication Date
- 2025-07-02
AI Technical Summary
Existing augmented reality applications are limited by the restricted field of view of real-world information and the inability to seamlessly integrate virtual objects with real-world interactions.
The method involves receiving and synchronizing image sequences from two real cameras, isolating real objects from their backgrounds, and rendering a composite augmented reality model in real time using a virtual camera, thereby combining real and virtual elements in a coherent and interactive manner.
This approach enhances the information content and interaction possibilities in augmented environments by seamlessly integrating real and virtual elements, providing a more flexible, efficient, and user-friendly experience without the need for additional processing.
Smart Images

Figure IB2023060954_08052025_PF_FP_ABST
Abstract
Description
[0001] Image synthesis method, apparatus, computer program and computer-readable storage medium
[0002] The invention relates to an image synthesis method, an image synthesis appa- ratus, a computer program and a computer readable storage medium for synthesizing images by means of a virtual camera.
[0003] Virtual reality (VR) refers to a computer-generated environment that is typically displayed on a dedicated user interface such as a head-mounted display, thereby giving a user a 360 degree view of an entirely virtual world shielded from the surrounding real world environment.
[0004] On the contrary, augmented reality (AR) is a special case of an extended reality environment that has virtual objects, i.e., computer-generated infor- mation, superimposed with the real world in the field of view of a user.
[0005] Thereby, an augmented reality application typically attempts to combine real and virtual elements by projecting virtual objects into a real world image captured with a real camera.
[0006] However, in such a case the information provided by the real world is limited to the field of view of the user. Moreover, virtual objects are merely projected into the real camera image, thereby restricting the interaction possibilities of the user with the virtual world.
[0007] Thus, an object of the invention is to overcome such limitations and provide an image synthesis method, apparatus, computer program and computer readable storage medium that improve upon the information content and interaction possibilities provided by an augmented environment.
[0008] This object of the invention is achieved by the image synthesis method, image synthesis apparatus, computer program and computer readable storage medium as described in claims 1, 13, 14 and 15. Advantageous developments and embodiments are described in the dependent claims.
[0009] The invention relates to an image synthesis method. The image synthesis method comprises the receiving of a first image sequence, wherein the first image sequence is captured by means of a first real camera and the receiving of a second image sequence, wherein the second image sequence is captured by means of a second real camera. The second image sequence is captured time-synchronously with the first image sequence.
[0010] The image synthesis method further comprises performing an initialization, wherein the initialization comprises estimating an initial pose of the first real camera based on at least a first image of the first image sequence and aligning a virtual three-dimensional model based on the estimated initial pose of the first real camera.
[0011] While receiving and / or capturing further images of the first and second image sequence after the initialization, the method further comprises: time-continu- ously isolating a real object in the further images of the second image sequence from its background in the further images of the second image sequence; and time-continuously rendering a composite augmented reality model by means of a virtual camera and thereby forming a third image sequence in real time and time-synchronously with the further images of the first and second image sequence, wherein the composite augmented reality model is formed by compositing the isolated real object from the second image sequence together with the virtual three-dimensional model aligned based on the estimated initial pose of the first real camera during initialization.
[0012] With the proposed image synthesis method information from two real and time-synchronous image sequences captured from two different real-world cameras can be combined with a three-dimensional virtual model in a single synthesized image sequence in a seamless and coherent manner, thereby extending the information space that can be provided to a user in a particular flexible, efficient and user-friendly manner without the need for additional pre-processing or post processing. Thereby, the first real image sequence provides information for the alignment of the virtual three-dimensional model while the second real image sequence provides the real (physical) object that is composited with the virtual three-dimensional model.
[0013] Preferably, there is substantially no time delay between the capturing of an image of the first and second real cameras and the receiving of an image of the first and second real cameras. The steps of initialization can be carried out while receiving and / or capturing the at least first image, i.e. at least a first or a first few images, of the first image sequence. The subsequent steps of isolating the real object and the rendering of the composite augmented reality model can be carried out based on (and while) receiving and / or capturing the further images of the first and second image sequence after initialization. The further images are received and captured after the at least first image is received and captured.
[0014] Preferably, there is also no substantial time delay between the steps of the method corresponding to the isolating of the real object in the second image sequence and the compositing and the rendering of the composite augmented reality model. The steps of compositing or overlaying the aligned virtual three-dimensional model and the isolated real object and the rendering of the composite augmented reality model may be performed simultaneously by the virtual camera. Thus, these steps of the method can be considered to be performed quasi-simultaneously and time-synchronously such that the image synthesis model can be performed in real time while receiving and / or capturing the first and second image sequence and / or the further images of the first and second image sequence. The receiving and the capturing of an image sequence may refer to the same method step. The first, second and third image sequence may each correspond to a video-stream.
[0015] Preferably, the pose of the first and second real cameras are different and fixed relative to each other during the capturing of the first and second image sequence, during the initialization and / or during the time-continuously rendering the composite augmented reality model. Optionally, the optical axes of the first and second real cameras may be parallel or identical to each other. The viewing / capturing directions of the first and second real camera may be directed in opposite directions.
[0016] The first and second real camera may be digital cameras. Advantageously, the first real and second real camera may be connected and / or fixed with respect to a same support structure. The support structure may define a plane perpendicular to the optical axes of the first and second real cameras. For example, the first real camera, the second real camera and / or the support structure may be part of a mobile device such as a handheld mobile device, e.g., a smartphone, a tablet or a laptop. The mobile device may also be a headmounted mobile device, e.g., a head mounted display or glasses. The first real camera may correspond to a back camera of the mobile device and the second real camera may correspond to a front camera of the mobile device. The support structure and / or the mobile device may further comprise a computing unit that carries out the steps of the image synthesis method.
[0017] Keeping the relative pose of the first and second real camera fixed during the receiving and / or capturing of the first and second image sequence, the initialization and the rendering allows for a particular efficient compositing and forming of the composite augmented reality model. For example, the pose of the second real camera does need to be estimated separately, but is determined as soon as the pose of the first real camera is estimated and determined, which considerably simplifies and accelerates the image synthesis method. The pose of the first real camera, the second real camera and / or the virtual camera may correspond to a position and / or an orientation, e.g., of the respective camera in a three-dimensional coordinate system. A three-dimensional coordinate system may be a (fixed) world coordinate system. The origin of a world coordinate system may coincide with a physical real world object. World coordinates may correspond to geographic, geodetic, cartesian and / or GPS coordinates.
[0018] A world coordinate system may stay fixed and may not change during the capturing and / or receiving of the first and second image sequence. If not explicitly stated otherwise, a pose according to the present application may be a pose in a world coordinate system.
[0019] In a camera coordinate system, the coordinates are defined relative to the optical center of the camera. The origin of the (local) first real camera coordinate system corresponds to the position of the first real camera, e.g., the optical center of the first real camera. The origin of the (local) second real camera coordinate system corresponds to the positions of the second real camera, e.g., the optical center of the first second camera. The origin of the (local) virtual camera coordinate system corresponds to the position of the virtual camera, e.g., the optical center of the virtual camera.
[0020] The world coordinate system may be identical to the (initial) first real camera coordinate system defined by the initial pose of the first real camera estimated during initialization.
[0021] The pose may be determined through a set of six parameters, i.e., three position coordinates and three orientation angles. The pose of the first real camera and / or the virtual camera may be a six degrees of freedom (6DOF) pose.
[0022] The virtual three-dimensional model is computer-generated and may be prestored, e.g., in a memory unit of the first real camera, the second real camera, the support structure or the mobile device or an external server. The virtual three-dimensional model may represent a 360-degree view of a virtual world or at least an angular section of it. The virtual three-dimensional model may comprise or may be identical to one or a plurality of virtual objects. During initialization, the aligning of the virtual three-dimensional model based on the estimated initial pose of the first real camera defines an initial pose of the virtual three-dimensional model, e.g., in a world coordinate system. For example, the distance and / or orientational relationship of the virtual three- dimensional model with respect to the initial pose of the first real camera may be pre-defined. The estimate of the initial pose of the first real camera, then also defines and sets the initial pose of the virtual three-dimensional model. The pose of the virtual three-dimensional model, e.g., in a world coordinate system, may stay fixed also after initialization and correspond to the initial pose of the virtual three-dimensional model also during the rendering of the composite augmented reality model based on further images of the first and second image sequence received and / or captured after the initialization and during the rendering.
[0023] The real object may also be referred to as a physical object in the real world as captured with the second real camera. For example, the time-continuously isolating the real object in the second image sequence or further images of the second image sequence from its background in the second image sequence or further images of the second image sequence may comprise isolating the real object by means of chromakeying technology. This may comprise placing a greenscreen or bluescreen behind the real object. The greenscreen or bluescreen may represent the background in the second real image sequence.
[0024] The use of chromakeying, greenscreen and / or bluescreen technology allows for a particular simple, accurate and efficient isolation of the real object as no extensive or complex computational resources, pre-processing (e.g. training) and / or image post processing are / is required. In particular, the use of chroma-keying, greenscreen and / or bluescreen technology allows for an efficient on-device and / or on-chip image synthesizing (e.g., by the computing unit of the support structure or the mobile device) without lagging in the third image sequence and other detrimental effects that may occur due to the necessity of extensive computations and / or the need to first transmit the first and second image sequences or parts thereof to a remote server via a mobile connection or long cable-based connections. Additionally, or alternatively, the real object may be isolated from its background by using a machine learning algorithm. This would improve upon the freedom of movement as no greenscreen or bluescreen need to be provided and captured by the second real camera.
[0025] Optionally, the isolated real object is a two-dimensional real object. For example, the isolated real object may be represented in a two-dimensional plane. The two-dimensional plane may be a normal plane with respect to the optical axis of the virtual camera.
[0026] The isolated real object may stay aligned in a fixed positional and / or orientational relationship with respect to the pose of the virtual camera during the time-continuously rendering the composite augmented reality model. Therefore, the pose of the isolated real object in the local virtual camera coordinate system can remain the same when the virtual camera is moved. In this case, the viewing angle of the virtual camera at the isolated real object stays the same during the rendering.
[0027] For example, the isolated real object may be rendered by the virtual camera in front or behind the aligned three-dimensional virtual model or in front or behind one or a plurality of virtual objects of the three-dimensional virtual model as seen from the virtual camera.
[0028] Alternatively, the isolated real object plane may also be positioned with a predefined distance and / or predefined orientation with respect to the virtual camera and / or the optical axis of the virtual camera. The isolated real object may also be positioned into the aligned three-dimensional virtual object such that at least parts of the aligned three-dimensional virtual object may appear in front of the isolated real object in the third image sequence.
[0029] Alternatively, the isolated real object may also be kept in a fixed positional and / or orientational relationship with respect to the virtual three-dimensional model and / or the initial pose of the first real camera during the receiving and / or capturing of the first and second image sequence and the rendering of the composite augmented reality model. In this case, when the first real camera and / or the virtual camera moves, the pose of the isolated real object may stay the same, but the viewing angle of the virtual camera at the isolated real object in the composite augmented reality model changes. During the time-continuously rendering the composite augmented reality model the pose of the virtual camera may be different from the estimated initial pose of the first real camera. Additionally, or alternatively, during the rendering the pose of the virtual camera may also be different from the current pose of the first real camera at a point of time after initialization. This may enable a different viewing angle at the virtual three-dimensional model in the composite augmented reality model of the third image sequence.
[0030] For example, the pose of the virtual camera may correspond to a translational and / or rotational virtual camera offset with respect to the first real camera. A virtual camera offset may be defined with respect to the initial pose of the first real camera. The translational and / or rotational virtual camera offset may be pre-defined or may be entered during the initialization and / or the rendering by a user by means of a user interface.
[0031] The translational and / or rotational virtual camera offset may also change during the receiving and / or capturing of the first and second image sequence and / or the rendering such that the viewing angle at the composite augmented reality model may changes as well even if the first and / or second real cameras are not moved during the rendering.
[0032] During the time-continuously rendering the composite augmented reality model the virtual camera may perform a virtual camera movement / drive along path nodes of a continuous or discrete space-time path while receiving and / or capturing the first and second image sequence, the further images of the first and second image sequence, and / or during the rendering, wherein each path node of the space-time path may correspond to a different translational and / or rotational virtual camera offset, e.g., with respect to the initial pose of the first real camera estimated during initialization.
[0033] A translational virtual camera offset may correspond to a difference between the position of the origin of a (local) virtual camera coordinate system and the initial position of the first real camera, i.e., the origin of the first real camera coordinate system during initialization.
[0034] The rotational virtual camera offset may correspond to a difference in orientation between corresponding axes of the (local) virtual camera coordinate system and the axes of the first real camera coordinate system during initialization.
[0035] The time component of the space time path may correspond to a velocity of the camera drive and may be defined individually for each segment of the space-time path, i.e., the velocity may be different between two path nodes. A camera drive between two path nodes with a finite velocity may be achieved through interpolation between two path nodes.
[0036] The velocity between two path nodes may also be infinite in which case there is an instantaneous switching or snapping from one path node to another path node, e.g., triggered by an input of the user via a user interface.
[0037] The path nodes, the space-time path and / or the corresponding virtual camera offsets may be pre-defined or may be entered by a user during the receiving and / or capturing of the first and second image sequence by means of a user interface. The virtual camera movement from one path node to another may also be triggered by an interaction of the user with a user interface of the mobile device.
[0038] Optionally, the first and / or second real camera maybe be rotationally and / or translationally moved during the receiving and / or capturing of the first and second image sequence after initialization or the further images of the first and second image sequence and / or during the rendering the composite augmented reality model.
[0039] After the initialization and during the time-continuously rendering the composite augmented reality model, the image synthesis method may further comprise also time-continuously estimating the pose of the first real camera and time-continuously and instantaneously adapting the pose of the virtual camera to the estimated pose of the first real camera. Consequently, by moving the first real camera after initialization and during the rendering, the pose of the virtual camera may be adapted / aligned instantaneously with respect to the pose of the first real camera. In this way, the user can move through the virtual three-dimensional model and / or change the viewing angle at the virtual three-dimensional model by moving the first real camera during the rendering. The virtual camera may adopt to the motion of the first real camera in real time. Preferably, when the first real camera is moved by a translation and / or rotation during the receiving and / or capturing of the first and second image sequence or the further images of the first and second image sequence the virtual camera performs a corresponding translation and / or rotation as seen from the current local virtual camera coordinate system, e.g., as defined by the current virtual camera offset or current path node, and adapts its pose instantaneously.
[0040] When the first real camera is moved by a translation along a first distance, the virtual camera may perform a corresponding translation by the same first distance. The translation can be performed from the virtual camera offset. For example, when the translation of the first real camera along a first distance is in a direction of a certain type of axis (x, y or z) of the current first real camera coordinate system, the virtual camera may perform a corresponding translation along the same first distance in a direction of the same type of axis of the current virtual camera coordinate system.
[0041] Additionally, and / or alternatively, when the first real camera is moved by a rotation along a first angle, the virtual camera may perform a corresponding rotation by the same first angle. The rotation can be performed from the current virtual camera offset. For example, when the rotation of the first real camera along a first angle is around a certain type of axis (x, y or z) of the current first real camera coordinate system, the virtual camera may perform a corresponding rotation along the same first angle around the same type of axis of the current virtual camera coordinate system.
[0042] The time-continuously estimating the pose of the first real camera to time- continuously and instantaneously adapt the pose of the virtual camera can be based on the tracking of feature points in the first image sequence or the further images of the first image sequence and / or can be based on the capturing of translational and / or rotational motion data from at least one motion sensor characterizing the motion of the first real camera.
[0043] The estimating the initial pose of the first real camera during initialization to align the virtual three-dimensional model can be based on tracking of a marker in the first image sequence or the at least first image of the first image sequence and / or be based on the capturing of translational and / or rotational motion data from at least one motion sensor characterizing the motion of the first real camera.
[0044] While receiving and / or capturing the first and second image sequence, the image synthesis method may also further comprise time-continuously tracking of a marker and / or feature points in the first image sequence to obtain tracking data and estimate the pose of the first real camera, e.g., in a world coordinate system, based on the tracking data.
[0045] Marker may represent pre-defined geometric structures that are externally placed in the real world at selected positions while feature points or feature point clouds may be identified based on intrinsic features of the real world such as surfaces, pins or holes of certain real-world objects. The motion of such marker and / or feature points or point clouds displayed in the first image sequence may then be tracked to provide information (i.e., tracking data) that allows to determine the pose of the first real camera based on the tracking.
[0046] It may be advantageous to use a combined marker and feature point tracking. In this case marker-based tracking may only be used during initialization for aligning the virtual three-dimensional model, e.g., based on the at least first or first few images of the first image sequence, while marker-free featurebased tracking may be used after initialization to adapt / align the pose of the virtual camera after initialization based on the further images of the first image sequence captured and / or received after initialization and during rendering. This has the advantage that the marker does not necessarily need to be captured by the first image sequence after the initialization.
[0047] Marker-based and feature-based tracking may also alternate such that marker-based tracking may additionally be performed after regular time intervals to stabilize the tracking and alignment of the virtual three-dimensional model after periods of feature-based tracking.
[0048] In addition to the tracking of a marker and / or feature points, the image synthesis method may also comprise time-continuously capturing translational and / or rotational motion data from at least one motion sensor characterizing the motion of the first real camera. The captured translational and / or rotational motion data can be used together with the tracking of a marker and / or feature points in the first image sequence to estimate the pose of the first real camera.
[0049] The pose of the first real camera may be a six degree of freedom pose (6 DOF pose). The estimation of a six degree of freedom pose of the first real camera can be achieved by marker and / or feature point tracking in combination with simultaneously capturing translational and / or rotational motion data of the at least one motion sensor.
[0050] The at least one motion sensors may be connected and / or spatially fixed with respect to the first real camera, the second real camera, the support structure and / or the mobile device.
[0051] As described by the various aspects of the image synthesis method, the virtual camera integrates real world information from two different real cameras and virtual or computer-generated information as represented by the virtual three-dimensional model into the third image sequence. It is noted that the real-world information from the first real camera is here used as an anchor for the virtual world during initialization, i.e., to align the virtual three-dimensional model with the initial pose of the first real camera. Further, the first image sequence can be used to estimate the pose of the first real camera also after initialization to be able to instantaneously adapt the pose of the virtual camera with respect to the pose of the first real camera after initialization and during the rendering.
[0052] Thereby, the first image sequence or the further images of the first image sequence itself do not need to be composited or overlaid with the virtual three- dimensional model and / or rendered into the third image sequence. The composite augmented reality model may thus not include real world objects from the first image sequence, but only from the second image sequence. In other words, the first image sequence may only be used for estimating the pose of the first real camera and may not be part of the composite augmented reality model and / or the third image sequence.
[0053] However, in some situations it may also be useful to additionally include the first image sequence or at least real (physical) world objects from the first image sequence or the further images of the first image sequence into the composite augmented reality model and the rendering of the third image sequence. In this case the third image sequence may also display real world objects or images captured with the first real camera.
[0054] The method may further comprise setting up an augmented reality camera as an augmented / virtual representation of the first real camera. The pose of the augmented reality camera may be kept fixed and identical to the pose of the first real camera while receiving and / or capturing the first and second image sequence. In other words, the local first real camera coordinate system and the local augmented reality camera coordinate system may be the same coordinate system during and / or after the initialization.
[0055] The virtual camera and the augmented reality camera represent softwarebased models of a camera. For example, the corresponding software-based models can be stored in a memory unit of the first real camera, the second real camera, the support structure and / or the mobile device or on an external server.
[0056] The software-based model may comprise a camera projection matrix. The camera projection matrix is the mapping of a point in the three-dimensional world coordinate system and a two-dimensional image point in the image plane of the respective camera. A camera projection matrix may comprise intrinsic parameters of the camera, e.g., configuration and / or calibration data. The intrinsic parameters may comprise values of at least one of a focal field parameter, an aperture parameter, a zoom parameter or exposure time parameter. The intrinsic parameters of the first real camera, the augmented reality camera and / or the virtual camera may be the same.
[0057] The camera projection matrix may also comprise extrinsic parameters such as rotation angles and translation vector components. During the receiving and / or capturing the first and second image sequence the method may also comprise determining the camera projection matrix of the first real camera and / or copying the camera projection matrix of the first real camera to the camera projection matrix of the augmented reality camera and / or the virtual camera. In this case, the camera projection matrix of the first real camera, the augmented reality camera and the virtual camera may be the same, e.g., during initialization and / or after the initialization during the rendering.
[0058] The camera projection matrix, the intrinsic parameters, the extrinsic parameters, the configuration data and / or the calibration data of the augmented reality camera, the first real camera and / or the virtual camera may be the same, e.g., during the initialization and / or after the initialization and during the rendering.
[0059] The simulating of the augmented reality camera, the virtual camera, the pose estimation of the first real camera and / or the aligning of the virtual three-dimensional model can be performed or at least supported by using augmented reality software, e.g., AR Foundation or AR core or similar.
[0060] The image synthesis method may additionally comprise displaying the third image sequence at a user interface substantially time-synchronously with the receiving and / or capturing of the first and second image sequence after initialization and the rendering. The user interface may be part of the mobile device. For example, the user interface may be a display of the mobile device. The display of the mobile device may be positioned in the plane perpendicular to the optical axes of the first and second real camera. The display may also be configured to display the first image sequence, the second image sequence and / or the third image sequence.
[0061] At least initially, i.e., during initialization, the pose of the virtual camera and / or the augmented reality camera may be identical to the pose of the first real camera. In this case, at least initially the local virtual camera coordinate system, the local augmented reality camera coordinate system and the local first real camera coordinate system may be the same.
[0062] At least initially, i.e., the capture / viewing direction of the first real camera and the augmented reality camera may be the same and opposite to the capture direction of the virtual camera. The capture direction of the virtual camera may at least initially coincide with the capture direction of the second real camera. For example, the capture direction of the first real camera and the augmented reality camera may initially correspond to the positive z-axis of the world coordinate system, while the capture direction of the virtual camera and the second real camera may initially correspond to the negative z-axis (i.e. rotated by 180 degrees).
[0063] In this way, the virtual camera may point in the direction of the second real camera and when the isolated real object corresponds to the user captured by the second real camera, e.g., a front camera of a mobile device, the user, by looking at the third image sequence, has the impression of capturing / creating an image sequence or movie / video stream through the second real camera, e.g., the front camera of the mobile device, while the tracking of marker and / or feature point cloud is still being achieved by using the first real camera, e.g., the back camera of the mobile device.
[0064] The image synthesis method may also comprise an additional step of mirroring the images of the third image sequence. The mirroring may be performed with respect to an axis or the camera plane of the virtual camera in each image of the third image sequence (i.e., the plane perpendicular to the normal vector of the virtual camera).
[0065] The invention may be applied in various fields of technology ranging from virtual and augmented based computer simulations to video conferencing and mobile communications, cinematography, medical imaging and gaming. The invention can also be applied in the fields of architecture, project development, for collaborations in research and teaching as well for applications in the area of sales and human ressources.
[0066] Preferably, the method is performed in a video conference application, wherein the isolated real object corresponds to a real person / moderator captured by the second real camera in the second image sequence or the further images of the second image sequence. The method may then further comprise displaying the third image sequence at a user interface substantially time-synchronously with the receiving and / or capturing of the first and second image sequence or the further images of the first and second image sequence.
[0067] For example, the invention may be used in a situation where the first and second real cameras are part of a handheld mobile device that is held by a real person or moderator as a user corresponding to the real object while receiving and capturing the first image sequence, wherein the first real camera is the back camera of the mobile device and the second real camera is the front camera of the mobile device facing the user. The virtual three-dimensional model may be a simulated environment such as a virtual space / room with virtual objects therein. The third image sequence may then be displayed on the display of the mobile device while capturing the first and second image sequence. With the help of the back camera, the real person or moderator can then freely move through the virtual three-dimensional world / model in real time. Thereby, it is not necessary to capture the marker time-continuously in the first image sequence. Instead, the marker can be used only during initialization and subsequent tracking during the rendering can be based on feature points.
[0068] The rendering of the composite augmented reality model with a virtual camera whose pose may be different from the pose of the first real camera allows for free virtual camera drives / movements. For example, the capture direction of the virtual camera may be aligned or at least substantially directed in the same direction as the capture direction of the second real camera and opposite to the capture direction of the first real camera. Then the real person or moderator, by viewing the third image sequence, has the impression of filming through the second real camera with a selfie-type experience. Such a mode of operation is intuitive for a user experienced with the handling of mobile devices.
[0069] The invention also relates to a computer program comprising instructions which, when the program is executed by a computer, cause the computer to carry out the image synthesis method or any combination of its embodiments as described above. The computer program (or a sequence of instructions) may use software means for performing the image synthesis method when the computer program runs in a computing unit. The computer program can be stored directly in an internal memory, a memory unit, or the computer.
[0070] The invention also relates to a computer-readable storage medium having stored there on the computer program described above. The computer program can be stored in machine-readable data carrier(s), preferably digital storage media.
[0071] The invention also relates to an image synthesis apparatus comprising means to carry out the steps of the image synthesis method or any combination of its embodiments described above. The image synthesis apparatus may further comprise the first real camera configured to capture the first image sequence and the second real camera configured to capture the second image sequence.
[0072] The image synthesis apparatus may comprise at least one computing unit and / or at least one electronic storage unit. The at least one computing unit may comprise at least one of processor, a CPU (central processing unit) or a GPU (graphical processing unit).
[0073] The at least one computing unit may be configured to activate and simulate the virtual camera and / or the augmented reality camera.
[0074] The at least one computing unit may be configured to receive the first and second image sequence and align the virtual three-dimensional model with the initial pose of the first real camera.
[0075] The at least one computing unit may also be configured to time-continuously isolate a real object in the second image sequence from its background in the second image sequence.
[0076] The at least one computing unit may also be configured to composite the aligned virtual three-dimensional model together with the isolated real object from the second image sequence to form the composite augmented reality model.
[0077] The at least one computing unit may also be configured to render the composite augmented reality model by means of a virtual camera and thereby form a third image sequence in real time and time-synchronously with the further images of the first and second image sequence.
[0078] The image synthesis apparatus can also comprise the support structure, the mobile device and / or the at least one motion sensor.
[0079] Exemplary embodiments of the invention are illustrated in the drawings and will now be described with reference to figures 1 to 7.
[0080] In the figures:
[0081] Fig. 1 shows an embodiment of the image synthesis apparatus, Fig. 2 shows an embodiment of the first, second and third image sequence,
[0082] Fig. 3 shows a schematic flow diagram of an embodiment of the image synthesis method,
[0083] Fig. 4 shows a virtual camera offset along a space-time path,
[0084] Fig. 5 shows an adaptation of the pose of the virtual camera when the first real camera is rotated,
[0085] Fig. 6 shows an adaptation of the pose of the virtual camera when the first real camera is translated,
[0086] Fig. 7 shows an adaptation of the pose of the virtual camera when the first real camera is rotated and translated.
[0087] Figure 1 shows an embodiment of the image synthesis apparatus. The image synthesis apparatus comprises the first real camera 1.2 and the second real camera 2.2. The first real camera 1.2 and the second real camera 2.2 are connected to a common support structure 5. The common support structure 5 is part of a handheld mobile device that corresponds to the image synthesis apparatus. The first real camera 1.2 corresponds to the back camera of the mobile device and the second real camera 2.2 corresponds to the front camera of the mobile device.
[0088] The image synthesis apparatus further comprises a computing unit configured to perform the steps of the image synthesis method. Figure 5 depicts the image synthesis apparatus with its pose during initialization SO (discussed further below). In particular, the computing unit is configured to set up and initialize the augmented reality camera 1.2.1 and the virtual camera 3.2 in the world coordinate system 6 shown in Figure 1 with z-axis and x-axis. Initially, the capture direction of the first real camera 1.2 and the augmented reality camera 1.2.1 is in the positive z-direction and the capture direction of the virtual camera 3.2 is in the negative z-direction (the virtual camera 3.2 is rotated clockwise by 180° around the y-axis with respect to the first real camera 1.2 and the augmented reality camera 1.2.1) such that at least initially the virtual camera 3.2 is directed in the same direction as the second real camera 2.2 and the augmented reality camera 1.2.1 is directed in the same direction as the first real camera 1.2.
[0089] The convention is that initially, i.e., during initialization SO, as shown in Figure 1 the world coordinate system 6, the first real camera coordinate system, the augmented camera coordinate system and the virtual camera coordinate system are the same.
[0090] The second real camera coordinate system is rotated clockwise by 180° around the y-axis (out of the x-z plane and normal to the x-z plane; not shown) of the world coordinate system 6.
[0091] Thus, initially, the capture direction of the first real camera 1.2 is in the positive z-direction of the first real camera coordinate system, the capture direction of the second real camera 2.2 is in the positive z-direction of the second real camera coordinate system, the capture direction of the augmented reality camera 1.2.1 is in the positive z-direction of the augmented reality camera coordinate system, but the capture direction of the virtual camera 3.2 is in the negative z-direction of the virtual camera coordinate system.
[0092] The virtual three-dimensional model 3.3 is kept fixed and aligned, i.e., in a fixed and predetermined positional and orientational relationship, with the initial pose of the first real camera 1.2.
[0093] The isolated real object 2.3 is two-dimensional and is represented in a two-dimensional plane normal to the optical axis of the virtual camera 3.2 with a predetermined distance from the optical centre of the virtual camera 3.2 chosen such that the isolated real object 2.3 appears in front of the virtual three- dimensional model 3.3 as seen from the virtual camera 3.2. The virtual camera 3.2 is then configured to composite and render the isolated real object 2.3 in front of the three-dimensional virtual model 3.3.
[0094] Recurring features are provided in the following figures with identical reference signs as in Figure 1. Figure 2 shows an embodiment of at least a part of the further images 1.1.1, 1.1.2 of the first image sequence 1.1 captured by the first real camera 1.2, at least a part of the further images of the second image sequence 2.1 captured by the second real camera 2.2 and at least a part of the third image sequence 3.1 rendered by the virtual camera 3.2. The further images of the first image sequence 1.1 comprise a first image 1.1.1 and a second image 1.1.2. The first image 1.1.1 of the first image sequence 1.1 comprises a feature point cloud 1.3 and a surrounding object 1.4. From the first image 1.1.1 of the first image sequence 1.1 to the second image 1.1.2 of the first image sequence 1.1 the first real camera 1.2 has moved such that the feature point cloud 1.3 appears larger and the surrounding object 1.4 is not captured anymore.
[0095] The further images 2.1.1, 2.1.2 of the second image sequence 2.1 comprise a first image 2.1.1 and a second image 2.1.2. The first 2.1.1 and second 2.1.2 image of the second image sequence 2.1 each comprise a real person (moderator or user) as the real object 2.3 and a greenscreen as a background 2.4. The real person 2.3 is holding the support structure 5 while the receiving and capturing of the first 1.1 and second 2.1 image sequence. The greenscreen 2.4 fills the whole background in both first 2.1.1 and second 2.1.2 image.
[0096] The movement of the first real camera 1.2 is triggered by the real person 2.3 moving the support structure 5 and causes an analogue movement of the second real camera 2.2 since both real cameras are fixed to the same support structure 5. As the user with the first real camera 1.2 moves closer to the feature point cloud 1.3, the second real camera 2.2 moves also closer to the feature point cloud 1.3. In the example of Figure 2, the real person 2.3 keeps the same distance and orientation with respect to the second real camera 2.2 during the movement.
[0097] Therefore, the first 2.1.1 and second 2.1.2 image of the second image sequence 2.1 show the real object / person 2.3 with almost the same pose. This, of course, is just an example and in a different exemplary embodiment, the pose of the real object / person 2.3 in the second image sequence 2.1 may also change during the movement of the first real camera 1.2.
[0098] The third image sequence 3.1 comprises a first image 3.1.1 and a second image 3.1.2. The first image 3.1.1 of the third image sequence 3.1 comprises the composite augmented reality model 3.4 as rendered by the virtual camera 3.2. The composite augmented reality model 3.4 comprises the isolated real person 2.3 overlaid / composited with the three-dimensional virtual model 3.3 that has been aligned based on the feature point cloud 1.3.
[0099] In the example of Figure 2, initialization SO has already been carried out based on at least a first or first few images (not shown) of the first image sequence 1.1 that were captured and received before the further images 1.1.1, 1.1.2 as shown in Figure 2 were received and captured. During initialization SO the virtual three-dimensional model 3.3 has been aligned with the initial pose of the first real camera 1.2 and the initial pose of the augmented reality camera 1.2.1.
[0100] During the movement of the first real camera 1.2 after initialization SO and during rendering S3 the virtual camera 3.2 adapts its pose accordingly and instantaneously with a corresponding movement. The distance between the isolated real object 2.3 in its two-dimensional plane and the virtual camera 3.2 itself is fixed and constant during the receiving and / or capturing of the first 1.1 and second 2.1 image sequence and during the movement of the first real camera 1.2 and the corresponding movement / adaptation of the virtual camera 3.2 during the rendering S3.
[0101] Thus, the movement of the first real camera 1.2 from the first image 1.1.1 of the first image sequence 1.1 to the second image 1.1.2 of the first image sequence 1.1 causes a relative movement of the real person 2.3 with respect to the virtual three-dimensional model 3.3 in the rendered composite augmented reality model 3.4. The result of the relative movement is shown in the second image 3.1.2 of the third image sequence 3.1. Here the virtual three-dimensional model 3.3 remains aligned with the initial pose of the first real camera 1.2 shown in Figure 1. As the first real camera 1.2 moves closer to the feature point cloud 1.3, e.g., in the direction of the positive z-axis of the world coordinate system 6 in Figure 1, the virtual camera 3.2 adapts its pose accordingly by instantaneously moving also closer to the feature point cloud 1.3, e.g., along the positive z-axis of the virtual camera coordinate system which in Figure 1 is identical with the world coordinate system 6. However, in this way, the virtual camera 3.2 moves away from the virtual three-dimensional model 3.3 in Figure 1 that remains aligned with the fixed initial pose of the first real camera 1.2. Therefore, in the second image 3.1.2 of the third image sequence 3.1, the virtual three-dimensional model 3.3 appears smaller as compared to the first image 3.1.1 of the third image sequence 3.1.
[0102] The further images of the first 1.1 and the second 2.1 image sequence and the images of the third 3.1 image sequence as shown in Figure 2 each are associated with a time stamp representing the time of capturing by the respective camera. No substantial time delay occurs between the capturing and the receiving of said images.
[0103] The time stamp of the first image 1.1.1 of the first image sequence 1.1 and the time stamp of the first image 2.1.1 of the second image sequence 2.1 are identical. The time stamp of the second image 1.1.2 of the first image sequence 1.1 and the time stamp of the second image 2.1.2 of the second image sequence 2.1 are also identical. Therefore, the second image sequence 2.1 and the further images 2.1.1, 2.1.2 of the second image sequence 2.1 are captured time-synchronously with the first image sequence 1.1 and the further images 1.1.1, 1.1.2 of the first image sequence 1.1.
[0104] The time difference between the time stamp of the first image 3.1.1 of the third image sequence 3.1 and the time stamp of the first image 1.1.1 of the first image sequence 1.1 is very small due to an efficient processing of the image synthesis method. Similarly, the time difference between the time stamp of the second image 3.1.2 of the third image sequence 3.1 and the time stamp of the second image 1.1.2 of the first image sequence 1.1 is very small due to the efficient processing of the image synthesis method. Therefore, the third image sequence 3.1 or, equivalently the images of the third image sequence 3.1, is / are formed in real time and time-synchronously with the first 1.1 and second 2.1 image sequence and the further images of the first 1.1 and second 2.1 image sequence.
[0105] Recurring features are provided in the following figures with identical reference signs as in Figure 2.
[0106] Figure 3 shows a schematic flow diagram of an embodiment of the image synthesis method performed by the image synthesis apparatus shown in Figure 1. The image synthesis method comprises receiving the first image sequence 1.1, wherein the first image sequence 1.1 is captured by means of the first real camera 1.2.
[0107] The image synthesis method also comprises receiving the second image sequence 2.1, wherein the second image sequence 2.1 is captured by means of the second real camera 2.2. The second image sequence 2.2 is captured time- synchronously with the first image sequence 1.1 as described above.
[0108] While receiving and capturing at least the first or the first few images of the first image sequence 1.1, the image synthesis method comprises an initialization SO. During initialization SO the initial pose of the first real camera 1.2 is estimated based on tracking of a marker in at least the first or the first few images of the first image sequence 1.1. For this purpose, a marker is placed in the capture field of the first real camera 1.2 during initialization SO such that the at least first image or few first images depict / comprise the marker. Based on the marker the initial pose of the first real camera 1.2 is estimated using a marker-based pose estimation algorithm. The estimated initial pose of the first real camera 1.2 is then used to align the virtual three-dimensional model 3.3 and determine its initial pose. The initial pose of the virtual three-dimensional model is then kept fixed also after initialization SO and during the rendering S3, i.e., during the receiving and capturing of the further images of the first 1.1 and second 2.1 image sequence for the rendering of the composite augmented reality model 3.4 as shown in Figure 2.
[0109] While receiving and capturing further images of the first 1.1 and second 2.1 image sequence after initialization SO, the image synthesis method further comprises:
[0110] Time-continuously estimating the pose of the first real camera 1.2 based on the tracking of the feature points 1.3 in the further images of the first image sequence 1.1 in step Sil in order to be able to instantaneously adapt the pose of the virtual camera 3.2 to the estimated pose of the first real camera 1.2 after initialization SO and during the rendering S3 of the composite augmented reality model 3.4 when the first real camera 1.2 is moved during the rendering S3. Time-synchronously with the estimating Sil of the pose of the first real camera 1.2, time-continuously isolating the real person (moderator) as the real object 2.3 in the further images of the second image sequence 2.1 from its background 2.4 in the further images of the second image sequence 2.1 in Step S12.
[0111] The method further comprises overlaying and compositing in Step S2 the virtual three-dimensional model 3.3 aligned based on the at least first or the first few images of the first image sequence 1.1 during initialization SO together with the isolated real person 2.3 from the further images of the second image sequence 2.1 to form the composite augmented reality model 3.4 and rendering in Step S3 the composite augmented reality model 3.4 by means of the virtual camera 3.2 and thereby forming the third image sequence 3.1 in real time and time-synchronously with the further images of the first 1.1 and second
[0112] 2.1 image sequence.
[0113] Step S3 also comprises displaying the third image sequence 3.1 at a display as part of the image synthesis apparatus.
[0114] After initialization SO based on the at least first image or first few images of the first image sequence 1.1, steps Sil, S12, S2 and S3 are performed in real-time and quasi-time-synchronously (up to negligible time delays) based on the further images of the first 1.1 and second 2.1 image sequence as shown in Figure 2 and discussed further above.
[0115] Figure 4 shows a space-time path 4.1 with a first path node 4.1.1 and a second path node 4.1.2. The position of the virtual camera 3.2 has moved with respect to the initial pose in Figure 1 and coincides with the first path node 4.1.1. The first path node 4.1.1 defines a translational and orientational offset of the virtual camera coordinate system with respect to the world coordinate system 6, i.e., the initial pose of the first real camera 1.2, the initial pose of the augmented reality camera 1.2.1 and also with respect to the initial pose of the virtual camera 3.2 during initialization SO. In Figure 4, the first real camera
[0116] 1.2 and the augmented reality camera 1.2.1 are still at the origin of the world coordinate system 6 and have not moved since initialization SO. However, the pose of the virtual camera 3.2 is different from the initial pose of the first real camera 1.2 to enable a different viewing angle at the composite augmented reality model 3.4 in the third image sequence 3.1.
[0117] In Figure 4 the space-time path 4.1 has been set up to enable a free virtual camera 3.2 movement / drive along path nodes 4.1.1, 4.1.2 of the discrete space-time path 4.1 while receiving and capturing the further images of the first 1.1 and second 2.1 image sequence, wherein each path node 4.1.1, 4.1.2 of the space-time path 4.1 corresponds to a different translational and rotational virtual camera offset with respect to the initial pose of the first real camera 1.2 and also with respect to the initial pose of the virtual camera 3.2 as shown in Figure 1.
[0118] Figure 5 shows the instantaneous adaption of the pose of the virtual camera 3.2 from its current virtual camera offset when the first real camera 1.2 is moved after the virtual camera 3.2 has moved to the first path node 4.1.1. In such a case the virtual camera 3.2 adapts its pose accordingly to the pose of the first real camera 1.2 and performs the same corresponding motion starting from the current virtual camera offset corresponding to the first path node 4.1.1.
[0119] In Figure 5 the first real camera 1.2 rotates by an angle of 35° clockwise around the y-axis of the world coordinate system 6 (perpendicular to the z-x- plane of Figure 5). Note, that the world coordinate system 6 is identical to the first real camera coordinate system during initialization SO as shown in Figure 1. Accordingly, the virtual camera 3.2 also rotates by 35° around the y-axis (not shown) of the current local virtual camera coordinate system as defined by the virtual camera offset according to path node 4.1.1, i.e., this rotational movement of the first real camera 1.2 is added to the (current) virtual camera offset as defined by the first path node 4.1.1.
[0120] In Figure 6 the first real camera 1.2 performs a translation by a given distance dxalong the x-axis of the world coordinate system 6. Accordingly, the virtual camera 3.2 also translates by dxalong the x-axis of the local virtual camera coordinate system as defined by the first path node 4.1.1, i.e., this translational movement is added to the (current) virtual camera offset as defined by the first path node 4.1.1.
[0121] In Figure 7 the first real camera 1.2 performs a translation by a given distance dxalong the x-axis of the world coordinate system 6 and a translation by a given distance dzalong the z-axis of the world coordinate system 6 and a rotation by an angle of -15° anti-clockwise around the y-axis (not shown) of the world coordinate system 6. Accordingly, the virtual camera 3.2 also translates by dxalong the x-axis and by dzalong the z-axis and rotates by -15° anti-clock- wise around the y-axis of the local virtual camera coordinate system as defined by the first path node 4.1.1, i.e., this translational and rotational movement is added to the (current) virtual camera offset as defined by the first path node 4.1.1.
[0122] Thus, performing, by the first real camera 1.2, a translation along an axis of the world coordinate system 6 by a given distance leads to a corresponding translation of the virtual camera 3.2 along the same axis of the (current) virtual camera coordinate system as defined by first path node 4.1.1 by the same given distance.
[0123] Performing, by the first real camera 1.2, a rotation around an axis of the world coordinate system 6 by a rotation angle leads to a corresponding rotation of the virtual camera 3.2 around the same axis of the (current) virtual camera coordinate system as defined by first path node 4.1.1 by the same rotation angle.
[0124] Features of the different embodiments which are merely disclosed in the exemplary embodiments as a matter of course can be combined with one another and can also be claimed individually.
Claims
Claims1. An image synthesis method comprising: receiving a first image sequence (1.1), wherein the first image sequence (1.1) is captured by means of a first real camera (1.2); and receiving a second image sequence (2.1), wherein the second image sequence (2.1) is captured by means of a second real camera (2.2), wherein the second image sequence (2.1) is captured time-synchro- nously with the first image sequence (1.1); and performing an initialization (SO), wherein the initialization (SO) comprises estimating an initial pose of the first real camera (1.2) based on at least a first image of the first image sequence (1.1) and aligning a virtual three-dimensional model (3.3) based on the estimated initial pose of the first real camera (1.2); and while receiving and / or capturing further images of the first (1.1) and second (2.1) image sequence after initialization (SO), the method further comprises: time-continuously isolating (S12) a real object (2.3) in the further images of the second image sequence (2.1) from its background (2.4) in the further images of the second image sequence (2.1); and time-continuously rendering (S3) a composite augmented reality model (3.4) by means of a virtual camera (3.2) and thereby forming a third image sequence (3.1) in real time and time-synchronously with the further images of the first (1.1) and second (2.1) image sequence, wherein the composite augmented reality model (3.4) is formed by compositing the isolated real object (2.3) from the second image sequence (2.1) together with the virtual three-dimensional model (3.3) aligned based on the estimated initial pose of the first real camera (1.2) during initialization (SO).
2. Image synthesis method of one of the preceding claims, wherein the pose of the first (1.2) and second (2.2) real cameras are different and fixed relative to each other during the receiving and / or capturing of the first (1.1) and second (2.1) image sequence and / or during the initialization (SO) and / or during the time-continuously rendering (S3) the composite augmented reality model (3.4).
3. Image synthesis method of one of the preceding claims, wherein the time-continuously isolating (S12) a real object (2.3) in the further images of the second image sequence (2.1) from its background (2.4) in the further images of the second image sequence (2.1) comprises isolating the real object (2.3) by means of chromakeying technology or machine learning.
4. Image synthesis model of one of the preceding claims, wherein the isolated real object (2.3) is a two-dimensional object in the composite augmented reality model (3.4) and stays aligned in a fixed positional and / or orientational relationship with respect to the pose of the virtual camera (3.2) during the time-continuously rendering (S3) the composite augmented reality model (3.4).
5. Image synthesis method of one of the preceding claims, wherein during the time-continuously rendering (S3) the composite augmented reality model (3.4) the pose of the virtual camera (3.2) is different from the estimated initial pose of the first real camera (1.2) and / or the pose of the first real camera (1.2) during the rendering (S3) to enable a different viewing angle at the virtual three-dimensional model (3.3) in the composite augmented reality model (3.4) of the third image sequence(3.1).
6. Image synthesis method according to any of the preceding claims, wherein the method additionally comprises: during the time-continuously rendering (S3) the composite augmented reality model (3.4) performing a virtual camera (3.2) movement along path nodes (4.1.1, 4.1.2) of a continuous or discrete space-time path(4.1) while receiving and / or capturing the further images of the first(1.1) and second (2.1) image sequence, wherein each path node (4.1.1,4.1.2) of the space-time path (4.1) corresponds to a different translational and / or rotational virtual camera offset with respect to the initial pose of the first real camera (1.2) estimated during initialization (SO).
7. Image synthesis method according to any of the preceding claims, wherein the method further comprises: after the initialization (SO) and during the time-continuously rendering (S3) the composite augmented reality model (3.4) also time-continu- ously estimating (Sil) the pose of the first real camera (1.2) based on the further images of the first (1.1) image sequence and time-continu- ously and instantaneously adapting (Sil) the pose of the virtual camera (3.2) to the estimated pose of the first real camera (1.2).
8. Image synthesis method according to claim 7, wherein instantaneously adapting (Sil) the pose of the virtual camera (3.2) to the estimated pose of first real camera (1.2) comprises: when the first real camera (1.2) is moved by a translation along a first distance, the virtual camera (3.2) performs a corresponding translation by the same first distance and / or when the first real camera (1.2) is moved by a rotation along a first angle, the virtual camera (3.2) performs a corresponding rotation by the same first angle.
9. Image synthesis method according to claim 7, wherein time-continu- ously estimating (Sil) the pose of the first real camera (1.2) to time- continuously and instantaneously adapt (Sil) the pose of the virtual camera (3.2) is based on the tracking of feature points (1.3) in the further images of the first image sequence (1.1) and / or the capturing of translational and / or rotational motion data from at least one motion sensor characterizing the motion of the first real camera (1.2).
10. Image synthesis method according to any of the preceding claims, wherein the estimating the initial pose of the first real camera (1.2) during initialization (SO) to align the virtual three-dimensional model(3.3) is based on tracking of a marker in the at least first image of the first image sequence (1.1) and / or the capturing of translational and / or rotational motion data from at least one motion sensor characterizing the motion of the first real camera (1.2).
11. Image synthesis method according to any of the preceding claims, wherein the pose of the first real camera (1.2) and / or the pose of the virtual camera (3.2) is a six degree of freedom pose.
12. Image synthesis method according to any of the preceding claims, wherein the method is performed in a video conference application and / or wherein the isolated real object (2.3) corresponds to a real person captured by the second real camera (2.2) in the further images of the second image sequence (2.1) and / or the method further comprises displaying the third image sequence (3.1) at a user interface substantially time-synchronously with the receiving and / or capturing of the further images of the first (1.1) and second (2.1) image sequence.
13. Computer program comprising instructions which, when the program is executed by a computer, cause the computer to carry out the image synthesis method according to any one of claims 1 to 12.
14. Computer readable storage medium comprising a computer program according to claim 13.
15. An image synthesis apparatus comprising means to carry out the steps of the image synthesis method according to any one of claims 1 to 12.