METHOD AND DEVICE FOR OVERLATING AN IMAGE OF A REAL SCENERY WITH VIRTUAL IMAGE AND AUDIO DATA AND A MOBILE DEVICE
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2017-11-07
- Publication Date
- 2026-04-02
AI Technical Summary
Existing augmented reality systems struggle to accurately and continuously superimpose virtual images and audio data onto real scenes, especially when markers become partially or fully obscured, and require constant online synchronization for marker identification.
A method and device that uses a combination of geometric markers, such as QR codes, and natural features for tracking, allowing virtual objects to be anchored to real-world objects, enabling precise positioning even when markers are not visible, using a combination of image and audio processing to maintain accurate overlay.
Enables continuous, perspective-correct superimposition of virtual images and audio data over a wide area and varying distances without constant marker visibility, reducing the need for online synchronization and improving robustness against environmental disturbances.
Description
[0001] The invention relates to a method and a device for superimposing an image of a real scene with virtual image and audio data, wherein the method can be carried out, for example, using a mobile device, and a mobile device such as a smartphone.
[0002] The basic concept of Augmented Reality (AR) has existed for several decades and refers to the overlaying of real-time images of reality (e.g., camera images) with virtual information.
[0003] ZHOU Z ET AL: "An experimental study on the role of 3D sound in augmented reality environment"; INTERACTING WITH COMPUT, BUTIERWORTH-HEINEMANN; GB, Vol. 16, No. 6, 1 December 2004 (2004-12-01), pages 1043-1068; XP004654624, ISSN: 0953-5438, DOI: 10.1016 / J.INTCOM.2004.06.016 deals with the effect of sound on depth perception in an augmented reality (AR) environment. Several markers are used, positioned at varying distances from a camera.
[0004] JAKA SODNIK ET AL: "Spatial sound localization in an augmented reality environment", COMPUTER-HUMAN INTERACTION, ACM, 2 PENN PLAZA, SUITE 701 NEW YORK NY 10121-0701 USA, November 20, 2006 (2006-11-20); Pages 111-118; XP058245079; DOI: 10.1145 / 1228175.1228197; ISBN: 978-1-59593-545-8) deals with the combination of sounds with virtual objects in a virtual environment.
[0005] RUMINSKI DARIUSZ: "Modelling spatial sound in contextual augmented reality environments", 2015 6TH INTERNATIONAL CONFERENCE ON INFORMATION, IN-TELLIGENCE, SYSTEMS AND APPLICATIONS (IISA), IEEE, July 6, 2015 (2015-07-06); Pages 1-6; XP032852099, DOI: 10.1109 / IISA.2015.7387982 [found on 2016-01-20]) deals with the combination of visual and acoustic content in a virtual environment.
[0006] ULRIC H NEUMANN ET AL: "Natural Feature Tracking for Augmented Reality", IEEE TRANSACTIONS ON MULTITIME DIA, IEEE, USA, Vol. 1, No. 1, March 1, 1999 (1999-03-01), XP011 036279, ISSN: 1520-9210 discloses an application of natural feature tracking as an automatic extension of the working area of an AR system by using naturally occurring features as additional reference points.
[0007] SATO T ET AL: "CAMERA PARAMETER ESTIMATION FROM A LONG IMAGE SEQUENCE BY TRACKING MARKERS ANO NATURAL FEATURES", SYSTEMS & COMPUTERS IN JAPAN, WILEY, HOBOKEN, NJ, US, Vol. 35, No. 8, July 1, 2004 (2004-07-01), Pages 12-20, XP001196294, ISSN: 0882-1666, DOI: 10.1002 / SCJ.10702 discloses a method in which predefined markers do not need to be visible in an input sequence, as 3D positions of natural features are detected and used instead of markers.
[0008] The invention aims to create an improved method and device for superimposing an image of a real scene with virtual image and audio data, as well as an improved mobile device, compared to the prior art.
[0009] This problem is solved by a method and a device for superimposing an image of a real scene with virtual image and audio data, as well as a mobile device according to the main claims. Advantageous embodiments and further developments of the invention are described in the dependent claims.
[0010] The described approach specifically addresses the field of optically and acoustically congruent augmented reality, in which virtual objects and audio data are linked to selected anchor points in the real scene and always superimposed into the three-dimensional scene with correct perspective, as if they were part of the real environment. According to one embodiment, each individual frame of a camera stream can be analyzed using image and / or audio processing methods, and the necessary three-dimensional position and orientation of the virtual object can be calculated to achieve this effect. Advantageously, the described approach allows for continuous tracking of the scene as the viewer moves.
[0011] Selecting a virtual object superimposed on the real-world scene, hereinafter also referred to as a virtual image and audio object or virtual image and audio data, can be advantageously accomplished using a marker present in the real-world scene, such as a QR code. The object can be stored as a three-dimensional data repository in a database. Additionally or alternatively, the object can consist of a sequence of recordings, such as photographs and / or audio recordings, taken from different angles (360°) and stored in the database. In the case of a three-dimensional data repository, points of the object can be defined by coordinates in a coordinate system, or a single point and vectors for determining all other points of the object. The sequence of recordings can be a series of two-dimensional images. Each of the recordings can depict the object.Positioning the virtual image and audio data within a representation of the real-world scene can be advantageously achieved using at least a segment of an object, such as an edge or face, located in the vicinity of the marker in the real-world scene. A representation of this object segment can then be used as a new and / or additional anchor point for the virtual object. The marker can occupy less than 1%, for example, only 0.6%, 0.1%, or even 0.01% of the representation of the real-world scene.
[0012] Using the marker ensures, with minimal effort, that the virtual image and audio data that match the real-world scene are selected. Using the object section guarantees that the virtual image and audio data can be positioned very precisely, even under adverse conditions such as poor lighting. This positioning remains possible even if the marker is no longer visible or only partially visible in subsequent renderings of the real-world scene.
[0013] An optical image of an object is the reflection, perceived by the eye, of optically visible waves with a typical wavelength of 400-800 nm. These waves initially strike the object, are reflected by it, and arrive at the observer's eye. In the case of light sources, the object itself emits visible light at specific points. Similarly, an acoustic "image" of an object or its surroundings can be created by the corresponding reflection of audible waves, for example, with a typical frequency of 20-20,000 Hz. These waves are reflected by the object or its surroundings and can be interpreted by the observer's ears as a spatial "image." Just as a light source can emit sound sources at different points, the object itself can also create a spatial impression (example: an orchestra).Similarly, blind people can create and reproduce a "spatial image" through clicking sounds and reflections from their surroundings. Echo sounders work in the same way; an electronic spatial image / image of the object is generated from the incoming sound waves and displayed on a screen; it is also possible to create a corresponding acoustic image of the environment in the mind of the observer.
[0014] The approach described here involves overlaying the virtual image or audio data onto the camera-captured representation of the environment displayed on the screen at any given time while the viewer is in motion. This overlay is applied in the correct scale, position, and angle relative to the marker (e.g., the QR code) and the image markers. The viewer then perceives this "overall image" as a seemingly real, unified image captured by the camera. Simultaneously, the virtual image and / or audio element should emit sound at precisely the same intensity and / or quality at any given time and from any direction of the viewer / listener as it would in reality. Naturally, the emitted sound waves are adjusted in frequency and / or volume according to the distance and angle of the emitting object.The movement of the emitting object results in a corresponding distortion (Doppler effect). As you walk around the object, some sound sources will "disappear" while others "appear." This very rendering process is controlled on the screen or in the headphones by the approach described here.
[0015] To determine a marker and its positioning within the image data, and to determine the image and audio data via the marker data and their positioning in relation to the image, suitable known methods can be used, with many ways to solve the corresponding sub-steps being known.
[0016] A method for overlaying an image of a real scene with virtual three-dimensional or two-dimensional image and audio data includes the following steps: Reading image data representing a visual representation of the real-world scene captured by at least one environmental sensing device of a mobile device; determining marker data from the image and audio data, where the marker data represents an image and position of a marker located in the real-world scene; reading virtual image and audio data selected using the marker data. The read data, consisting of multiple virtual three-dimensional and / or two-dimensional image and audio data, also includes a rendering instruction for displaying a virtual image, a positioning instruction for positioning the virtual image, and a positioning instruction for playing back acoustic data and / or a trigger for playing the audio data.Determining object data from the image data, wherein the object data consists of an optical and / or acoustic three-dimensional image or a series of two-dimensional photographs and / or sound recordings from different angles and a positioning of an object segment of an object located in the vicinity of the marker in the real-world scene; determining a positioning instruction for positioning the virtual image and the acoustic data or additional virtual audio data associated with this virtual image in relation to the image of the object segment using the marker data, the object data, and the virtual image and audio data.
[0017] The real-world scene could be, for example, an area surrounding the mobile device that lies within the detection range of one or more environmental sensing devices. The environmental sensing device could be an optical image capture device, and an optional additional environmental sensing device could be an acoustic sound capture device, such as one or more cameras or microphones. The virtual representation can also be referred to as a virtual image. The virtual representation can include virtual image and audio data.The virtual image and audio data can include a rendering instruction for the visual and / or acoustic representation of a three-dimensionally defined object and / or for the representation of a selection of visual and / or acoustic recordings captured from different angles, for example, in the form of two-dimensional photographs or sound recordings of an object. The rendering instruction can be used to overlay the visual and acoustic representation of the real-world scene with the virtual three-dimensional or two-dimensional image and audio data. The representation from which the object data is determined in the determination step can depict image data and, optionally, audio data of the real-world scene captured using the environment sensing device(s), which can be displayed or output, for example, using the display and output devices of the mobile device.Virtual image and audio data can be understood as any visual and acoustic representation, such as a graphic, symbol, text, conversation, music, or other sounds, that can be inserted into the representation of the real-world scene. The virtual image and audio data can represent a three-dimensional or two-dimensional image along with associated audio data, or a single point or sound source. The virtual image and audio data can be selected from a range of sources. An overlay of the visual and acoustic representation of the real-world scene with the virtual image and audio data can include an area in which at least one region is completely or, for example, partially obscured by the virtual image and audio data.According to one embodiment, the virtual audio data comprises stereo audio data, which, for example, is provided to a stereo loudspeaker via a suitable interface and can be output by the stereo loudspeaker. Stereo audio data has the advantage of providing the listener with a sense of direction in which a virtual sound source associated with the virtual audio data appears to be located. The virtual audio data can include the acoustic data that can be used for overlaying. A marker can be understood as an artificial marker placed in the scene, for example, a geometric marker in the form of a code or pictogram. The marker can be implemented as an artificial marker in the form of a one-dimensional or two-dimensional code. For example, the marker can be implemented as a matrix with light and dark areas. The marker can represent optoelectronically readable text.The marker can represent data in the form of a symbol. The marker data can include information about the marker's representation and its position within the representation of the real-world scene. In subsequent steps of the process, the marker data can be used in whole or in part, and optionally in a further processed form. The positioning instruction for positioning the virtual image and audio data can be used to position the virtual image and audio data relative to the marker's representation within the representation of the real-world scene. The object segment can be a part, section, or area, such as an edge, surface, or even an acoustically defined area of a real-world object. An object can be any item, such as a building, a piece of furniture, a vehicle, a musical instrument, or a piece of paper.The object segment can be, for example, an outer edge or an edge between angled surfaces of such an object. The object data can include information about the optical and acoustic representation of the object segment and the positioning of this representation within the representation of the real-world scene. In the subsequent steps of the process, the object data can be used in whole or in part, and possibly in a further processed form. The positioning instruction can be suitable for positioning the virtual image and audio data relative to the optical and acoustic representation of the object segment in the corresponding representation of the real-world scene or another representation of the real-world scene.The positioning instruction can be determined using the positioning of the marker image, the positioning of the optical and / or acoustic image of the object section, and the positioning instruction.
[0018] The aforementioned object segment, or rather its representation, can be considered an anchor point. Such an anchor point can be used in addition to, or as an alternative to, the marker for positioning the virtual representation and the acoustic data. Therefore, it is not always necessary to use the marker itself, such as the QR code, to position the virtual object—that is, the virtual representation and the acoustic data. Instead, the marker can be extended with one or more anchor points from its surroundings, allowing the marker to be tracked even when it is no longer visible in the image—that is, the representation of the real-world scene displayed on the mobile device's screen.
[0019] Thus, during the data acquisition step, the acquired image data can represent or include audio data in addition to the image data. The audio data is also referred to as sound data. This audio data can represent an acoustic representation of the real-world scene, captured by at least one additional environmental sensing device of the mobile device. In this way, for example, background noise associated with the captured optical image data can be recorded and processed. The additional environmental sensing device can include, for example, one or more microphones. When using multiple microphones or a directional microphone, a sound source emitting the captured audio data can be located. This localization information can then be compared with the captured image data.
[0020] According to one embodiment, the method for superimposing an image of a real scene with virtual image and audio data comprises the following steps: Reading in optical and / or acoustic image and audio data, wherein the image and audio data represent an image of the real scene captured by an environment sensing device of a mobile device; determining marker data from the image and audio data, wherein the marker data represent an image and a position of a marker arranged in the real scene; reading in virtual image and audio data, wherein the virtual data represent three-dimensional or sequence of two-dimensional recordings of image and audio data selected from a plurality of virtual data using the marker data, wherein the virtual image and audio data include a rendering instruction for displaying the virtual image and a positioning instruction for positioning the virtual image, as well as a trigger position for playing the virtual audio data;Determining object data from the image and sound data, wherein the object data represents a representation and a position of an object segment of an optically or acoustically recognizable object located in the vicinity of the marker in the real-world scene; determining a positioning instruction for positioning the virtual image in relation to the representation of the object segment and to the starting position for playing the audio data using the object data and the virtual image and sound data.
[0021] In general, the image and audio data can consist of real three-dimensional or a sequence of two-dimensional image and sound data, the object data of real object data, and the object segment of a real object segment.
[0022] According to one embodiment, the positioning instruction can be determined in the determination step using the marker data or at least a part of the marker data. By defining further anchor points and / or anchor lines in a defined sequence, the optical and acoustic representation of the real scene can be tracked in the real scene, even if the actual marker can no longer be detected by the environment sensing device of the mobile device.
[0023] According to one embodiment, a continuous iteration of the steps of reading, determining, and ascertaining can be performed at short intervals, in particular several times per second. For example, the steps can be executed between ten and two hundred times per second (i.e., every tenth of a second or every 5 / 1000 of a second).
[0024] The described approach enables the positioning of the virtual optical / acoustic object in a perspective-correct representation from a great distance and from a relatively unrestricted position of the mobile device. Advantageously, it is no longer necessary for the mobile device to recognize the marker and position the associated virtual object in a fixed position relative to this marker, but rather in a defined position relative to these additional anchor points / lines. A great distance can be understood as a distance between ten and five thousand times the side length of the marker, for example, the QR code. According to one embodiment, the range between ten and five hundred times the side length of the marker is preferred. With an edge length of 2 cm for the marker, this corresponds to a distance of up to 100 m (5000 times the edge length).The relatively unrestricted position can be understood as deviations between 0.1° and 180° in all three axes. This is intended to cover 360° in all directions. It is also not necessary for the marker to be constantly within the field of view (environment detection device) of the mobile device.
[0025] According to one embodiment, the described approach uses the measuring devices located in the mobile device – in addition to image acquisition – to measure the change in the relative position – after the marker has been detected – compared to the position fixed during the initial detection of the marker. Additionally, data from a real object, extracted from the actual image and sound data, is used as an object segment, also referred to as a "secondary marker," so that the actual marker no longer needs to be within the detection range of the environmental detection device.
[0026] The following devices, also known as detection devices or measuring sensors, can be used in mobile devices, such as smartphones or tablets, to determine deviations from the initial position after the marker has been detected once. Individual measuring sensors or any combination thereof can be used.
[0027] Accelerometer: on the one hand for measuring translational movements of the mobile device, on the other hand for determining the direction of the earth's gravity relative to the device and thus the orientation / rotation of the device.
[0028] Rotation sensor: for measuring rotational movements of the mobile device.
[0029] Magnetometer: used to measure the Earth's magnetic field and thus the horizontal rotation of the mobile device.
[0030] GPS receiver: optional for very large distances and for rough positioning with an accuracy of ± 2 meters.
[0031] Microphone: for detecting and measuring individual sound sources and / or general background noise. Frequencies in the audible range (20 - 20000 Hz) are preferred, but frequencies in the ultrasound range can also be used.
[0032] The use of an accelerometer and a rotation sensor as a supplement to the image capture device is preferred.
[0033] The image acquisition device may be limited to visible light (400-800nm), but may also additionally or exclusively capture other spectral ranges (e.g. additionally or exclusively IR or UV light).
[0034] For example, measurements from a suitable measuring device can be used to determine a displacement of the object section or the image of the object section caused by movement of the mobile device. According to one embodiment, a value representing the displacement is used in the determination step to determine the positioning instruction for positioning the virtual image relative to the image of the object section.
[0035] Thus, the positioning instruction can be determined, for example, using a measurement from one or more measuring devices, such as an accelerometer, a rotation sensor, a magnetometer or a GPS receiver, of the mobile device.
[0036] This solves a further technical problem that arises when the virtual object is supposed to move in reality. If the marker disappears from the field of view of the environmental detection device while tracking this movement, the virtual representation does not collapse. This now allows image sequences to be displayed over a wide area.
[0037] Additionally, audio data can now be played at various freely chosen positions for a more realistic representation of the virtual object.
[0038] According to one embodiment, the method includes a step of providing at least a portion of the marker data to an interface with an external device. In this case, during the step of reading virtual three-dimensional or selected two-dimensional, or a sequence of, such image and audio data, the virtual image and audio data can be read via the interface to the external device, for example, a server. This interface could, for example, be a wireless interface. Advantageously, the selection of the virtual image and audio data can be performed using the external device. This saves storage space on the mobile device and ensures that up-to-date virtual image and audio data are always available.
[0039] The process can include a step of selecting the appropriate virtual image and audio data from a plurality of virtual image and audio data using the marker data. This selection step can be performed using an external device or on the mobile device itself. The latter offers the advantage that the process can run autonomously on the mobile device. The virtual image and audio data can be selected by, for example, comparing the marker image or marker identification with the multiple virtual images associated with it or with the identifications of potential markers, and selecting the virtual image that matches. In this way, the appropriate virtual image and audio data can be selected with a high degree of certainty.
[0040] The process can include a step of determining the marker's identification using the marker data. In the selection step, the virtual image and audio data can then be selected using this identification. An identification could be, for example, a code or a string of characters.
[0041] For example, the marker can represent a machine-readable code that includes a corresponding marker identification. In this case, the marker identification can be determined as part of the marker data during the marker data determination step. Using a machine-readable code, the marker representation can be evaluated very easily.
[0042] The process can include a step of using the positioning rule to overlay another image of the real-world scene with the virtual image and audio data. Advantageously, the positioning rule, once determined, can be used to overlay the virtual image and audio data onto successive images of the real-world scene.
[0043] For example, the use step may include a step of reading in further image data representing the further image of the real-world scene captured by the mobile device's environment sensing device; a step of determining the positioning of a further image of the object segment from the further image data—which may be in the form of three-dimensional points in a coordinate system, points and vectors, or selections of two-dimensional photographs; and a step of creating superimposed image and audio data using the further image data, the further image of the object segment, and the positioning instruction, with the superimposed image and audio data representing a superposition of the further image of the real-world scene with the virtual image and audio data.In the positioning step, the positioning of the further representation of the object segment within the further representation of the real scene can be determined. Thus, optical and acoustic representations of the object segment can be used as anchor points for the virtual image and audio data within temporally and spatially appropriate representations of the real scene. In the step of creating superimposed image and audio data, the virtual image and audio data can be displayed using the rendering instructions.
[0044] The process can include a step of displaying a superimposed image of the real-world scene with the virtual image and audio data using a display and playback device on the mobile device. For example, the aforementioned superimposed image and audio data can be provided to the display and playback devices. The display device could be, for example, a screen or a display, and the playback device a speaker or a stereo playback interface.
[0045] The process can include a step of capturing image and, optionally, audio data using at least one environmental sensing device of the mobile device. For example, image and audio data can be captured continuously over time, so that virtual representations of the real-world scene can be provided. The virtual image and audio data can then be superimposed onto each of these virtual representations of the real-world scene.
[0046] Depending on the specific implementation, multiple virtual three-dimensional objects or two-dimensional images and audio data can also be used for superimposition. In this case, several virtual image and audio data can be read in during the import step, or the virtual image and audio data can include rendering and positioning instructions for displaying and positioning the multiple virtual images and audio data.
[0047] Similarly, multiple object sections of one or different objects can be used. In this case, multiple object data points can be determined in the object data definition step, or the object data can represent images and positions of the majority of object sections. Accordingly, in the positioning rule determination step, multiple positioning rules for positioning the virtual image relative to individual object sections can be determined. Alternatively, a single positioning rule can be determined that is suitable for positioning the virtual image and audio data relative to the images of the majority of object sections.The use of multiple object sections offers the advantage that the virtual image and audio data can be positioned very precisely, and can still be positioned even if not all object sections used are depicted in an image of the real scene.
[0048] The approach presented here further creates a device designed to perform, control, and implement the steps of a variant of the method presented here in appropriate facilities. This embodiment of the invention in the form of a device also allows the problem underlying the invention to be solved quickly and efficiently.
[0049] The device can be configured to read input signals and, using these input signals, determine and provide output signals. An input signal can, for example, be a sensor signal readable via an input interface of the device. An output signal can be a control signal or a data signal that can be provided at an output interface of the device. The device can be configured to determine the output signals using a processing instruction implemented in hardware or software. For example, the device can include a logic circuit, an integrated circuit, or a software module and may be implemented as, or comprised of, a discrete component.
[0050] It is also advantageous to have a computer program product with program code that can be stored on a machine-readable medium such as semiconductor memory, hard disk memory or optical memory and is used to carry out the method according to one of the embodiments described above, if the program product is executed on a computer or device.
[0051] Exemplary embodiments of the invention are shown in the drawings and explained in more detail in the following description. It shows: Fig. 1 is an overview of a method for superimposing an image of a real scene with virtual image and audio data according to an embodiment; Fig. 2 is an overview of a method for creating an assignment rule according to an embodiment; Fig. 3 is a schematic representation of a mobile device according to an embodiment; Fig. 4 is a flowchart of a method for superimposing an image of a real scene with virtual image and audio data according to an embodiment; and Fig. 5 is a QR code placement square with binary contours according to an embodiment.
[0052] Fig. 1 shows an overview of a method for superimposing an image of a real scene with virtual image and audio data according to an exemplary embodiment.
[0053] In the left half of Fig. 1 A mobile device 100, for example a smartphone, is shown, comprising an environment sensing device 102, a further environment sensing device 103, a display device 104, and an output device 105. According to this embodiment, the environment sensing devices 102 and 103 are configured as a camera and a microphone, respectively, designed to capture a real scene 106, also referred to as the real environment, located within the detection range of the environment sensing devices 102 and 103. According to this embodiment, the display devices 104 and 105 are configured as a display and a loudspeaker, respectively, designed to show an image 108 of the real scene 106, captured by the environment sensing devices 102 and 103, to an operator of the mobile device 104.
[0054] In the real-world scene 106, according to this embodiment, an object 110 is arranged, on whose outer surface a marker 112 is positioned. For example, the object 110 can be any image or object. The object 110 lies partially and the marker 112 completely within the detection range of the environmental detection devices 102, 103. In particular, at least one object segment 114 of the object 110 lies within the detection range of the environmental detection devices 102, 103. Thus, the image 108 comprises an image 116 of the marker 112 and at least one image 118 of the object segment 114.
[0055] In the right half of Fig. 1 The mobile device 100 is shown at a later point in time compared to the representation on the left half. Due to an interim movement of the mobile device 100, the real scene 106 has changed slightly from the perspective of the environmental detection devices 102, 103, so that the display device 104 shows a further image 120 that is slightly modified in relation to the image 116. For example, the further image 120 can depict the real scene 106 from a different perspective, including a different audio perspective, or a different section of the real scene 106 compared to the image 108. For example, the different section is such that the further image 120 includes a further image 122 of object section 114 but not a further image of marker 112. Nevertheless, virtual image and audio data 124, 125 can be superimposed on the further image 120 using the described method.According to one embodiment, the virtual image and audio data 124, 125 are to be superimposed on the further image 120 in a predetermined position and / or orientation. Such a predetermined superimposition is possible, according to one embodiment, as long as the further image 120 includes a suitable further image 122 of the object section 106, which can be used as an anchor point for the virtual image and audio data 124, 125.
[0056] The steps of the procedure can be performed exclusively using the facilities of the mobile device 100 or additionally using at least one external facility, which is represented here as an example of a cloud. For example, the external facility 130 can be connected online to the mobile device 100.
[0057] According to one embodiment, the virtual image and audio data 124, 125 are generated only using data acquired by the environment sensing device 102, i.e. no real audio data are used.
[0058] The procedure can be executed continuously or started with a content request or a view of the real scene 106 requested by the operator using the display devices 104.
[0059] Image 108 is based on image and audio data provided by the environmental sensing devices 102, 103, or one of the evaluation devices downstream of the environmental sensing devices 102, 103. For example, using an object recognition method or another suitable image and audio processing method, marker data 132 and object data 134 are determined from the image and audio data, as shown schematically here. The marker data 132 is determined by suitable extraction from the image and audio data and includes identification data 136 assigned to marker 112, for example, an identification ID assigned to marker 112 and / or an address or pointer assigned to marker 112, for example, in the form of a URL.The marker data 132, or parts thereof, or data derived therefrom, such as the identification assigned to the marker, can be used to select virtual image and audio data 140 assigned to the marker 112 from a plurality of virtual image and audio data using an assignment rule 138, for example, an assignment table, which, according to this embodiment, is stored in a storage device of the external device 130. The plurality of virtual image and audio data can be stored in the assignment table 138 in the form of AR content. The virtual image and audio data 140 are transmitted to the mobile device 100 and used to display or play back the virtual image 124.According to one embodiment, the selection of the virtual image and audio data 140 is only carried out if a new marker 112 is found, i.e., for example, if the image 116 of the marker 112 or the identification data 136 of the marker 112 has been extracted for the first time from the image and audio data representing the image 108.
[0060] The object data 134 are determined by appropriately extracting suitable image and / or sound features from the image and audio data. These suitable image / sound features are used to create a positioning instruction 142, also called a new AR marker, for example, for temporary and local use. The positioning instruction 142 is used by the mobile device 100 to display the virtual image and audio data 124 as a superimposition of the image 106 or the further image 120, even when no image 116 of the marker 112 is available. No online synchronization is required for using the positioning instruction 142. According to this embodiment, the positioning instruction 142 refers to the object segment 114, which represents a natural marker.
[0061] According to one embodiment, this enables a secure assignment of the AR content using a URL and stable 3D tracking using a new, and therefore up-to-date, natural marker.
[0062] According to one embodiment, at least two natural markers, for example object section 114 and another object section 144 of object 110, are used to position the virtual image and audio data 124, 125 in the further image 120. In this case, the positioning instruction 142 refers to the two object sections 114, 144 or their images 118, 122, 146. In the further image 120 of the real scene 106, the in Fig. 1 In the illustrated embodiment, the further object section 144 is not shown. Nevertheless, the virtual image and audio data 124, 125 can be positioned using the further image 122 of the object section 114.
[0063] According to one embodiment, the described approach is based on a combination of two methods that can be used to extract three-dimensional positions of objects from camera images.
[0064] In the first of these methods, predefined geometric shapes are used as markers 112, which are placed within the camera image area, e.g., QR codes. Based on the known shape of such a marker 112 and its image 116 in the camera image 108, its three-dimensional position in space can be determined using image processing. Advantages of the first method are that, due to predefined design rules for the marker 112, it can be unambiguously identified in the camera image 108, and that additional information can be directly encoded in the appearance of the marker 112, such as the ID of a marker 112 or a web link via QR code. Thus, a very large number of different markers can be unambiguously distinguished from one another visually by means of a uniquely defined encoding scheme, e.g., black and white bits of the QR code.
[0065] A disadvantage, however, is that these markers 112, due to their precisely defined shape, are hardly robust against minor disturbances in the camera image 108. Such minor disturbances can, for example, represent slight focus blurring, motion blur, or a steep viewing angle. This means that the three-dimensional position of one of these markers 112 can only be correctly extracted if it is fully in focus, parallel to the image plane, and unobstructed in the camera image 108, and if the camera 102 is almost stationary relative to the marker 112. This makes, for example, the continuous, correctly oriented AR overlay of a virtual 3D object 124 based on a marker 112 in the form of a QR code virtually impossible.By making a geometric marker 112 sufficiently large, this problem is slightly improved, but this comes with the further disadvantage that it then has to be placed very prominently and large in the scene 106, which is unsuitable for most applications.
[0066] In the second of these methods, which can also be called Natural Feature Tracking or NFT, images of objects 110 located in the real environment 106, e.g., the cover image of a flyer, are defined in advance as markers, and their natural optical features 114, e.g., distinctive points, edge profiles, or colors, are first extracted from the original in a suitable form by an algorithm—in other words, trained on them. For AR position determination, i.e., for determining the position of a virtual image 124 to be superimposed, the camera image 108 is then searched for these previously trained natural features 114. Optimization procedures determine, firstly, whether the object 110 being sought is currently located in the camera image 108, and secondly, its location and position are estimated based on the arrangement of its individual features 114. The advantage of this is that the optimization-based method offers high robustness against disturbances.Thus, the positions of marker objects 114 can still be detected even in blurred camera images 108, 120, with partial occlusion, and at very steep angles. More advanced methods (e.g., SLAM) even make it possible, based on the initial detection of a marker object 114 in the camera image 108, 120, to continuously expand its model with features from the current environment, so that its position in space can still be correctly determined even when it is no longer visible in the camera image 120. However, this method has significant disadvantages, especially when a large number of different markers are to be detected. For example, each marker object 114 must first meet certain optical criteria regarding its natural optical appearance in order to be detectable in the camera image 108, 120 at all.
[0067] Furthermore, for unambiguous identification, all recognizable markers 114 must be clearly distinguishable from one another in their visual appearance – the greater the number of recognizable markers 114, the higher the probability of misidentification. This is particularly problematic when many visually similar objects 100, e.g., business cards, are to be distinguished within a database. Additionally, at the time of recognition, a database containing the natural characteristics of all recognizable markers must already exist, and this complete database must be compared with the camera image 108, 120 to determine whether one of the markers 114 is present in the camera image.In a system like a smartphone AR app with a constantly growing marker database, this requires keeping the latest version of the database in a central location (online), while each smartphone has to send a computationally intensive image search request to this database to analyze each individual camera image.
[0068] The approach described here, according to one embodiment, is based on a combination of the two methods described above. For the detection and 3D positioning of marker objects in the camera image 108, 120, both methods are carried out in successive, interconnected stages: In the first stage, for the pure identification of virtual image and audio data 140 of a virtual image 124, referred to here as AR content 124, a geometric, predefined marker design, e.g., a QR code or a barcode, is used as an image 116 of the marker 112 in the camera image 108. For example, the image 116 of the marker 112 can occupy only 0.6%, 0.1%, or even 0.01% of the image 108 of the real scene 106. This corresponds to a side length of 0.5 cm for the image 116 of the marker 112 on a DIN A4 sheet.
[0069] The detection of a marker 112 in the form of a QR code in the respective camera image under examination will later be performed using Fig. 5 described in detail.
[0070] According to one embodiment, the microphone 103 or the loudspeaker 105, or, if available, multiple microphones and / or multiple loudspeakers of the smartphone 100, are thus included. The selection of the virtual data 140 therefore depends on the detection of a primary marker 116 (QR codes / barcodes) by the camera 102 of the smartphone 100. The selected virtual data 140 now consists not only of image data but also of sound data, which are played back depending on the further movement of the virtual object 124 superimposed on the real scene.
[0071] To simplify the concept: imagine a three-dimensional television film (shot with a series of cameras in a 360° field of view – for example, 36 cameras spaced 10° apart or 72 cameras spaced 5° apart) that takes place in the open space of a living room. Naturally, the virtual image and sound objects 140 are displayed with correct perspective, even as the smartphone 100 moves around the scene, i.e., when secondary markers 122 are used. For the correct rendering of the sound objects, it is particularly desirable to play the audio data through stereo headphones. Suitable stereo headphones can be connected to the smartphone 100 via an appropriate interface. In a further implementation, these secondary markers 122 will include not only image features but also sound features of the real scene.This includes, for example, singular sound sources of specific tones or the specific arrangement of musical instruments.
[0072] Fig. 2 Figure 1 shows an overview of a method for creating an allocation rule 138 according to an exemplary embodiment. The allocation rule 138 can, for example, be implemented in the following. Fig. 1 The external device shown will be stored.
[0073] An operator 250 provides 3D AR content 252, for example, in the form of multiple virtual image and audio data. A web interface 254 is used to create or update the mapping rule 138 based on the 3D AR content 252. According to one embodiment, the mapping rule 138 includes a link to a specific, unique URL for each piece of 3D AR content.
[0074] Fig. 3 Figure 1 shows a schematic representation of a mobile device 100 according to an exemplary embodiment. The mobile device 100 could, for example, be the one described in Figure 1. Fig. 1 The mobile device shown is a mobile device. The mobile device 100 has environment sensing devices 102, 103 and display devices 104, 105 for displaying an image of a real scene detected by the environment sensing device 102. Virtual image and audio data can be superimposed on the image. According to this embodiment, the mobile device 100 includes an interface 360, for example, an interface for wireless data transmission, to an external device 130. According to one embodiment, the environment sensing device 102 is arranged on a rear side and the display device 104 on a front side of the mobile device 100.
[0075] The mobile device 100 has a reading device 362 coupled to the environmental sensing devices 102, 103, which is configured to read image and audio data 364, 365 from the environmental sensing devices 102, 103 as raw data or already processed data. For example, the reading device 362 is an interface to the environmental sensing devices 102, 103. The image and audio data 364, 365 represent a representation of the real scene as captured by the environmental sensing devices 102, 103. The image and audio data 364, 365 read in by the reading device 362 are further processed in a determination device 366 of the mobile device 100. In particular, marker data 132 and object data 134 are determined, for example, extracted from the image data 364 and optionally from the audio data 365. The marker data 132 represents an image and a position of a marker arranged in the real-world scene, for example, the one in Fig. 1 The geometric marker 112 shown represents the object data 134. This data represents a representation and positioning of an object segment of an object located in the vicinity of the marker in the real-world scene. For example, the object segment could be the one shown in Fig. 1 The object section 114 shown can be used as a natural marker. For this purpose, the detection device 366 is designed to first recognize the image of the marker in the image of the real scene and then to determine the marker data associated with the image of the marker from the image and audio data 364, 365. Similarly, the detection device 366 is designed to first recognize one or more suitable images of object sections in the image of the real scene and then to determine the object data associated with the image(s) of the suitable object section(s) from the image and audio data 364, 365. According to one embodiment, only the image data 364 and not the audio data 365 are used for this purpose.
[0076] According to this embodiment, the marker data 132 are provided to the external interface 360 and transmitted via the external interface 360, for example a radio interface, to the external device 130, for example in the form of an external unit. The external device 130 has a selection unit 368 configured to select virtual image and audio data 140 assigned to the marker data 132 from a plurality of virtual image and audio data using an assignment rule and to provide them to the external interface 360 of the mobile device 100. Alternatively, only parts of the image and audio data 132 or the image and audio data 132 in a further processed form can be provided to the input devices 360 and / or the external device 130. The external interface 360 is configured to provide the virtual image and audio data 140 to a destination device 370.The virtual image and audio data 140 comprise a display instruction for displaying a virtual image and a positioning instruction for positioning the virtual image or the image of an object, as well as an instruction for the playback positioning of the virtual audio data. The determination device 370 is further configured to receive the marker data 132 and the object data 134. The determination device 370 is configured to determine a positioning instruction 142 for positioning the virtual image relative to the image of the object section using the marker data 132, the object data 134, and the virtual image and audio data 140.
[0077] According to this embodiment, the mobile device 100 comprises a control unit 372 for controlling the display unit 104. The control unit 372 is configured to provide superimposed image and audio data 376, for example in the form of a control signal for controlling a display shown by the display unit 104, to the display unit 104, 105. The superimposed image and audio data 376 represent a superimposition of a further image of the real scene with the virtual image and audio data. The control unit 372 is configured to generate the superimposed image and audio data 376 using the positioning instruction 142 provided by the determination unit 370, further image and audio data 376, and further object data 378. The further image and audio data 376 represent a further image of the real scene acquired by the environment detection units 102, 103.The further object data 378 include at least a positioning of the object section within the further image of the real scene.
[0078] According to one embodiment, the positioning instruction 142 includes the display instruction for displaying the virtual image, which is contained within the virtual image and audio data 140. Alternatively, the display instruction can be transmitted separately to the control unit 372 in addition to the positioning instruction 142.
[0079] According to one embodiment, the selection device 368 is part of the mobile device 100. In this case, the external device 130 is not required and the external interface 360 can be implemented as an internal interface.
[0080] The in Fig. 3 The devices 360, 362, 366, 370, and 372 shown are only an exemplary arrangement of devices of a apparatus 379 for superimposing an image of a real scene with virtual image and audio data. To implement the process steps of a method for superimposing an image of a real scene with virtual image and audio data, some or all of the devices 360, 362, 366, 370, and 372 can, for example, be combined into larger units.
[0081] Fig. 4 Figure 1 shows a flowchart of a method for overlaying an image of a real-world scene with virtual image and audio data according to an exemplary embodiment. The method can be executed using the features of a mobile device as described in the preceding figures.
[0082] In step 480, image and audio data are read in, representing a representation of a real-world scene captured by the mobile device's environmental sensing devices. This image and audio data may have been acquired by the environmental sensing devices in an optional preceding step 482. In step 484, marker data is determined from the image and audio data, representing the image and position of a marker located in the real-world scene. Similarly, in step 486, object data is determined from the image and audio data, representing the image and position of a section of an object located in the vicinity of the marker in the real-world scene.In step 488, virtual image and audio data are read in. These data represent selected images and audio from a plurality of virtual image and audio data using the marker data and include a rendering instruction for displaying the virtual image, a positioning instruction for positioning the virtual image, and for playing the audio data. In an optional step 490, which can be executed on the mobile device or an external device, the virtual image and audio data are selected using the marker data. Using the marker data, the object data, and the virtual image and audio data, a positioning instruction is determined in step 492 that is suitable for displaying the virtual image and audio data in relation to the image of the object section, for example, as a superimposition of another image of the real-world scene.
[0083] In an optional step 494, the positioning instruction is used to represent the overlay of the further image of the real scene with the virtual image and audio data, for example on the display and playback device of the mobile device.
[0084] Step 494 can, for example, include a step 496 of reading in additional image and audio data representing a further representation of the real-world scene, a step 498 of determining the position of a further representation of the object segment from the additional image and audio data, and a step 499 of creating superimposed image and audio data using the additional image and audio data, the further representation of the object segment, and the positioning instruction, where the superimposed image and audio data represent a superimposition of the further representation of the real-world scene with the virtual image and audio data. In the step of determining the positioning, the positioning of the further visual and acoustic representation of the object segment within the further representation of the real-world scene can be determined.Thus, images of the object segment in successive images of the real-world scene can be used as anchor points for the virtual image and audio data. During the creation of superimposed image and audio data, the virtual image and audio data can be displayed using the rendering instruction.
[0085] Step 494 can be repeated continuously, using the positioning rule each time to overlay further images of the real scene with the virtual image and audio data. The preceding steps do not need to be repeated, as it is sufficient to define the positioning rule once.
[0086] According to one embodiment, in step 486, object data is determined from the image and audio data. This object data represents images and positions of multiple object segments, for example, two, three, four, or more object segments, of one or more objects located in the vicinity of the marker in the real-world scene. This allows the number of anchor points for anchoring the virtual image in the subsequent image(s) of the real-world scene to be increased. In this case, the positioning rule can be determined in step 492 such that it is suitable for representing the virtual image and audio data in the subsequent images of the real-world scene in relation to the optical and acoustic images of the object segments. To implement this representation, the positions of the individual images of the object segments are determined from the subsequent image and audio data in step 498.Advantageously, in this case, the virtual image and audio data can still be positioned according to the specifications stored in the virtual image and audio data even if not all images of the object sections are included by the other image and audio data.
[0087] According to one embodiment, the positioning instruction is determined in step 492 using a measured value from a measuring device, in particular an accelerometer, a rotation sensor, a magnetometer, a GPS receiver or one or more microphones of the mobile device.
[0088] Fig. 5 shows a QR code placement square 500 with binary contours according to an embodiment in which a QR code is used as a marker.
[0089] To recognize the QR code, the camera image being examined is first binarized, converting all pixels of the image into pure black or white values. Then, contours—that is, straight lines between black and white pixels—are searched for in the resulting image and filtered according to the visual properties of the three placement squares of a QR code. This results in a closed black contour (502) within a closed white contour (504), which in turn is a closed black contour (506).
[0090] Once the three placement squares 502, 504, 506 of the QR code have been found, the pixels between them are read and, according to the distribution of black and white pixels with a previously determined encoding, a bit sequence is determined, which in turn is converted into a string or URL.
[0091] In the next step, the position and orientation of the QR code relative to the camera are determined. For this, the Perspective-n-Point method "RANSAC," well-known in the literature, is used. Essentially, the camera is approximated with a simple pinhole camera model, assuming appropriate calibration. This allows the mapping of 3D points in the camera's real-world environment to their corresponding points in the 2D camera image to be described by a system of linear equations. This system is populated with the points of the three QR code placement squares in the camera image and extended with the known constraints regarding the relative positions of the squares, enabling its solution through linear optimization.
[0092] The following will be partly based on Fig. 1The reference symbols used are employed to further describe the process: Simultaneously, for example, at the precise moment of marker 112's detection (e.g., in the form of a code), the immediate surroundings of marker 112 are captured in camera image 108. Natural features 114 are extracted from this image, and a new natural marker 118 is created in real time according to the second method. For this purpose, the "SURF" (Speeded Up Robust Features) method, known from the literature, is used. This method stores features in two-dimensional objects in a transformation-invariant manner and can recognize them in subsequent images. The entirety of the features identified by SURF at the time of creation, as well as their relative positions, are stored as a single "marker." Additionally, the previously calculated position of the QR code within this image is stored in relation to this newly created marker.
[0093] In all subsequent camera images 120 and movements of camera 102 or marker 114, the three-dimensional position determination of the AR content 124 can now be carried out using the new, robust natural marker 114.
[0094] For this purpose, the SURF algorithm is applied again to each subsequent camera image, and the features found are compared with the previously stored features. If there is sufficient agreement, the previously stored marker, linked to the initial QR code, is considered recognized in the subsequent image. Furthermore, its position can be determined again using a Perspective-n-Point method (see above).
[0095] To display augmented reality, the data determined in this way regarding the position and orientation of the QR code are used, according to one example implementation, to transform the representation of virtual objects, which are available, for example, as 3D CAD models, and then to calculate a 2D representation of these objects using a virtual camera. In the final step, the transformed 2D view of the virtual object is superimposed onto the real camera image, thus creating the impression in the composite image that the virtual object is located directly on the QR code within the camera image of the real environment.
[0096] As the distance or rotation of the camera relative to the initially identified QR code increases, the positioning procedure described above can be repeated as often as desired to continuously create new "markers" in the real environment and store them along with their relative position to the QR code. This continuous iteration is known in the literature as "SLAM" (Simultaneous Location and Mapping). Depending on the expected scene (e.g., predominantly flat surfaces or uneven structures, glossy or rough materials, stationary or moving images), several other feature descriptors can be used in addition to the aforementioned SURF method to uniquely and reliably identify features.
[0097] Thus, a consistently stable display and movement and acoustically correct representation of three-dimensional virtual objects as virtual images 124 is possible, or, in contrast to geometric markers, these can still be tracked even if they are placed small and discreetly in the real scene 106.
[0098] Furthermore, the visual distinguishability of the newly created marker 114 is completely irrelevant compared to other markers, since its assignment to an AR content 124 has already been determined by the linked code, i.e., the marker 112. Directly extracting a URL from the linked code also avoids the need for constantly searching an online feature database and increases the number of distinguishable markers within an application to virtually unlimited levels. Moreover, unlike previous AR methods, creating the natural AR marker 114 immediately at the time of use allows objects 100 that frequently change their appearance, such as building facades at different times of day or seasons, to be used as natural markers 114.
[0099] One extension is the augmented reality display of objects for which no 3D CAD data exists, but only photos from different viewpoints.
[0100] The main problem here is that without 3D CAD data, no transformation of the virtual object can be performed, and conventional methods cannot be used to calculate a virtual 2D image that creates the impression of the virtual object's correct positioning in the real environment. As a solution, a method is presented here that achieves this impression solely based on previously taken photographs of an object, with the camera's viewing angle to the object known at the time of capture. For this purpose, the position and orientation of the QR code relative to the camera, determined as described above, are used: First, the image of the object whose viewing angle at the time of capture best matches the viewing angle of the augmented reality camera relative to the QR code is selected from the available images.Optionally, a new image is interpolated from several images, which corresponds even better to the viewing angle. This image is then scaled according to the distance of the QR code from the augmented reality camera and positioned according to the position of the QR code in the camera image, so that the composition of both images continuously creates the impression that previously photographed objects are standing in the environment subsequently viewed with the augmented reality camera.
Claims
1. Method of overlaying an optical and acoustic reproduction of a real scene with virtual three-dimensional or two-dimensional image and audio data, the method comprising the following steps: capturing (482) reproduction data using an environment sensor (102) and a further environment sensor (103) of a mobile device (100), wherein the reproduction data represent image data (364) which represent an image reproduction (108) of the real scene (106) captured by the environment sensor (102) of the mobile device (100), and wherein the reproduction data represent audio data (365) which represent an acoustic reproduction of the real scene (106) captured by the further environment sensor (103) of the mobile device (100); reading (480) the reproduction data; determining (484) marker data (132) from the image data (364), wherein the marker data (132) represent a reproduction (116) and a positioning of a marker (112) arranged in the real scene (106); reading (488) virtual image and audio data (140) which represent image and audio data selected from a plurality (252) of virtual image and audio data (140) using the marker data (132), only when a new artificial marker has been found as the marker (112), wherein the virtual image and audio data (140) comprise a representation instruction for representing a three-dimensional the defined object and / or a selection of captures of an object captured from various angles as a virtual reproduction (124), a positioning instruction for positioning the virtual reproduction (124) and a positioning instruction for replay of acoustic data; determining (486) object data (134) from the reproduction data (364), wherein the object data (134) consist of a three-dimensional reproduction (118) or a series of two-dimensional photographs and / or audio recordings from various angles and a positioning of object portion (114) of an object (110) arranged in the environment of the marker (112) in the real scene (106); ascertaining (492) a positioning rule (142) for representing the virtual reproduction (124) and the acoustic data with reference to the reproduction (118) of the object portion (114) using the object data (134) and the virtual image and audio data (140); wherein in the step of ascertaining (492), the positioning rule (142) is ascertained using the positioning of the reproduction (116) of the marker (112), the positioning of the reproduction of the object portion (114), and the positioning instruction, wherein the reproduction (118) of the object portion (114) is stored as a new natural marker for positioning the virtual reproduction (124) and the acoustic data in further reproductions (120) of the real scene (106) captured by the environment sensor (102) and is used in addition to or as an alternative to the marker (112); and a continuous iteration of the steps of reading (480, 488), of determining (484, 486) and of ascertaining (492) is performed several times per second in order to continuously create new natural markers, wherein the step of determining (484) is performed in order to recognize the previously stored new natural marker; using (494) the positioning rule (142) in order to overlay a further optical and acoustic reproduction (120) of the real scene (106) with the virtual image and audio data (124), wherein an overlay of the further optical and acoustic reproduction (120) of the real scene (106) with the virtual image and audio data (124) includes the further optical and acoustic reproduction (120) of the real scene, in which at least a portion is masked completely or in a semitransparent manner by the virtual image and audio data (124); displaying (498) the overlay of the further reproduction (120) of the real scene (106) with the virtual image and audio data (124) using a display device and a replay device (104 and 105) of the mobile device (100).
2. Method according to claim 1, wherein the image data (364) and / or the audio data (365) represent real image and audio data, the object data (134) represent real object data, and the object portion (114) represents a real object portion.
3. Method according to one of the preceding claims, wherein in the step (492) of ascertaining the positioning rule (142) is ascertained using the marker data (132) or at least part of the marker data (132).
4. Method according to one of the preceding claims, wherein in the step (492) of ascertaining the positioning rule (142) is ascertained using a measured value of a measuring device, in particular an acceleration sensor, a rotation sensor, a magnetometer or a GPS receiver, of the mobile device.
5. Method according to one of the preceding claims, comprising a step of providing at least part of the marker data (132) to an interface (360) to an external device (130), wherein in the step of reading (488) virtual image and audio data (140) the virtual image and audio data (140) are read via the interface (360) to the external device (130).
6. Method according to one of the preceding claims, wherein the marker (112) represents machine-readable code comprising an identification (138) of the marker (112), wherein the step of determining (484) marker data (132) the identification (138) of the marker (112) is determined as part of the marker data (132).
7. Method according to one of the preceding claims, wherein the step (494) of using comprises a step of reading (495) further image and audio data (376), wherein the further image and audio data (376) represent the further image (120) of the real scene (106) captured by the environment sensors (102) of the mobile device (100), a step (496) of determining a positioning of a further reproduction (122) of the object portion (114) from the further image and audio data (376), and a step of creating (497) overlaid image and audio data (374) using the further image and audio data (376), the positioning of the further reproduction (122) of the object portion (114) and the positioning rule (142), wherein the overlaid image and audio data (374) represent an overlay of the further reproduction (120) of the real scene (106) with the virtual image and audio data (124).
8. Method according to one of the preceding claims, wherein the reproduction (116) of the marker (112) takes up less than 1% of the reproduction (108) of the real scene (106).
9. Apparatus (379) for overlaying a reproduction of a real scene (106) with virtual image and audio data, wherein the apparatus (379) comprises devices for implementing the steps of the method according to one of the preceding claims.
10. Mobile device (100), in particular smartphone, comprising an apparatus (379) according to claim 9.
11. Computer program product with program code for performing the method according to one of the preceding claims 1 to 8, when the computer program product is executed on an apparatus.