Method and digital processing system for producing digitally enhanced camera images

The method and system address XR integration issues by using a camera vignette model to correct for vignetting, improving the quality and realism of augmented images by ensuring seamless transitions and accurate virtual content integration.

JP2025528601APending Publication Date: 2025-08-28ガリャセビッチ セナド
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025537293
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-07-14
Filing Date
2023-07-11
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

Existing extended reality (XR) technologies using green screens and LED walls face challenges in achieving seamless integration of virtual content with real-world footage due to lighting inconsistencies, color distortions, and vignetting effects, leading to visual artifacts.

Method used

A method and system that utilize a camera vignette model to correct for vignetting effects, allowing accurate determination of augmented regions and application of virtual content image data to match the real-world image, thereby reducing visual artifacts.

Benefits of technology

Improves the quality of digitally augmented camera images by ensuring seamless transitions and accurate integration of virtual content, reducing visual artifacts and enhancing realism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025528601000001_ABST
    Figure 2025528601000001_ABST
Patent Text Reader

Abstract

The present invention relates to a method, preferably a computer-implemented method, for creating digitally augmented camera images based on input images captured by a camera, comprising the steps of providing a camera vignette model, capturing at least one input image by the camera, determining an augmented region in the input image to be digitally augmented, providing virtual content image data for at least the augmented region, and generating a digitally augmented output image by augmenting the input image with the virtual content image data in at least the augmented region, wherein the determination of the augmented region and / or the provision of the virtual content image data is based on the camera vignette model. The present invention further relates to a corresponding digital image processing system.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method for creating a digital augmented camera image based on an input image taken by a digital camera according to the subject matter of claim 1 and to a corresponding digital image processing system according to the subject matter of claim 10. [Background technology]

[0002] Extended reality is an umbrella term covering a variety of technologies that combine real-world camera footage captured by a digital camera with virtual content image data, such as computer-generated virtual graphics. These technologies include virtual reality, augmented reality, and mixed reality. To achieve this effect, real camera footage must be combined with rendered virtual graphics or virtual content image data. The virtual content image data must be rendered with the correct perspective so that the combination of real camera footage and virtual content image data appears consistent. To ensure accurate rendering of the virtual content image data, the camera position and / or orientation must be determined with high spatial and temporal precision.

[0003] In XR, the goal is often to combine real-world camera footage with virtual content image data in a way that seamlessly blends the virtual elements together. This way, viewers can enjoy the added virtual elements while still feeling like everything they see is real. To achieve this goal, the virtual content image data must be rendered to match the real-world image data as closely as possible. For example, if a real-world camera uses optical lenses that distort the images it captures, the virtual content image data must also be rendered with distortion to mimic the effect of the real lens on the images it captures.

[0004] Virtual reality effects are achieved using so-called green screens. These are areas uniformly covered in a particular shade of green placed in the background of a scene being filmed by a camera. Green is the most common, although other colors have also been used. Specific shades or hues of the background color can be recognized in the filmed image and replaced with other video content. This allows virtual content image data to be inserted as a background for actors or objects filmed in the scene. This process involves selecting areas of the filmed image that are set transparent based on the frequency of a particular color or hue present in the background screen (e.g., a particular shade of green in the case of a green screen), allowing virtual content image data to be inserted into the scene. This process is called chromakeying.

[0005] A problem arising from the use of green screens and chromakeying is related to the color-based selection of image regions to be replaced with virtual content image data. In real-world scenarios, the green background in a camera image does not have a perfectly uniform tone or hue due to uneven lighting and color distortion caused by the camera itself. Given this, it is appropriate to allow a certain degree of latitude in the color selection for chromakeying. However, the degree of latitude must be carefully selected. If the color or hue selection is too narrow, not all areas to be augmented by the virtual content image data will be selected, resulting in green artifacts appearing in the augmented image. If the color or hue selection is too lax, image content may be unintentionally removed.

[0006] Another challenge with using green screens is adjusting the lighting so that the real image matches the virtual graphics. This is especially difficult when the graphics are changing over time. For example, if the actual real-world light source is stationary but the light source in the virtual graphics moves, the blend of real and virtual elements will look unnatural.

[0007] For these reasons, LED walls are preferred over green screens when producing high-quality virtual studios. An LED wall is a digital display made up of many individual light-emitting diodes (LEDs) arranged in a grid (i.e., an LED grid).

[0008] Extended reality technology using LED walls was recently popularized by Lucasfilm in the hit Star Wars series, "The Mandalorian," and has since rapidly grown in popularity. Especially with the COVID-19 pandemic, the film industry has had to rethink how movies are made and comply with safety regulations that reduce the number of people on set. This new XR technology is a modern alternative to green screen studios.

[0009] Problems arising from the use of LED walls are related to the limited dimensions of the LED wall. LED walls of satisfactory quality are quite expensive. Therefore, budgetary constraints often dictate the use of smaller LED walls, which do not fill the entire background of the scene captured by the camera, especially when the camera is positioned at a certain distance from the LED wall. In such cases, it is necessary to augment the captured image with virtual content image data in areas outside the LED wall to artificially stretch the LED wall graphics into areas of the camera image where the LED wall is not yet visible. If the colors of the virtual content image data do not exactly match the colors on the LED wall, undesirable artifacts will appear in the digitally augmented camera image at the boundaries of the LED wall. Summary of the Invention

[0010] From the above, it is clear that there is still a need for improvements in XR technology using green screens or LED walls. Therefore, the present invention aims to provide a solution for further improving the quality of digitally augmented camera images. In particular, the present invention aims to improve the realism of digitally augmented camera images and enable the reduction of visual artifacts in digitally augmented camera images.

[0011] The above object is solved by a method for producing a digital enhanced camera image according to the subject matter of claim 1 and by a digital image processing system according to the subject matter of claim 10. Preferred embodiments of the invention are defined by the subject matter of the dependent claims.

[0012] In particular, the above problem is solved by a method, preferably a computer-implemented method, for creating a digital augmented camera image based on an input image taken by a camera, the method comprising: a) providing a camera vignette model; b) capturing at least one input image with a camera; c) determining a dilated region in the input image to be digitally dilated; d) providing virtual content image data for (at least) the extension region; e) generating a digitally augmented output image by augmenting the input image with the virtual content image data, at least in the augmentation region; The determination of the extension area and / or the provision of the virtual content image data is solved by a method based on a vignette model of the camera.

[0013] The present invention is based on the surprising discovery that the quality of digitally enhanced camera images is significantly improved if the vignetting effect of the camera used to capture the image being enhanced is taken into account during image processing.

[0014] Enhanced quality of digitally enhanced camera images can be achieved in application scenarios where digital enhancement replaces image regions depicting background elements, such as a green screen. In such application scenarios, a vignette model can be used in the context of determining the augmented regions. Using the vignette model during the determination of the augmented regions improves the accuracy of the augmented region determination. This aspect of the present invention is particularly useful when used in the context of chromakeying. The term "background element" refers to an element that can be used to identify an augmented region in an input image through chromakeying, such as, but not limited to, a green screen. The term "background element" also encompasses a blue screen or other background element that can be used to identify an augmented region in an input image.

[0015] Another context in which the present invention provides significant improvements in digitally augmented images is a setting that includes an LED wall, where the input image does not fill the entire background of the image. Determining the augmented region in such a case is relatively straightforward; any region of the input image that does not depict either the LED wall or a foreground object, such as a person, can be determined to be the augmented region. The image displayed on the LED wall can then be augmented by overlaying corresponding virtual content image data onto the augmented region. In this way, the image displayed on the LED wall and captured in the input image is virtually stretched or augmented by the virtual content image data.

[0016] A problem that arises in this scenario, and that is solved by the present invention, is the difficulty of ensuring a seamless transition between the input image region where the LED wall was photographed and the virtual content image data. Surprisingly, we have found that correcting the virtual content image data based on a camera vignette model can significantly reduce the transition error between the region of interest and the digitally augmented image region with the virtual content image data.

[0017] Because the image content on the LED wall captured in the input image is subject to camera photography aberrations and errors, modifying the virtual content image data based on the camera's vignette model (thereby virtually subjecting it to the same photography errors as the data in the input image) can significantly improve the quality of the digitally augmented camera image. In this scenario, the vignette model is preferably used to apply a vignette effect to the virtual content image data, such that the virtual content image data is subject to the same vignette effect as the rest of the content of the input image that was captured by the camera and that is not replaced by the virtual content image data.

[0018] Vignetting refers to a decrease in brightness or saturation of an image toward the periphery that occurs in all lenses and lens systems used in cameras. The vignetting effect varies depending on the type of lens used and can also vary with various camera settings such as aperture. The image effect caused by vignetting, or the vignette effect, is characterized by a decrease in brightness in areas of the image away from the center of the image and a darker border around the image.

[0019] To allow for the vignetting effect in the creation of digitally enhanced images, the present invention utilizes a vignette model. The vignette model is a function that outputs, for a given 2D point coordinate in an image, a reduction in brightness or saturation relative to the center of the image at this particular 2D point coordinate. The vignette model may be configured by a radial function indicating the reduction in brightness or saturation depending on the (radial) distance from the center of the image. Alternatively, the vignette model may be provided as a look-up table in which the reduction in brightness or saturation is provided for each pixel of the image frame. The reduction in brightness or saturation may be represented, for example, by a value between 0 and 1, where 0 indicates that the brightness or saturation is reduced completely to zero and 1 indicates that the brightness or saturation is not reduced (e.g., in the case of the center of the image frame).

[0020] The vignette model preferably depends on the aperture size of the camera: the smaller the aperture of the camera, the less noticeable the vignetting effect. Therefore, the vignette model preferably includes a set of multiple sub-models corresponding to different aperture sizes of the camera.

[0021] In the context of the present invention, the camera is preferably configured as a digital camera with an image sensor for capturing images in a digital memory, which allows image enhancement in real time. However, the present invention can also be implemented with an analog camera, where the captured images are digitized. Therefore, in the context of the present invention, the term "camera" refers to any camera system suitable for providing a digital input image that can be enhanced with digital image content.

[0022] An input image taken by a (digital) camera is typically provided as a digital image made up of pixels, each pixel having a finite, discrete numerical representation for its intensity or grey level. These quantities are typically output as spatial coordinates, denoted x and y on the x-axis and y-axis, respectively. In the context of this disclosure, the 2D region spanned by the spatial coordinates representing the input image is also referred to as the image frame.

[0023] In the context of the present invention, two basic types of image content are distinguished in the input image: The input image contains regions depicting objects or scenes of interest, such as a person standing in front of a green screen or a person and an object located in front of an LED wall (in this case both the person and the object, as well as the image depiction of the LED wall, may belong to the regions of interest). These regions in the input image may hereinafter also be referred to as regions of interest. In addition to these regions of interest, the input image also contains augmented regions, which are replaced by image information not contained in the input image. In the case of a green screen, the augmented regions essentially correspond to the green regions of the input image that depict the green screen. In the case of an LED wall scene, the augmented regions include areas outside the LED wall where neither the LED wall nor the person or object of interest is visible in the input image.

[0024] The input image is modified by extending the input image with virtual content image data, at least in the extension region. The virtual content image data is not particularly limited. In a green screen application scenario, the virtual content image data may include a background image, such as a landscape where a person in front of the green screen appears to be located. The virtual content image data may also include information for extending the input image, such as a map or text information for a weather forecast. In an LED wall application, the virtual content image data may be image data that complements the image data displayed on the LED wall, thereby virtually extending the image displayed on the LED wall beyond the boundaries of the LED wall.

[0025] By combining the virtual content image data (at least in the augmentation region) with the image data contained in the region of interest, a digitally augmented output image is obtained. Various methods for combining the virtual content image data with other image information in the input image are available and known in the art. For example, after the augmentation region is determined, the corresponding pixels in the input image can be set to transparent. The resulting modified input image can then be overlaid on the image containing the virtual content image data. To smooth the transition between the virtual content image data in the augmentation region and the input image information in the region of interest, the virtual content image data can be provided to partially overlap the region of interest, and the overlapping region can be overlaid with a transparency gradient.

[0026] The method for creating digitally enhanced camera images according to the present invention is preferably computer-implemented. The method can be performed on a dedicated computer to which the input image(s) taken by the camera are provided. Alternatively, the method can be implemented on a microcomputer included in the camera or a separate device.

[0027] It should be noted that this method is not strictly limited to the order of the steps described above. These steps may be performed partially and simultaneously. Step d) may be performed before or simultaneously with any of steps a) to c).

[0028] Providing the virtual content image data may be accomplished by receiving pre-compiled virtual content image data from a data source and / or by creating or rendering the virtual content image data. Providing the virtual content image data based on a vignette model may include applying a vignette effect to the virtual content image data based on or from the vignette model.

[0029] According to a preferred embodiment, the step of determining the extended region in the input image includes correcting the input image based on a vignette model. Preferably, the step of determining the extended region in the input image includes chromakey processing. More preferably, the step of correcting the input image based on the vignette model is performed before determining the extended region in the input image.

[0030] This embodiment of the present invention is particularly advantageous when the selection of the enhancement region is based on image information in an input image distorted by a vignetting effect. A typical example is chromakeying, in which the enhancement region in the input image is determined based on a predetermined color and / or hue. The vignetting effect darkens peripheral areas of the input image. To accurately account for these areas in the chromakeying process, the selection tolerance must be increased to account for the color and / or hue changes caused by the vignetting effect. If the vignetting effect in the input image is corrected and / or compensated for based on a vignetting model before chromakeying, the color and hue of the peripheral areas of the image where the vignetting effect is most pronounced can be corrected before keying. This allows the selection tolerance to be reduced during chromakeying. Surprisingly, it has been found that correcting the vignetting effect in the input image based on a camera's vignetting model can significantly improve the accuracy of the chromakeying process, thereby significantly improving the quality of the subsequently digitally enhanced camera image. This significantly reduces visual artifacts in the digitally enhanced camera image.

[0031] Particularly preferably, the step of determining the extension region comprises: providing a 3D model of a background element, such as a green screen, that is photographed in the input image; capturing a plurality of reference images of the background element with a camera from a plurality of camera positions and / or camera poses; correcting the reference image based on a vignetting model to obtain a devignetted reference image; assigning color information of the devignetted reference image to the 3D model of the background element to obtain a 3D reference model of the background element; and determining, as an augmented region, a region in the input image in which the background element is captured based on a comparison between the input image and a 3D reference model of the background element.

[0032] To improve the sensitivity and selectivity of the chromakeying process for selecting areas in the input image where the green screen is visible as an augmented region, it is preferable to obtain the color information required for the chromakeying process from an actual camera image rather than defining theoretical color and / or hue values ​​in the chromakeying process. This directly takes into account color distortions caused by the camera's optics and / or image sensor. In this context, it may be further advantageous to provide a 3D model of the background element defining the augmented region (e.g., the green screen) in the input image and use a reference image of the background element captured by the camera to create color information assigned to the 3D model of the background element. If the color and / or hue of the background element is not perfectly uniform, for example, due to uneven lighting or contamination of the background element, such unevenness can be taken into account by constructing a 3D model and assigning color information obtained from the actual camera image to the 3D model.

[0033] One problem with this approach is that the color and / or hue of a region in the background image may vary depending on the region's location in the reference image. For example, if a corner of a background element is captured in a central region of a first reference image, the corner may appear brighter due to a vignetting effect than in a second reference image in which the corner is captured in a peripheral region. If color information from the first reference image is assigned to a 3D model of the corner of the background element and the input image contains the corner in a peripheral region (or vice versa), using the color information in the 3D model for chromakeying may result in poor quality due to color and / or hue mismatches between the 3D model and the input image.

[0034] Therefore, it is preferable to correct the reference image based on the vignette model to improve the accuracy of the color information assigned to the 3D model of the background element and ultimately used in the chromakey process to determine the augmented region in the input image. In this context, it is even more preferable to also correct the input image based on the vignette model before performing the chromakey process. This further improves the quality of the augmented image. Alternatively, if the position and orientation of the camera relative to the background element are known, the 3D reference model of the background element can be used to create a virtual input image corresponding to the (actual) input image captured by the camera. A vignetting effect based on the vignette model is applied to the virtual input image, and the resulting modified virtual input image is compared with the (actual) input image to determine the augmented region in the input image.

[0035] The color information may be provided as RGB, YUV, or CYMK values, or in other formats suitable for identifying the color and / or hue of a background element, such as a green screen for chromakeying.

[0036] The 3D (reference) model of the background element preferably includes 3D (Cartesian) coordinates of at least some of the partition features of the background element. If the background element consists of a rectangular screen, for example, the 3D (reference) model may include at least the (Cartesian) coordinates of the corners of the screen.

[0037] When capturing an input image including a background element, the position and orientation of the camera are preferably determined. This allows the camera's viewpoint relative to a 3D (reference) model of the background element to be determined. This information allows the determination of portions of the 3D (reference) model of the background element that correspond to areas captured in the input image, which can be used for further processing of the input image, in particular chromakey processing. A camera tracking system can be employed to determine the position and orientation of the camera. Suitable camera tracking systems include systems based on odometry, GPS, or optical tracking methods such as SLAM. Optical tracking systems suitable for determining the position and orientation of the camera are described, for example, in EP 3 985 609 A1.

[0038] The camera position indicates the camera's absolute 3D position in space. The camera pose indicates the combination of the camera's position and orientation, which can be described by a vector parallel to the camera's optical axis and pointing in the camera's viewing direction.

[0039] According to a further preferred embodiment, the vignette model is used in the step of creating the virtual content image data by applying a vignette effect to the virtual content image data based on the vignette model, whereby the step of creating the virtual content image data is based on the vignette model in the sense that the vignette effect is applied to the virtual content image data using the vignette model.

[0040] This embodiment is particularly useful for digitally augmenting an input image capturing a scene with an LED wall, especially when the LED wall does not fill the background of the entire input image. To virtually augment the scene displayed on the LED wall with virtual content image data, a region in the input image that does not include the LED wall (a region that does not include other objects of interest, such as a person positioned in front of the LED wall) can be determined as the augmented region. Because the LED wall captured in the input image is subject to a camera vignette effect, applying the vignette effect to the virtual content image data based on the camera's vignette model ensures that the brightness and / or saturation of the virtual content image data in the augmented region of the output image matches the brightness and / or saturation of the remaining image region, thereby preventing abrupt transitions from appearing in the output image between the virtual content image data and other regions of the image.

[0041] Here, preferably, the step of determining the extension region includes: providing a 3D model of an object of interest captured in an input image; establishing a camera position and pose; calculating an image position of the object of interest in the input image based on a 3D model of the object of interest and the established camera position and pose; determining an extension region based on the image location of the object of interest in the input image.

[0042] The object of interest is preferably a stationary object near the camera, and may consist of, for example, one or more LED walls. Given a 3D model of the object of interest, regions in the input image that do not depict the object of interest can be determined with high reliability, accuracy, and speed. The position and orientation of the camera at the time the input image was captured are established, and the position of the object of interest in the input image can be calculated based on the 3D model of the object of interest and the position and orientation of the camera.

[0043] The 3D model of the object of interest preferably includes 3D (Cartesian) coordinates of at least some local features of the object of interest. If the object of interest consists of a rectangular screen, such as an LED wall, for example, the 3D model may include the (Cartesian) coordinates of at least the corners of the screen.

[0044] The position and pose of the camera can be determined by any suitable tracking means, as described above. In a further preferred embodiment, the virtual content image data is subjected to a lens distortion effect corresponding to the lens distortion imparted by the camera before generating the digitally augmented output image, thereby further improving the quality of the digitally augmented output image.

[0045] In a further embodiment, the vignette model is applying a devignetting effect based on a vignetting model to the input image, preferably at least to an area outside the extended area of ​​the input image; The virtual content image data is used in the step of creating the virtual content image data by combining it with a devignetted input image, preferably with a region of the input image that is outside the extended region and to which the devignetting effect has been applied.

[0046] In addition, in this embodiment, by devignetting image data in the input image, the abrupt transition between the input image data (vignetted by capturing it with a camera) and the virtual content image data is mitigated or eliminated, thereby improving the image quality of the digitally enhanced camera image.

[0047] According to a preferred embodiment, the vignette model includes a function that maps a vignette value, indicative of the reduction in brightness and / or saturation due to the camera's vignetting effect, to each pixel location in the camera's image frame. The mapping may be provided in a look-up table, in which each pixel in the camera's image frame is assigned a value indicative of the reduction in brightness and / or saturation. This allows for the provision of a conceptually simple and computationally easily accessible vignette model.

[0048] The values ​​of the vignette model, which indicate the reduction in brightness and / or saturation due to the camera's vignetting effect, may be provided as numbers on a predetermined scale. For example, a value of 0 in the vignette model may indicate that the corresponding pixel is completely covered by the vignetting effect and always has a brightness of 0, i.e., is completely black. The values ​​may also be normalized. For example, a value of 1 in the vignette model may indicate that the corresponding pixel has not experienced any reduction in brightness and / or saturation due to the vignetting effect, or that it is the pixel with the least reduction in brightness relative to other pixels in the input image frame.

[0049] When the vignette model uses values ​​between 0 and 1, the applicability of the vignette model is facilitated. The vignette model can be used to apply a vignette effect to image data, particularly virtual content image data, by, for example, multiplying the luminance and / or saturation values ​​of the pixel by the corresponding values ​​of the vignette model. Conversely, image data, particularly input image data, can be de-vignetted by dividing the luminance and / or saturation values ​​of the pixel by the corresponding values ​​of the vignette model.

[0050] According to a further embodiment, the vignette model includes a function that maps a vignette value, which indicates a reduction in brightness and / or saturation due to the camera's vignetting effect, to a distance relative to the center of the camera's image frame. This vignette model requires less storage space than a pixel-based vignette model, while still providing a good representation of the camera's vignetting effect. The vignette value is provided as a numerical value on a predetermined scale, as in the previous embodiment, and may be further normalized. Preferably, the vignette value uses a value between 0 and 1.

[0051] When image data is modified using this type of vignette model, the distance between the pixel coordinate to be modified and the center of the image is calculated to obtain the corresponding vignette value. The geometric center of the image frame may be used as the image center. If the image frame has a resolution of W × H pixels in the x and y directions, the image center may be determined as (W / 2; H / 2). Alternatively, to account for asymmetric or off-axis alignment of the camera's lens system, the image center may be offset relative to the geometric center of the image frame. In this case, the center of the camera's image frame preferably includes an offset value indicating the offset between the geometric center of the camera's image frame and the center of the vignette model where the minimum or maximum vignette value occurs. The offset value preferably comprises a 2D vector indicating the offset in the x and y directions in the image frame.

[0052] The object of the present invention is also to provide a digital image processing system for creating a digital augmented camera image based on an input image taken by a camera, preferably using a method as described above, comprising: Receives an input image from a camera; determining a dilated region in the input image to be digitally dilated; providing virtual content image data for (at least) the augmented region; a processing unit configured to generate a digitally augmented output image by augmenting the input image with virtual content image data in at least the augmented region; Remember the camera's vignette model, a modeling unit configured to provide a vignette model to the processing unit; The processing step is solved by a digital image processing system configured to determine the extension region based on a vignette model of the camera and / or to provide the virtual content image data.

[0053] The technical advantages achieved by the digital image processing system of the present invention correspond to the technical advantages achieved by the method of creating digital augmented camera images described above. Aspects, features, and advantages described in the context of the method of the present invention are also applicable to the image processing system of the present invention, and vice versa. In particular, any task that a processing or modeling unit of the image processing system of the present invention is configured to perform can be performed as a method step in the method of the present invention, and vice versa. Preferably, the method of the present invention is implemented using the image processing system of the present invention as described above.

[0054] The processing unit and the modeling unit may be implemented as separate physical entities or as software and storage modules contained and executed on a common computing unit.

[0055] According to a preferred embodiment, the processing unit is configured to correct the input image based on a vignette model when determining the extended regions in the input image.

[0056] The processing unit is configured to use chromakey processing when determining the extended region in the input image, and preferably is further configured to correct the input image based on a vignette model before determining the extended region in the input image.

[0057] More preferably, the processing unit Provides 3D models of background elements, such as green screens, captured in the input image, receiving a plurality of reference images of a background element captured by a camera from a plurality of camera positions and / or camera poses; correcting the reference image based on a vignetting model to obtain a devignetted reference image; assigning color information of the devignetted reference image to the 3D model of the background element to obtain a 3D reference model of the background element; The corrected input image is compared with a 3D reference model of the background element to determine an area in the input image where the background element is captured as the augmented area.

[0058] According to a further preferred embodiment, the processing unit is configured to create the virtual content image data using the vignette model by applying a vignette effect to the virtual content image data based on the vignette model.

[0059] More preferably, the processing unit Providing a 3D model of the object of interest captured in the input image; Establishing the camera position and orientation; Calculating an image position of the target of interest in the input image based on a 3D model of the target of interest and the established camera position and orientation; The device is configured to determine the extension region based on the image position of the object of interest in the input image.

[0060] The object of the present invention is further solved by a computer-readable medium comprising instructions, which, when executed by at least one processor, cause the at least one processor to perform a computer-implemented method, including the steps according to the above-described inventive method. The technical advantages achieved by the inventive method correspond to the technical advantages achieved by the above-described computer-implemented method and the corresponding computer-readable medium. Aspects, features, and advantages described in the context of the inventive method are also applicable to the inventive computer-implemented method and computer-readable medium. [Brief explanation of the drawings]

[0061] The above and further features and advantages of the present invention will become more readily apparent from the following detailed description of preferred embodiments of the invention, taken in conjunction with the accompanying drawings. [Figure 1] 1 is a schematic diagram of a camera with a digital image processing system for capturing a scene with a person and an LED wall according to one embodiment of the present invention; [Figure 2] FIG. 2 is a schematic diagram of an input image including a region of interest and an extended region, captured using the settings of FIG. 1. [Figure 3] FIG. 1 is a schematic diagram of a camera with a digital imaging system according to one embodiment of the present invention for capturing a scene with two people and a green screen. [Figure 4] FIG. 4 is a schematic diagram of an input image including a region of interest and an extended region, captured by the camera of FIG. 3. [Figure 5] FIG. 1 is a schematic diagram of a vignetting effect in an image frame. [Figure 6] 1 illustrates the radial coordinates used in the vignette model according to a preferred embodiment of the present invention. [Figure 7] FIG. 1 is a schematic diagram of a setup for creating a vignette model for a camera. [Figures 8a-8d] FIG. 10 is a schematic diagram of an image taken during creation of a vignette model. [Figure 9] FIG. 1 is a schematic diagram of an image frame showing the image position of a visual reference target for creating a vignette model. DETAILED DESCRIPTION OF THE INVENTION

[0062] FIG. 1 shows a schematic diagram of a scenario in which the method for producing digitally enhanced camera images and the digital image processing system according to the present invention can be used.

[0063] 1 shows an image processing system 3 according to one embodiment of the present invention. The image processing system 3 includes a processing unit 31 and a modeling unit 32, which will be described in detail below.

[0064] The image processing system 3 is connected to the camera 2. This may be constituted by a separate unit or may be included in the electronics of the camera 2. The camera 2 is provided to capture a scene including the LED wall 11. The LED wall 11 constitutes an object of interest that is included in the image captured by the camera 2 and that is intended to be at least largely retained in the camera image after image processing performed by the image processing system 3. Another object of interest, such as a person 12, may be located in front of the LED wall 11.

[0065] The camera 2 may be a photo camera for capturing still images, or more preferably a cinema camera for capturing a cinema sequence consisting of a plurality of sequential images. The area captured by the camera 2 is indicated by a dotted line in FIG.

[0066] The scenario shown in Figure 1 may occur when filming a movie scene in which a person 12 acts in front of an LED wall 11 that displays the background of the scene, or may occur in a television production studio in which a person 12 is an anchorman delivering a news or weather report in front of an LED wall 11 that displays additional information such as a weather map or a virtual studio background.

[0067] FIG. 2 is a schematic diagram of an image captured by camera 2 in the setup shown in FIG. 1. The area captured by camera 2, indicated by the dotted line in FIG. 1, is included in the image frame of camera 2, resulting in an input image as shown schematically in FIG. 2. The input image is preferably provided as a pixel-based digital image. In the input image, an LED wall 11 and a person 12 in front of it are visible. Both the LED wall 11 and the person 12 constitute objects of interest that are at least largely retained in the input image after digital processing. The entire input image area covered by the objects of interest may be referred to as area of ​​interest 1, and is delimited by a dash-dotted line in FIG. 1.

[0068] Figure 2 illustrates a scenario in which part of person 12 (his feet) is not included in the region of the input image bounded by LED wall 11. This part may be omitted from region of interest 1 for simplicity. Alternatively, this part may be automatically determined, for example by a suitable AI algorithm, and included in region of interest 1, as shown in Figure 2.

[0069] The input image of FIG. 2 further includes regions that do not depict objects of interest, such as the LED wall 11 or the actor 12, i.e., regions outside the region of interest 1. These cross-hatched regions in FIG. 2 may be designated as augmented regions 4. The augmented regions are augmented with virtual image content that is not visible in the input image of FIG. 2. For clarity, the entire augmented region 4 in the input image is referred to as augmented region 4.

[0070] In the case of a movie scene, the extended area 4 may be extended with scenery images that complement the scenery displayed on the LED wall 11. In the case of a television production studio, it may be desirable to display virtual studio content in the extended area 4.

[0071] FIG. 3 illustrates a further application scenario of the present invention. As in FIG. 1, a camera 2 equipped with an image processing system 3 is provided. In the scenario of FIG. 3, the camera 2 is used to capture a scene with two people 12 positioned in front of a background object 41. The background object 41 may consist of a green screen. FIG. 4 illustrates an input image captured by the camera 2 of FIG. 3. The input image includes a depiction of two people 12 in front of a background element 41. The people 12 constitute a region of interest 1 in the input image. The area covered by the background element 41 constitutes an augmented region 4 of the input image. This augmented region 4 may be overlaid with virtual content image data to replace the (usually uniform) background provided by the background object 41 with a desired background, such as a movie scene or a virtual studio background.

[0072] The digital image processing system 3 shown in Figures 1 and 3 comprises a processing unit 31 configured to receive an input image from a camera 2 shown in Figure 2 or 4. The processing unit 31 is further configured to determine an augmentation region 4 in the input image. Furthermore, the processing unit 31 is configured to receive virtual content image data to be included in the digital augmented output image, at least in the augmentation region 4. The virtual content image data may be provided from an external source, such as a computer system, or may be generated by the processing unit 31, the modeling unit 32, or other components of the image processing system 3. The processing unit 31 is configured to generate the digital augmented output image by augmenting the input image with the virtual content image data, at least in the augmentation region 4.

[0073] In conventional image processing methods, an augmented region 4, such as that shown in Figures 2 and 4, is identified in the input image, which is then overlaid with virtual content image data in at least the augmented region 4 to generate a digital augmented output image. The present invention improves upon conventional image processing methods by taking into account the vignetting effect of the camera 2 in the image processing process to obtain a digital augmented output image of improved quality.

[0074] In the context of the present invention, the vignetting effect of camera 2 is taken into account by providing a vignette model that models the vignetting effect of camera 2. This allows for correcting the vignetting effect in the (input) image captured by camera 2 to obtain a de-vignetted camera image and / or for applying the vignetting effect to virtual content image data, as will be explained in more detail below.

[0075] The vignette model is stored in the modeling unit 32. It quantifies the reduction in brightness and / or saturation at each position within the image frame. The vignette model of the camera 2 is configured to quantify the reduction in brightness and / or saturation at each position within the image frame of the camera 2, which is affected by the limited aperture of the camera 2, particularly its lens system. FIG. 5 schematically illustrates a typical vignetting effect that occurs in a camera 2 with a lens system. As indicated by the increasing hatching density in FIG. 5, areas up to a certain radial distance around the center of the image frame may not exhibit significant vignetting, while areas around the periphery of the image frame gradually decrease in brightness and / or saturation. Note that the camera's lens system may also cause a vignetting effect, with areas with a stronger vignetting effect located in the center of the image frame and the vignetting effect becoming weaker as the image moves away from the center.

[0076] The vignette model may be configured by a lookup table that contains, for each pixel in the image frame, a vignette value that quantifies the reduction in brightness and / or saturation. To determine the vignette value for a location in the image frame of camera 2, the corresponding pixel coordinate is identified in the lookup table and the corresponding vignette value is looked up.

[0077] Alternatively, the vignette model can provide a (scalar) function that indicates the vignette value as a function of the (radial) distance d from the image center c of the image frame. Figure 6 shows a schematic representation of such a vignette model. The image center c is depicted along with three exemplary image (frame) positions p1, p2, and p3. To determine the vignette values ​​for image frame positions p1, p2, and p3, their respective distances d1, d2, and d3 to the image frame center c are determined, and the corresponding vignette values ​​are determined from the scalar function included in the vignette model. Image position p1 has the smallest distance d1 to center c, resulting in the smallest vignetting effect. The most vignetting effect is achieved at image position p3, which is located in the lower-left corner of the image frame and has the greatest distance d3 to the image center c.

[0078] The vignette model may consist of a linear function, an nth order polynomial function, or any other function suitable for modeling the vignetting effect of camera 2. The asymmetric alignment of the lens system of camera 2 with respect to the image sensor can be taken into account by an offset of the frame center c that can be provided in the vignette model.

[0079] Preferably, the vignette value is a number in the interval [0;1], where 0 indicates that the corresponding image location is completely obscured and 1 indicates no reduction in brightness (or minimal reduction in brightness relative to other image locations).

[0080] The vignette model stored in the modeling unit 32 can be used in a variety of ways to improve the quality of the digitally enhanced output image.

[0081] 1 and 2, determining the augmented region 4 is relatively straightforward. In this scenario, it is preferable to modify the virtual content image data using a vignette model to reduce visual artifacts in the digitally augmented output image.

[0082] To provide a high-quality digitally augmented output image, it is necessary to ensure that the virtual content image data matches the brightness and / or saturation of the image data in the input image's region of interest 1. If the brightness and / or saturation of the virtual content image data differs from the brightness and / or saturation of the image data in the region of interest 1 in the input image, visible seams will result in the boundary area between the region of interest 1 and the augmented region 4.

[0083] Due to the vignetting effect of camera 2, significant discontinuities in brightness and / or saturation may occur if the input image of FIG. 2 is extended with virtual content image data in extension region 4 without properly considering the vignetting effect of camera 2. The image data contained in the input image will clearly be subject to the vignetting effect of camera 2 due to the input image being captured by camera 2.

[0084] To seamlessly complement the input image with the virtual content image data, the processing unit 31 is configured to receive a vignette model from the modeling unit 32. The processing unit 31 is configured to create virtual content image data for enhancing the input image by applying a vignette effect based on the vignette model to the virtual content image data. Here, a designated position of each pixel of the virtual content image data in the image frame can be determined, and the brightness and / or saturation value of each pixel of the virtual content image data can be modified based on the vignette model. When a vignette model with a vignette value of 0 to 1 is used, as described above in detail, the modification based on the vignette model can include multiplying the brightness and / or saturation value of each pixel of the virtual content image data by the corresponding vignette value.

[0085] By applying a vignetting effect to the virtual content image data based on the vignette model of camera 2, the virtual content image data is modified to appear as if it had been captured by camera 2. Subsequent augmentation of the input image with the modified virtual content image data results in a digitally augmented output image of improved quality, as visual artifacts at the boundary between the region of interest 1 and the augmented region 4 are suppressed.

[0086] In the scenarios shown in Figures 3 and 4, a vignette model of camera 2 may be used in an alternative or additional processing step to improve the quality of the digitally augmented output image. The main challenge in the scenarios of Figures 3 and 4 is to reliably determine the augmented region 4. Because the region of interest 1 consists of moving objects, such as the person 12 in Figures 3 and 4, the augmented region 4 needs to be dynamically determined for each input image. In the case of Figures 3 and 4, the background elements 41 have a nearly uniform color. Therefore, chromakey processing is used to determine the augmented region 4 in the input image.

[0087] To improve the reliability and accuracy of the chromakey processing, it is necessary to minimize color artifacts generated by the camera 2 in the input image. For this reason, a vignette model is used to determine the extended region 4. Specifically, the input image is corrected based on the vignette model. In other words, the input image is corrected by performing de-vignetting based on the vignette model. When a vignette model with a vignette value between 0 and 1 is used, as described above, the correction may include multiplying the luminance and / or saturation value of each pixel in the input image by the inverse of the corresponding vignette value.

[0088] The vignette model can be used to further refine the chromakeying setup. For this additional refinement, the camera 2 is preferably equipped with a tracking system 21, as shown in Figure 3, configured to establish the position and pose of the camera 2 when the image is taken. Furthermore, a 3D model of the background element 41 is provided, e.g. stored in the modeling unit 32.

[0089] The 3D model of the background element 41 is used to create a virtual reference of the background element 41 for use in chromakeying. Thus, the camera 2 is used to capture multiple reference images of the background element 41 at multiple positions and / or orientations of the camera 2. For each reference image, the position and orientation of the camera 2 is determined by the tracking system 21. The position and orientation of the camera 2 relative to the background element 41 is determined based on the 3D model of the background element 41 and the position and orientation of the camera 2 established by the tracking system 21. The 3D model of the background element 41 is then supplemented with color information by assigning color values ​​captured in the reference images to corresponding positions in the 3D model of the background element 41 to obtain a 3D reference model of the background element 41.

[0090] To improve the accuracy of the color information in the 3D model, the reference image is devignetted before assigning color information to the 3D model. The devignetting of the reference image may be performed in a similar manner to the devignetting of the input image, as described above. By devignetting the reference image before assigning color information to the 3D model of the background object 41, the accuracy of the color information in the 3D model is significantly improved.

[0091] The resulting 3D reference model of the background element 41 can be used for an improved chromakeying process. For each input image, the corresponding position and orientation of the camera 2 is established. In combination with the 3D reference model of the background element 41, the portion of the background element 41 captured in the input image can be determined. The input image may be further devignetted. The chromakeying process is then performed by comparing the input image (preferably devignetted) with the corresponding portion of the 3D reference model of the background object 41.

[0092] As an alternative approach, the vignette model can be used in a different way to improve the accuracy of the chromakeying process. When the camera 2 captures an input image, the position and orientation of the camera 2 are determined (using the tracking system 21). The position and orientation of the camera 2 relative to the background element 41 are determined based on the 3D model of the background element 41 and the position and orientation of the camera 2 established by the tracking system 21. Using the 3D reference model, the relative position and orientation information of the camera 2, and the 3D model of the background element 41, a virtual input image of the background element 41 is created, including the portion of the background element 41 captured in the actual input image. Next, the virtual input image is modified by applying a vignette effect to the virtual input image based on the camera's vignette model. This modified virtual input image is then used as a reference for chromakeying the (unmodified) input image.

[0093] The tracking system 21 described for use in the scenarios of Figures 3 and 4 can also optionally be used in the scenarios of Figures 1 and 2 to improve the determination of the augmented region 4. In this case, the image processing system is provided with a 3D model of the object of interest 1 of Figure 1. When capturing an input image, the tracking system 21 establishes the corresponding position and orientation of the camera 2. Based on the 3D model of the object of interest 1 and the established position and orientation of the camera 2, the image position of the object of interest 1 in the input image shown in Figure 2 can then be calculated. The augmented region 4 can then be determined based on the calculated image position of the object of interest 1 in the input image.

[0094] The vignette model preferably takes into account different aperture sizes of the camera 2. Thus, the vignette model preferably includes multiple vignette models for different aperture sizes of the camera 2.

[0095] The vignette model may be provided by the manufacturer of the lens system and / or camera 2, or may be established for a given camera 2. Figure 7 shows a setup for creating a vignette model for a camera 2. The camera 2 is aimed at a visual reference target 5 consisting of a limited area of ​​uniform color. Preferably, the visual reference target 5 is a regular surface of uniform color, for example a white or green rectangle.

[0096] The camera 2 is mounted so that it can rotate freely, as shown by the three main rotation axes in Figure 7. The camera 2 is then rotated in different orientations and used to capture multiple images in which the visual reference target 5 appears in different regions of the field of view of the camera 2. This is shown schematically in Figures 8a to 8d, which show exemplary images with the visual reference target 5 at different positions in the image.

[0097] To establish a vignette model for camera 2, the image position of the visual reference target 5 is determined in each image. For example, the geometric center of the visual reference target in each image is determined as the image position of the visual reference target 5. Furthermore, for example, the luminance or saturation value of the visual reference target 5 is determined in each image by averaging the luminance or saturation value of the entire region of the visual reference target 5 in each image. Next, a vignette model for camera 2 is established based on the determined luminance or saturation value and the image position of the visual reference target 5 in the captured image.

[0098] To improve the quality of the resulting vignette model, the position of the visual reference target 5 in the image can be systematically varied. This is shown diagrammatically in FIG. 9. First, the camera 2 is oriented so that the visual reference target 5 appears in the center of the image frame. Next, the operator of the camera 2 is informed of the position in the field of view of the camera 2 where the visual reference target 5 should be placed. The operator then moves the camera 2 accordingly, preferably by rotating it around its point of view, to position the visual reference target 5 at the informed position in the field of view of the camera 2 before each of the multiple images is taken. In this way, a sequence of images for the calculation of the vignette model is systematically established.

[0099] For each image, the location of the visual reference target 5 and its luminance or saturation in each image are determined. The location of the visual reference target 5 in each image is preferably indicated by its distance from the center of the image. A vignette model can then be established by interpolation between associated value pairs of distance from the center of the image and corresponding luminance or saturation values. [Explanation of symbols]

[0100] 1. Areas of Interest 11. Attention (LED wall) 12 people 2 Cameras 21 Tracking System 3. Image Processing System 31 Processing section 32 Modeling Department 4. Expansion Area 41 Background Elements (Green Screen) 5 Visual Reference Targets p x Image (frame) position d x Distance from the center of the image c Image (frame) center

Claims

1. 1. A method, preferably a computer-implemented method, for creating a digital augmented camera image based on an input image taken by a camera (2), comprising: a) providing a vignette model of said camera (2); b) taking at least one input image with said camera (2); c) determining a dilated region (4) in the input image to be digitally dilated; d) providing virtual content image data for at least said extension area (4); e) generating a digitally augmented output image by augmenting said input image with said virtual content image data at least in said augmentation area (4), 10. The method of claim 9, wherein the determination of the extension area (4) and / or the provision of the virtual content image data is based on the vignette model of the camera (2).

2. 2. The method of claim 1, wherein the step of determining the extended region (4) in the input image comprises correcting the input image based on the vignetting model, and / or the method comprises applying a de-vignetting effect to the input image, preferably at least to areas of the input image outside the extended region (4), based on the vignetting model.

3. 3. The method of claim 2, wherein the step of determining the extended region (4) in the input image comprises a chromakey process, and preferably, before determining the extended region (4) in the input image, a step of correcting the input image based on the vignette model is performed.

4. The step of determining the extension area (4) comprises: providing a 3D model of a background element (41), such as a green screen, captured in the input image; taking a plurality of reference images of the background element (41) with the camera (2) from a plurality of camera positions and / or camera poses; correcting the reference image based on the vignetting model to obtain a devignetted reference image; - assigning color information of the devignetted reference image to the 3D model of the background element (41) to obtain a 3D reference model of the background element; and determining, based on a comparison of the input image with the 3D reference model of the background element (41), an area in the input image in which the background element (41) is photographed as an augmented area (4).

5. 5. A method according to claim 1, wherein the vignette model is used in the step of creating the virtual content image data by applying a vignette effect to the virtual content image data based on the vignette model.

6. The step of determining the extension area (4) comprises: providing a 3D model of the object of interest (11) captured in said input image; establishing the position and attitude of the camera (2); calculating an image position of the object of interest (11) in the input image based on the 3D model of the object of interest (11) and the established position and orientation of the camera (2); The method according to any one of claims 1 to 5, in particular the method according to claim 5, further comprising determining the extension region (4) based on the image position of the object of interest (11) in the input image.

7. 7. The method according to claim 1, wherein the vignette model comprises a function that maps a vignette value, which indicates a reduction in brightness and / or saturation due to a vignetting effect of the camera, to each pixel position in an image frame of the camera.

8. 8. The method according to claim 1, wherein the vignette model comprises a function that maps a vignette value, which indicates a reduction in brightness and / or saturation due to a vignetting effect of the camera (2), to a distance relative to the center of the image frame of the camera (2).

9. 9. The method of claim 8, wherein the center of the image frame of the camera (2) includes an offset value indicating an offset between a geometric center of the image frame of the camera (2) and a center of the vignette model at which the vignette value is minimum or maximum.

10. A digital image processing system (3) for creating digital augmented camera images based on an input image taken by a camera (2), preferably using a method according to any one of claims 1 to 9, comprising: receiving an input image from a camera (2); determining a dilated region (4) in the input image to be digitally dilated; providing virtual content image data for at least the extension area (4); a processing unit (31) configured to generate a digital augmented output image by augmenting the input image with the virtual content image data at least in the augmentation area (4); storing a vignette model of said camera (2); a modeling unit (32) configured to provide the vignette model to the processing unit (31), The digital image processing system (3) is configured to determine the extension area (4) and / or provide the virtual content image data based on the vignette model of the camera (2).

11. 11. The digital image processing system (3) of claim 10, wherein the processing unit (31) is configured to correct the input image based on the vignette model when determining the extended region (4) in the input image.

12. 12. The digital image processing system (3) of claim 11, wherein the processing unit (31) is configured to use chromakey processing when determining the extension region (4) in the input image, and preferably configured to correct the input image based on the vignette model before determining the extension region (4) in the input image.

13. The processing section (31) providing a 3D model of a background element, such as a green screen, captured in the input image; receiving a plurality of reference images of the background element taken by the camera (2) from a plurality of camera positions and / or camera poses; correcting the reference image based on the vignetting model to obtain a devignetted reference image; assigning color information of the devignetted reference image to the 3D model of the background element to obtain a 3D reference model of the background element; 13. The digital image processing system (3) of claim 10, configured to compare the corrected input image with the 3D reference model of the background element to determine an area in the input image in which the background element is captured as an augmented area (4).

14. 14. The digital image processing system (3) of claim 10, wherein the processing unit (31) is configured to create the virtual content image data using the vignette model by applying a vignette effect to the virtual content image data based on the vignette model.

15. The processing section (31) providing a 3D model of an object of interest captured in the input image; Establishing the position and orientation of the camera (2); calculating an image position of the object of interest in the input image based on the 3D model of the object of interest and the established position and orientation of the camera (2); A digital image processing system (3) according to any one of claims 10 to 14, in particular claim 14, configured to determine the extension region (4) based on the image position of the object of interest in the input image.