Video processing system for imperceptible blending of objects with a video patch

EP4802714A1Pending Publication Date: 2026-09-09UNIQFEED AG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
EP2024759126
Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-30
Filing Date
2024-08-21
Publication Date
2026-09-09

AI Technical Summary

Technical Problem

Existing video processing systems lack the capability for advanced image adjustment and manipulation that allows for seamless integration of country-specific content into global video feeds, particularly in real-time applications.

Method used

A video processing system that includes a video patch generation unit, which produces video patches based on input video images and assigned masks, allowing for real-time manipulation of video frames to create a seamless and non-perceptible overlap of objects or advertisements.

Benefits of technology

Enables improved image adjustment and manipulation, allowing for the creation of country-specific video feeds by seamlessly integrating or removing advertisements and content, while maintaining a real-time processing capability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024073422_08052025_PF_FP_ABST
    Figure EP2024073422_08052025_PF_FP_ABST
Patent Text Reader

Abstract

The invention describes a video processing system (10), in particular as part of a television transmission system, for processing video frames (VE, VEc) of a video image sequence, in particular a video stream, wherein the video processing system (10) is configured to provide at least one associated image mask (PM, VM) and associated camera data (KD) for each video frame (VE, VEc), wherein the video processing system (10) has: a video frame manipulation unit (12); a video patch generation unit (14); which has: - a texture patch generating unit (16) configured to calculate a texture patch (TP) based on a video frame (VE) and at least one image mask (PM, VM) associated with this video frame (VE); - a reference storage unit (18) configured to store at least two reference images (RB1, RB2) with respective reference texture patches (RT1, RT2), wherein a reference image (RB1, RB2) corresponds to a video frame (VE) on the basis of which the texture patch (TP) has been calculated, wherein the texture patch is stored as a reference texture patch (RT1, RT2); - an adjustment unit (20) configured to calculate the associated video patch (VP) for each video frame (VEc) entering the video processing system (10), wherein the at least two stored reference images (RB1, RB2) and reference texture patches (RT1, RT2) are taken into account in the calculation.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Video processing system for imperceptible blending of objects with a video patch

[0002] DESCRIPTION:

[0003] The invention relates to a video processing system, in particular as part of a television transmission system, for processing individual video images of a video image sequence, in particular of a video stream, wherein the video processing system is configured to provide at least one associated image mask and associated camera data for each individual video image, and wherein the video processing system has a video image manipulation unit configured to adapt each individual video image received on the input side and to output it as a manipulated individual video image on the output side, such that a video image sequence is formed from manipulated individual video images.

[0004] As examples of the prior art, reference is made to DE 10 2016 119 637 A1, DE 10 2016 119 639, and DE 10 2016 119 640 A1 of the applicant. These disclose techniques for the so-called enrichment of individual video images, particularly in real time. Starting from a so-called world feed, a country-specific feed can be generated, in which country-specific adaptations or enrichments are made. For example, an advertisement in a sports venue, which is displayed on the individual video images of the world feed, can be overlaid with a country-specific advertisement.

[0005] The object underlying the invention is to provide a video processing system that enables improved image adaptation or image manipulation that is imperceptible to the consumer of a manipulated video stream. This object is achieved by a video processing system having the features of the independent patent claim. Advantageous embodiments with expedient further developments are specified in the dependent patent claims.

[0006] A video processing system is therefore proposed, in particular as part of a television transmission system, for processing individual video images of a video image sequence, in particular of a video stream, wherein the video processing system is configured to provide at least one associated image mask and associated camera data for each individual video image, and wherein the video processing system comprises: a video image manipulation unit configured to adapt each individual video image arriving on the input side and to output it as a manipulated individual video image on the output side, such that a video image sequence is formed from manipulated individual video images;a video patch generation unit configured to provide an associated video patch for each video frame input to the video processing system, wherein the video frame manipulation unit is configured to generate the manipulated video frame based on the video patch; wherein the video patch generation unit comprises:;

[0007] - a texture patch generation unit configured to calculate a texture patch based on a video frame and at least one image mask associated with that video frame;

[0008] - a reference storage unit configured to store at least two reference images with respective reference texture patches, wherein a reference image corresponds to a video frame on the basis of which the texture patch has been calculated, wherein the texture patch is stored as a reference texture patch;

[0009] - an adaptation unit configured to calculate the associated video patch for each video frame input to the video processing system, wherein the at least two stored reference images and reference texture patches are taken into account during the calculation. In the video processing system, an image mask can be a patch mask, wherein the patch mask is a binary mask with the same pixel resolution as the video frame, wherein the patch mask specifies pixels to be manipulated in the video frame. The patch mask is assigned to a respective video frame, such that the pixels to be manipulated are specified for this video frame. For example, patch pixels to be manipulated can be marked with "1," while all other pixels are marked with "0."

[0010] In the video processing system, an image mask can be a foreground mask. The foreground mask is a binary mask with the same pixel resolution as the video frame. The foreground mask specifies pixels that represent recorded objects that cover areas identified by the patch mask. The foreground mask is assigned to each video frame, so that for this video frame, it is determined which (patch) pixel to be manipulated is covered by a foreground pixel. For example, foreground pixels can be marked with "1," while all other pixels are marked with "0."

[0011] In the video processing system, a fused mask can be formed from the patch mask and the foreground mask. The fused mask is a binary mask with the same pixel resolution as the video frame. The fused mask specifies pixels that form a patch pixel and / or a foreground pixel. For example, patch pixels and foreground pixels to be manipulated can be labeled "1," while all other pixels are labeled "0."

[0012] In the video processing system, the texture patch generation unit can comprise a precondition module configured to check whether one or more preconditions are met and to calculate the texture patch only if the precondition(s) are met. In other words, the precondition module calculates texture patches only under specific or desired circumstances, wherein these circumstances or preconditions are adjustable or adaptable.

[0013] The precondition(s) can be one or more of the following list: a) the texture patch generation unit has free computing capacity; b) an image edge of a video frame is located more than one pixel number away from a nearest patch pixel of the patch mask; c) a foreground mask has a areal overlap with a patch mask that is less than an overlap threshold; d) a foreground mask and a patch mask have no areal overlap; e) a foreground mask has a areal overlap with a patch mask for a time period that is greater than a time threshold; f) a current value for a camera movement that is part of the camera data of the video frame is less than a movement threshold.

[0014] Precondition a) checks whether the texture patch generation unit has already finished calculating a previous texture patch and is able to calculate a new result or texture patch.

[0015] Precondition b) checks whether the pixels to be manipulated, as defined by the patch mask, are located within the video frame.

[0016] Precondition c) ensures that a majority of the pixels to be manipulated, i.e. the patch pixels of the patch mask, are not covered by foreground pixels.

[0017] Precondition d) checks whether there is no overlap at all between the foreground mask and the pixel mask. Precondition e) ensures, in the case of an overlap, as in precondition c), that this overlap does not only occur briefly, but for a longer period of time, for example, several seconds. A time limit of 10 seconds or more can be set as an example.

[0018] In precondition f), the camera data for the video frame can be evaluated to ensure that a texture patch is only calculated if the camera does not move too quickly.

[0019] In the video processing system, the texture patch generation unit may be configured to store a video frame, the associated patch mask, and the camera data until the precondition(s) are met a next time.

[0020] The texture patch generation unit may be configured to calculate a texture video frame that is the same as the video frame except in the areas with patch pixels based on the patch mask, and to calculate pixel colors for the patch pixels such that the texture video frame is perceived by a viewer as a real camera image.

[0021] The texture patch generation unit may comprise a texture patch extraction module configured to generate the texture patch in geometrically normalized format based on the calculated texture video frame.

[0022] In this context, it should be noted that the texture patch generation unit can potentially also calculate a result (texture video frame) that manipulates a larger area than specified by the patch mask. In other words, the manipulation is not necessarily limited to the patch pixels, but can also occur adjacent to them or in other image areas. However, in such a case, the texture patch extraction module will only extract the area specified by the patch mask from this result and save the extracted texture patch in a geometrically normalized format.

[0023] In the video processing system, the adaptation unit may be configured to determine a first video patch based on a first reference image and a first reference texture patch, determine a second video patch based on a second reference image and a second reference texture patch, calculate a video patch applicable to the video frame to be output based on the first video patch and the second video patch, such that during a certain period of time, the first video patch and the second video patch are mixed in terms of colors.

[0024] In the video processing system, the applicable video patch can be calculated using the formula: applicable video patch = alpha * first video patch +

[0025] (1 -alpha) * second video patch where alpha increases continuously from 0 to 1 from video frame to video frame of the video image sequence during a predetermined period of time, the period of time being in particular a few seconds.

[0026] In this context, it should be noted that a single video image can also be referred to as a video frame. Furthermore, it should be noted that the single video image manipulation unit can operate in real time, such that the adjustment or manipulation of the input video image occurs so quickly that there is no delay in the transmission or broadcast of the manipulated video image. This ensures that a video stream containing no manipulated video images and a video stream containing manipulated video images can be transmitted or provided simultaneously.

[0027] Further advantages and details of the invention will become apparent from the following description of embodiments with reference to the figures.

[0028] Fig. 1 is a schematic diagram of a video processing system;

[0029] Fig. 2 is a schematic representation of individual video images and associated image masks;

[0030] Fig. 3 is a schematic representation of video frames and manipulated video frames.

[0031] Fig. 1 shows a simplified and schematic representation of a video processing system 10. The video processing system 10 may be part of a television transmission system 200, which is shown in a highly simplified form in Fig. 3.

[0032] The video processing system 10 is used to process individual video images VE of a video image sequence, in particular a video stream. The video processing system 10 is configured to provide at least one associated image mask PM, VM and associated camera data KD for each individual video image VE.

[0033] The video processing system 10 has a video frame manipulation unit 12 which is configured to adapt each incoming video frame VE on the input side and to output it as a manipulated video frame manVE on the output side, such that a video frame sequence is formed from manipulated video frames manVE.

[0034] The video processing system 10 further comprises a video patch generation unit 14 which is configured to provide an associated video patch VP for each video frame VE entering the video processing system 10, wherein the video frame manipulation unit 12 is configured to generate the manipulated video frame manVE based on the video patch VP.

[0035] The video patch generation unit 14 has a texture patch generation unit 16 which is configured to calculate a texture patch TP based on a video frame VE and at least one image mask PM, VM associated with this video frame VE.

[0036] Furthermore, the video patch generation unit 14 has a reference storage unit 18 which is configured to store at least two reference images RB1, RB2 with respective reference texture patches RT1, RT2, wherein a reference image RB1, RB2 corresponds to a single video image VE on the basis of which the texture patch TP has been calculated, wherein the texture patch TP is stored as a reference texture patch RT1, RT2.

[0037] Furthermore, the video patch generation unit 14 has an adaptation unit 20 which is configured to calculate the associated video patch VP for each video frame VE input into the video processing system 10, wherein the at least two stored reference images RB1, RB2 and reference texture patches RT1, RT2 are taken into account in the calculation.

[0038] Before referring back to Fig. 1 later to describe the functioning of the video processing system 10 or the video frame manipulation unit 12 in more detail, some terms relating to the individual video images VE to be processed, which can also be referred to as frames, are explained below with reference to Figs. 2 and 3.

[0039] In Fig. 2, the first row shows, as an example, three individual video images VE with the respective individual image numbers or frame numbers N, N+x, N+y. The individual video images VEN, VEN+x, VEN+y are part of a video image sequence VBA, which can also be referred to as a video stream. The first row shows the sequence of three individual video images recorded by a camera during a video image sequence. These three individual video images VE can follow one another directly, i.e., with x=1 and y=2, but they cannot follow one another directly within the video stream, for example, with x=20 and y=40.

[0040] In each individual video frame, the camera captures: a background UG, a boundary BG, for example, a wall or a barrier or the like, and an object OB represented as an ellipse, which is located in front of or on the boundary BG. The object OB can, for example, be a design represented on the boundary BG, such as text, a logo, or the like. Furthermore, the camera has captured an object VB located in the foreground, which is represented here as an exemplary contour arrow.

[0041] As can be seen from the sequence of video frames VE (N, N+x, N+y), object OB remains in place and forms a static part of the scene. In this example, the camera pans to the left, and at the same time, object VB, located in the foreground, moves in front of object OB, partially covering it.

[0042] In the three video frames VE (N, N+x, N+y) shown, a change in lighting conditions is represented by the narrowing hatching of the background UG, the object OB, and the foreground object VB, as well as the increasing dot density on the boundary BG. In other words, it gets slightly darker from the video frame VEN to the video frame VEN+y.

[0043] The second row shows the patch mask PM (N, N+x, N+y) assigned to the respective video frame VE (N, N+x, N+y). The PM assigned to the video frame VE is a binary mask with the same pixel resolution as the video frame VE, with the patch mask PM specifying pixels to be manipulated in the video frame VE. The pixels to be manipulated are referred to as patch pixels and are shown in black in the second row of Fig. 2. In the example shown, the object OB, represented as an ellipse, is to be manipulated.

[0044] The third row shows an optional foreground mask VM assigned to the respective video frame VE (N, N+x, N+y). The foreground mask VM assigned to the video frame VE is a binary mask with the same pixel resolution as the video frame VE. The foreground mask VM specifies pixels representing recorded objects VB that can cover areas identified by the patch mask PM, in this case, the object OB. Such pixels of the foreground mask VG are referred to as foreground pixels and are shown in black in the third row of Fig. 2.

[0045] The fourth row shows an optionally assigned or computable fused mask FM or fuse mask for the respective video frame VE (N, N+x, N+y). The fuse mask FM is a binary mask with the same pixel resolution as the video frame VE, with the fuse mask specifying pixels that form a patch pixel and / or a foreground pixel. Such pixels can also be referred to as fuse pixels, which are shown in black in the fourth row of Figure 2.

[0046] Fig. 3 is intended to illustrate how the respective input video frame VE is to be manipulated. The first row shows the video frames VE (N, N+x, N+y) already known from Fig. 2. By using the video processing system 10 or the video frame manipulation unit 12, the object OB represented as an ellipse is to be adapted to the appearance or texture of the boundary BG. In other words, the object OB should no longer be visible in the manipulated video frame without a viewer of the manipulated video frame or the video image sequence noticing the fading out or cross-fading of the object OB. This is shown in a simplified manner in the second and third rows, with the contour of the object OB being shown in the second row purely for illustrative purposes. The third row shows the end result of a manipulated video frame manVE.

[0047] After the adaptation or manipulation of individual video images VE has been graphically illustrated and explained with reference to Figs. 2 and 3, the operation of the video processing system 10 will now be further explained with reference to Fig. 1.

[0048] The texture patch generation unit 16 already mentioned above has a precondition module 22 which is configured to check the fulfillment of one or more precondition(s) and to calculate the texture patch TP only if the precondition(s) are fulfilled.

[0049] The precondition can be one or more of the following: a) the texture patch generation unit 16 has free computing capacity; b) an image edge of a video frame VE is located more than one pixel number away from a nearest patch pixel of the patch mask PM; c) a foreground mask VM has a surface overlap with a patch mask PM that is less than an overlap limit; d) a foreground mask VM and a patch mask PM do not have a surface overlap; e) a foreground mask VM has a surface overlap with a patch mask PM for a time period that is greater than a time limit; f) a current value for a camera movement that is part of the camera data KD of the video frame VE is less than a movement limit.

[0050] If the required precondition(s) is / are met, a video frame VE, camera data KD, patch mask PM and optionally a fuse mask FM are stored until they are overwritten again by the precondition module 22. The video frame VE and the fuse mask FM are inputs to a texture module 24. If no fuse mask FM is available, the patch mask PM can be used as input to the texture module 24. The texture module 24 generates or calculates a texture video frame TVE that is the same as the video frame VE except in the areas with patch pixels based on the patch mask PM. Pixel colors for the patch pixels are calculated in such a way that the texture video frame TVE is perceived by a viewer like a real camera image.

[0051] The texture module 24 is particularly designed such that it takes into account, for example, information about the boundary BG explained above with reference to Figs. 2 and 3, in particular its color configuration, in order to determine appropriate adjustments for the patch pixels. Changing lighting conditions can also be taken into account via the individual video frames recorded by a camera.

[0052] The texture video frame TVE is provided to a patch extraction module 26. The patch extraction module 26 uses the camera data KD and the patch mask PM as input variables and stores the color information of the texture video frame TVE in areas specified by the patch mask PM in geometrically normalized form, which is referred to as a texture patch TP.

[0053] Texture patches (TP) store only the color information at pixel locations indicated by the patch mask (PM), whereby their geometric representation is independent of the camera data (KD) of the specific video frame (VE). This storage format of the texture patch (TP) enables the geometrically correct application of a texture patch to another video frame (VE) that has different associated camera data (KD).

[0054] In this context, it should be noted that the texture module 24 can potentially also calculate a result (texture video frame TVE) that manipulates a larger area than specified by the patch mask PM. In other words, the manipulation is not necessarily limited to the patch pixels, but can also occur adjacent to them or in other image areas. In particular, it can be helpful for optimal perception if image content present in the vicinity of the patch pixels is also manipulated. However, in such a case, the texture patch extraction module 26 will only extract the area specified by the patch mask from this result and save the extracted texture patch in a geometrically normalized format.

[0055] Since the calculation of a texture patch TP cannot or must not necessarily be done in real time, a texture patch TP calculated for a video frame VE cannot be used for this video frame to generate the manipulated video frame manVE.

[0056] Therefore, a texture patch TP is permanently stored until an updated texture patch TP is calculated. This is done by the above-mentioned reference storage unit 18, which is configured to store at least two reference images RB1, RB2 with respective reference texture patches RT1, RT2. A reference image RB1, RB2 corresponds to a single video image VE on the basis of which the texture patch TP was calculated. The texture patch TP is stored as a reference texture patch RT1, RT2.

[0057] If two reference data sets, each containing a reference texture patch RT1 or RT2, and a reference image RB1 or RB2, are stored in the reference storage unit 18, an update can be carried out, for example, as follows.

[0058] As soon as, after fulfilling corresponding preconditions, a new video frame VE is used in the texture patch generation unit 16 to calculate a new texture patch TP, the reference image RB2 is overwritten with the data from the reference image RB1. The reference texture patch RT2 is overwritten with the data from the reference texture patch RT1. The new video frame VE is used as the reference image RB1, and the newly calculated texture patch TP is used as the reference texture patch RT1.

[0059] In other words, the reference texture patches RT1 , RT2 and the reference images RB1 , RB2 are updated or replaced as needed, whereby the older reference data set (RB2, RT2) is replaced by the younger reference data set (RB1 , RT1 ).

[0060] The adaptation unit 20 shown in Fig. 1 is configured, for example using a patch adaptation module 28, to determine a first video patch VP1 based on the first reference image RB1 and the first reference texture patch RT1, to determine a second video patch VP2 based on the second reference image RB1 and the second reference texture patch RT2, and to calculate the video patch VP applicable to the video frame VE to be output based on the first video patch VP1 and / or the second video patch VP2, such that during a specific period of time the first video patch VP1 and the second video patch are mixed in terms of colors.

[0061] The video patch VP applicable to a current video frame VEc can be calculated according to the formula: applicable video patch VP = alpha * first video patch VP1 + (1 - alpha) * second video patch VP2 where alpha increases continuously from 0 to 1 from video frame VE to video frame VE of the video frame sequence during a predetermined period of time, the period of time being in particular a few seconds.

[0062] In other words, the adaptation unit 20 is configured to adapt the color of the reference texture patches RT1, RT2 to generate video patches VP1, VP2 that enable imperceptible video image manipulation for a current video frame VEc. It is possible to use only the first video patch VP1 directly to perform manipulation on a current video frame VE and to calculate the manipulated video frame manVE.

[0063] For example, a current video frame VEc can be merged with the video patch VP in a rendering unit 30 so that the manipulated video frame manVE can be output.

[0064] The video processing system 10 presented here can be used, for example, in a sports broadcast or the like. The application will be explained below using a concrete example, with reference to Figures 1 to 3 as needed to connect concrete examples with the generalized terminology used here.

[0065] At a tennis event, a rights holder / organizer of the event may wish to make certain advertising permanently installed in the stadium (e.g., certain perimeter advertising) invisible in the television signal transmitted to certain countries. Referring to Figure 2, the advertising is the ellipse OB, and the stadium includes the underground area UG and the boundary BG.

[0066] Such a measure is necessary, for example, because there are country-specific legal regulations that prohibit certain types of advertising (e.g., alcohol or gambling). If such advertising is installed in the stadium and the broadcaster still wants to transmit the signal in a country with a corresponding advertising ban, the advertising must be visually removed from the television signal before transmitting the television signal.

[0067] In addition, in the example of a tennis event considered here, the color scheme of the advertising hoardings is often specified by the organizer. With reference to Fig. 2, this particularly applies to the color scheme of the boundary BG, on which the advertising (ellipse OB) is placed. The advertisers often cannot freely choose the colors of their advertising (ellipse OB).

[0068] Therefore, if an existing advertisement (Ellipse OB) is to be removed, the color scheme that is physically present in the stadium must be taken into account when calculating an image texture (Video Patch VP) that makes the advertisement (Ellipse OB) appear to have disappeared.

[0069] In addition, color variations caused by changes in lighting conditions (sunlight, floodlights, clouds, shadows) or the properties of the advertising material (fabric, plastic, etc.) must be taken into account when calculating the texture. Finally, a manipulated television signal (manVE) is generated in which the advertisement (ellipse OB) is masked by a calculated texture (video patch VP) in such a way that the advertisement (ellipse OB) appears to be absent from the television signal to the TV viewer. Reference is again made to Fig. 3 and its description.

[0070] Using the video processing system 10 presented here, an advertisement (ellipse OB) can be covered or blended using a calculated texture, which is ultimately provided as a video patch VP for a current video frame VE. The video patch VP is continuously adjusted or adjusted under certain conditions, taking into account the surroundings of the advertisement (ellipse OB) to be blended. Of course, it is also conceivable not only to blend an advertisement (ellipse OB) with a texture, but also to blend in another advertisement instead of the original advertisement, possibly in combination with the calculated texture.

Claims

PATENT CLAIMS:

1. Video processing system (10), in particular as part of a television transmission system, for processing individual video images (VE, VEc) of a video image sequence, in particular a video stream, wherein the video processing system (10) is configured to provide at least one associated image mask (PM, VM) and associated camera data (KD) for each individual video image (VE, VEc), and wherein the video processing system (10) comprises: a video image manipulation unit (12) configured to adapt each individual video image (VE, VEc) received on the input side and to output it as a manipulated individual video image (manVE) on the output side, such that a video image sequence is formed from manipulated individual video images (manVE);a video patch generation unit (14) configured to provide an associated video patch (VP) for each video frame (VE, VEc) entering the video processing system (10), wherein the video frame manipulation unit (12) is configured to generate the manipulated video frame (manVE) based on the video patch (VP); wherein the video patch generation unit (14) comprises:; - a texture patch generation unit (16) which is configured to calculate a texture patch (TP) based on a video frame (VE) and at least one image mask (PM, VM) associated with this video frame (VE); - a reference storage unit (18) which is configured to store at least two reference images (RB1, RB2) with respective reference texture patches (RT1, RT2), wherein a reference image (RB1, RB2) corresponds to a single video image (VE) on the basis of which the texture patch (TP) has been calculated, wherein the texture patch is stored as a reference texture patch (RT1, RT2); - an adaptation unit (20) which is designed to calculate the associated video patch (VP) for each video frame (VEc) received in the video processing system (10), wherein Calculation takes into account at least two stored reference images (RB1, RB2) and reference texture patches (RT1, RT2).

2. Video processing system (10) according to claim 1, wherein an image mask is a patch mask (PM), wherein the patch mask (PM) is a binary mask with the same pixel resolution as the video frame (VE), wherein the patch mask specifies pixels to be manipulated in the video frame (VE).

3. Video processing system (10) according to claim 2, wherein an image mask is a foreground mask (VM), wherein the foreground mask (VM) is a binary mask with the same pixel resolution as the video frame (VE), wherein the foreground mask (VM) indicates pixels representing captured objects (VB) covering areas (OB) identified by the patch mask (PM).

4. Video processing system (10) according to one of the preceding claims, wherein the texture patch generation unit (16) comprises a precondition module (22) configured to check the fulfillment of one or more precondition(s) and to perform the calculation of the texture patch (TP) only if the precondition(s) are fulfilled.

5. The video processing system (10) of claim 4, wherein the preconditions are one or more of the following: the texture patch generation unit (16) has free computing capacity; an image edge of a video frame (VE) is located more than a pixel number away from a nearest patch pixel of the patch mask (PM); a foreground mask (VM) has a surface overlap with a patch mask (PM) that is less than an overlap threshold; a foreground mask (VM) and a patch mask (PM) do not overlap; a foreground mask (VM) overlaps with a patch mask (PM) for a period of time that is greater than a time threshold; a current value for a camera movement that is part of the camera data (KD) of the video frame (VE) is less than a movement threshold.

6. Video processing system (10) according to claim 4 or 5, wherein the texture patch generation unit (16) is configured to store a video frame (VE), the associated patch mask (PM) and the camera data (KD) until the precondition(s) is / are fulfilled a next time 7. Video processing system (10) according to claim 6, wherein the texture patch generation unit (16) is configured to calculate a texture video frame that is the same as the video frame (VE) except in the areas with patch pixels based on the patch mask (PM), and To calculate pixel colors for the patch pixels in such a way that the texture video frame is perceived by a viewer as a real camera image.

8. The video processing system (10) of claim 7, wherein the texture patch generation unit (16) comprises a texture patch extraction module (26) configured to generate the texture patch (TP) in geometrically normalized format based on the calculated texture video frame.

9. Video processing system (10) according to one of the preceding claims, wherein the adaptation unit (20) is configured to determine a first video patch (VP1) based on a first reference image (RB1) and a first reference texture patch (RT1), to determine a second video patch (VP2) based on a second reference image (RB2) and a second reference texture patch (RT2), to calculate a video patch (VP) applicable to the video frame (VE, VEc) to be output based on the first video patch (VP1) and the second video patch (VP2), such that during a certain period of time the first video patch (VP1) and the second video patch (VP2) are mixed in terms of colors.

10. Video processing system (10) according to claim 9, wherein the applicable video patch (VP) is calculated according to the formula: applicable video patch = alpha * first video patch + (1 - alpha) * second video patch where alpha increases continuously from 0 to 1 from video frame to video frame of the video image sequence during a predetermined period of time, wherein the period of time is in particular a few seconds.