Updating the background layer during encoding
By classifying objects and using threshold-based updates for background layers with associated depth models, the method addresses the challenge of maintaining accurate and real-time depth perception in dynamic environments, optimizing computational efficiency and reducing false alarms.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- AXIS
- Filing Date
- 2025-10-29
- Publication Date
- 2026-05-19
AI Technical Summary
Existing depth perception systems struggle to maintain accurate and real-time depth information in dynamic environments where both static and dynamic elements are present, often requiring significant computational resources and time to update.
A method for updating background layers during scene encoding by classifying objects as foreground or background, using ordered background layers with associated depth models, and applying threshold-based comparisons to selectively update only significant changes, reducing unnecessary recalculations.
This approach optimizes computational efficiency and memory usage by minimizing redundant updates, ensuring accurate and timely depth perception in dynamic scenes, enhancing object classification and reducing false alarms.
Smart Images

Figure 2026082729000001_ABST
Abstract
Description
Technical Field
[0001] The embodiments presented in this specification relate to a method, an image processing device, a computer program, and a computer program product for updating a background layer during scene encoding.
Background Art
[0002] Depth perception is an essential aspect for understanding and interpreting the surrounding environment in various fields, especially in applications where three-dimensional (3D) spatial information is required. The ability to accurately determine the distance, shape, and size of objects within a scene enables more accurate analysis, improved object detection, and enhanced decision-making capabilities. Depth information is particularly useful in environments where it is important to distinguish objects based on distance or size, such as security systems, robotics, autonomous vehicles, and other systems that rely on visual data.
[0003] Conventionally, depth perception has been achieved by 3D cameras or other special sensors that provide a detailed understanding of the environment. The ability to perceive depth provides several advantages, including more accurate object detection and the ability to reduce false alarms by excluding objects that may appear larger or closer than they actually are. For example, depth perception can be useful in situations where objects may be misidentified based on a two-dimensional (2D) image, as depth information provides a more comprehensive view of the actual spatial relationships within a scene.
[0004] There are several methods by which depth information can be extracted. In some cases, a monocular camera system can estimate depth from a single viewpoint using an advanced computational model. In other cases, depth information is derived from disparity measurements obtained from overlapping images captured by multiple sensors, such as those used in multi-sensor panoramic systems. These systems calculate the difference in the positions of objects between images, enabling the determination of their relative distances. Another method involves sampling data from a laser point, such as those used in pan-tilt-zoom (PTZ) cameras equipped with lasers, to measure the distance of objects. Furthermore, self-learning techniques based on object tracking can provide depth information by analyzing how objects move and change position over time.
[0005] While these methods may be effective in relatively static environments, they face challenges when applied to more dynamic scenes. Each method typically requires a certain amount of time to process the data and compute accurate depth information. For example, monocular models often involve computationally intensive processes, and systems relying on PTZ cameras may require time for the camera to physically sweep or pan across the scene to collect sufficient data. This delay can hinder the ability to provide real-time or near real-time depth updates, especially in scenarios where large objects move quickly, causing rapid changes in their depth.
[0006] Therefore, in dynamic environments where objects can move at unpredictable or fluctuating speeds, keeping depth perception systems updated in real time becomes increasingly difficult. This challenge is particularly pronounced when the depth of large objects shifts dramatically, as the system may not be able to adjust quickly enough to provide accurate and up-to-date information.
[0007] As a result, there is a need for more efficient methods to maintain accurate depth perception, especially in situations where both static and dynamic elements are present in a scene. [Overview of the project]
[0008] The object of the embodiments described herein is to address the above-mentioned problems.
[0009] The specific objective is to provide computationally efficient techniques for maintaining accurate depth perception in scenes that contain both static and dynamic elements.
[0010] According to a first aspect, a method is presented for updating background layers during scene encoding, which is performed by an image processing device. The scene is encoded based on classifying objects depicted in the scene as either foreground or background. The background is an ordered background layer, where each background layer is divided into ordered background layers associated with their respective depth models. The method comprises detecting a change in an image portion of one background layer. The method comprises calculating the difference between the image portion and a corresponding image portion of a background layer ordered after the one background layer. The method comprises selecting the background layer ordered after the one background layer to represent the image portion when the difference is less than a threshold.
[0011] According to a second aspect, an image processing device is presented for updating background layers during scene encoding. The scene is encoded based on classifying objects depicted in the scene as either foreground or background. The background is an ordered background layer, where each background layer is divided into ordered background layers associated with their respective depth models. The image processing device comprises a processing circuit, which is configured to cause the image processing device to detect changes in an image portion of one background layer. The processing circuit is configured to cause the image processing device to calculate the difference between the image portion and the corresponding image portion of a background layer ordered after the one background layer. When the difference is less than a threshold, the processing circuit is configured to cause the image processing device to select the background layer ordered after the one background layer to represent the image portion.
[0012] According to a third aspect, a computer program is presented for updating background layers during scene encoding. The scene is encoded based on classifying objects depicted in the scene as either foreground or background. The background is an ordered background layer, where each background layer is divided into ordered background layers associated with their respective depth models. The computer program comprises computer code that, when executed on the processing circuitry of an image processing device, causes the image processing device to perform actions. One action comprises the image processing device detecting a change in an image portion of a background layer. Another action comprises the image processing device calculating the difference between the image portion and a corresponding image portion of a background layer ordered after the first background layer. Another action comprises the image processing device selecting the background layer ordered after the first background layer to represent the image portion when the difference is less than a threshold.
[0013] According to a fourth aspect, a computer program product is presented comprising a computer program according to a third aspect and a computer-readable storage medium in which the computer program is stored. The computer-readable storage medium may be a non-temporary computer-readable storage medium.
[0014] Advantageously, these embodiments provide a structured method for updating the background layer during encoding.
[0015] Advantageously, these embodiments enable efficient processing of background changes. More specifically, by detecting changes in individual image portions of the background layer and comparing them to corresponding portions of deeper background layers, the image processing device can efficiently identify minimal changes and avoid unnecessary recalculations or updates. This selective update mechanism reduces the computational load associated with continuous recalculation of the entire scene or background, especially in static or slowly changing environments. It ensures that the image processing device updates only the portions showing significant differences, optimizing both memory usage and processing power.
[0016] Advantageously, these embodiments allow for the use of layered depth models for better scene representation. More specifically, dividing the background into ordered layers, each associated with its own depth model, allows for a finer-grained and more accurate representation of the scene's depth. Each layer represents a different depth zone, enabling clear distinctions between objects at varying distances. This technique improves the accuracy of depth perception, especially in dynamic scenes, by maintaining more consistent and structured background data. Unlike flat depth maps or single-resolution systems, layered techniques can handle complex depth relationships more accurately and improve object classification between foreground and background.
[0017] Advantageously, these embodiments can leverage threshold-based decisions for computational efficiency. By using threshold-based comparisons to determine whether background layers need to be updated, image processing devices can skip redundant updates. When the difference between two background layers is smaller than a set threshold, a deeper background layer is selected, avoiding unnecessary recalculation of the depth model. This selective updating reduces the frequency of computationally intensive updates and ensures that image processing devices only process significant changes. This optimizes both time and energy consumption, making image processing devices more computationally efficient compared to solutions that continuously compute depth data regardless of scene changes.
[0018] Other purposes, features, and advantages of the attached embodiments will become apparent from the following detailed disclosure, from the attached dependent claims, and from the drawings.
[0019] In general, all terms used in the claims should be interpreted according to their ordinary meanings in the art unless otherwise expressly defined herein. All references to “a / an / the element, apparatus, component, means, module, step, etc.” should be broadly interpreted as referring to at least one instance of an element, apparatus, component, means, module, step, etc. unless otherwise expressly stated. The steps of the methods disclosed herein do not need to be performed in the exact order disclosed unless expressly stated otherwise.
[0020] Next, the concept of the present invention will be explained with reference to the attached drawings. [Brief explanation of the drawing]
[0021] [Figure 1] This is a schematic diagram showing an image processing device according to an embodiment. [Figure 2]It is a diagram schematically showing the foreground and background representation of a scene according to an embodiment. [Figure 3] It is a flowchart of a method according to an embodiment. [Figure 4(a)] It is a diagram showing a first example in which changes occur in an image portion of one background layer according to an embodiment. [Figure 4(b)] It is a diagram showing a first example in which changes occur in an image portion of one background layer according to an embodiment. [Figure 5(a)] It is a diagram showing a second example in which changes occur in an image portion of one background layer according to an embodiment. [Figure 5(b)] It is a diagram showing a second example in which changes occur in an image portion of one background layer according to an embodiment. [Figure 6] It is a schematic diagram showing the structural units of an image processing device according to an embodiment. [Figure 7] It is a diagram showing an example of a computer program product including a computer-readable storage medium according to an embodiment.
Embodiments for Carrying Out the Invention
[0022] Next, the concept of the present invention will be more fully described below by referring to the accompanying drawings showing specific embodiments of the concept of the present invention. However, the concept of the present invention can be embodied in many different forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided as examples so that this disclosure is thorough and complete and fully conveys the scope of the concept of the present invention to those skilled in the art. Throughout the description, like numbers refer to like elements. Steps or features shown by dashed lines should be regarded as optional.
[0023] As disclosed above, there is a need for a more efficient way to maintain accurate depth perception, especially in situations where both static and dynamic elements are present within a scene. Conventional background modeling techniques often rely on a single background depth model and can be inefficient when dealing with dynamic scenes with frequent changes. To address this problem, the use of multiple background image models has been proposed. Each background image model is associated with a corresponding depth model. The background image models are arranged in layers with different background merge times. Using such a cascaded background depth model can improve depth perception and accuracy in dynamic scene analysis.
[0024] Embodiments disclosed herein relate particularly to techniques for updating background layers during the encoding of a scene. To obtain such techniques, there is provided an image processing device, a method implemented by the image processing device, and a computer program product comprising code, in the form of, for example, a computer program, that when executed on the image processing device causes the image processing device to execute the method.
[0025] Figure 1 is a schematic diagram illustrating a scenario in which the image processing device 110 is used to capture a video sequence of scene 120. Different examples of scene 120 are provided below with reference to Figures 4 and 5. The image processing device 110 comprises a camera device 112. The camera device 112 is configured to capture a video sequence consisting of image frames. In some examples, the camera device 112 is a digital camera device and / or capable of pan, tilt and zoom (PTZ), and thus may be considered a (digital) PTZ camera device. Furthermore, the image processing device 110 is configured to encode image frames so that the video sequence can be decoded using any known video coding standard, such as, to name a few, High Efficiency Video Coding (HEVC), also known as H.265 and MPEG-H Part 2; Advanced Video Coding (AVC), also known as H.264 and MPEG-4 Part 20; Multipurpose Video Coding (VVC), also known as H.266, MPEG-I Part 3 and Future Video Coding (FVC); VP9, VP10, and AOMedia Video 2 (AV1). In this regard, the encoding is performed either in direct cooperation with the camera device 112 that captures the image frames, or in another entity such as a dedicated image encoder 114, and can then be stored at least temporarily in the database 116 for later retrieval, decoding, and viewing.
[0026] The image processing device 110 maintains several background image models, each having a different background merge time that represents the frequency with which the background is expected to change. Each background image model has a corresponding depth model that reflects the depth information associated with that particular background image. An example of a foreground-background representation 200 of a scene is shown in FIG. 2. As further disclosed below, the foreground-background representation can be encoded in the encoded video stream of the scene. The foreground-background representation 200 is composed of one foreground layer 210, or simply the foreground, and N background layers 220, or simply the background. Each background layer is associated with its own depth model indicated by dm1, dm2, dm3, ..., dmN. Thus, each image portion 225 has a depth value given by the depth model of the background layer representing the image portion 225. Further, each background layer is associated with its own merge time indicated by mt1, mt2, mt3, ..., mtN. The representation of an object in scene 120 can be merged into a given background layer when the object remains stationary in scene 120 for longer than the merge time of the given background layer. Thus, an object in scene 120 is first merged into the background layer with the shortest merge time (i.e., mt1). Thus, each background layer can be associated with its respective merge time. In this regard, the background layers are ordered according to their merge times, and the background layer with the shortest merge time is ordered closest to the foreground 210. That is, according to the notation in FIG. 2, mt1 < mt2 < mt3 < ... < mtN. In this regard, the merge time of a given background layer defines the time it takes for the immediately preceding ordered background layer to merge with this given background layer, as indicated by the dotted arrow in FIG. 2. Thus, a given background layer can be merged with the background layer ordered behind it after existing for longer than the merge time of the background layer ordered behind it.
[0027] Figure 3 is a flowchart illustrating an embodiment of a method for updating background layers 410a-410c, 510a-510c during the encoding of scene 120. Scene 120 is encoded based on classifying objects 420, 520 depicted in scene 120 as either foreground 210 or background 220, where background 220 is divided into ordered background layers 410a-410c, 510a-510c, where each background layer 410a-410c, 510a-510c is associated with its respective depth model indicated by dm. This method is carried out by an image processing device 110. This method is advantageously provided as a computer program.
[0028] S102: The image processing device 110 detects a change in the image portion 225 of one background layer.
[0029] S104: The image processing device 110 calculates the difference between the image portion 225 and the corresponding image portion 225' of the background layer ordered behind this one background layer.
[0030] S106a: When the difference is less than a threshold, the image processing device 110 selects a background layer ordered after the one background layer to represent the image portion 225.
[0031] In this way, when one background layer changes (as in step S102), it is checked (as in step S104) whether the change represents a new background model or a change that reverts to a background model with a longer merge time. If it is a change that reverts to a background model with a longer merge time, the current depth map is updated using the corresponding background depth model (as in step S106a). Thus, when a change occurs in the background, the image processing device 110 first checks whether the change matches an existing background image model with a longer background merge time. This comparison may be performed using techniques such as visual similarity or segmentation map comparison. One non-restrictive example of how visual similarity can be used in this context is provided in "Background modeling using mixture of Gaussians and Laplacian pyramid decomposition" by Minyong Wan et al., published in the minutes of the 2011 International Conference on Soft Computing and Pattern Recognition (SoCPaR), DOI:10.1109 / SoCPaR.2011.6089091. If the change matches an existing model, the current depth map is updated using the corresponding depth model.
[0032] Next, with continued reference to Figure 3, embodiments relating to further details of the updating of background layers 410a-410c and 510a-510c during the encoding of scene 120 performed by the image processing device 110 are disclosed.
[0033] If the detected change does not match the existing background image model, the image processing device 110 treats it as a new background change. A new (partial) background image model with a different background merge time may then be created. This triggers a process for creating a new background depth model if the change represents the new background model. Thus, in some embodiments, the image processing device 110 is configured to perform (optional) step S106b.
[0034] S106b: When the difference is not less than the threshold, the image processing device 110 creates a new background layer to represent the image portion 225.
[0035] As a result, when a new background change is detected, the image processing device 110 creates a new background depth model corresponding to the new image portion 225. This ensures that the depth map is always updated accurately in real time, even when a new background change occurs. This new background model can be treated as a temporary "branch" of the background and is retained in the image processing device 110 until further changes occur or until it stabilizes as part of the background.
[0036] If time has elapsed and the new partial background image model remains stable, the image processing device 110 merges the new background layer with the existing background layer which has a longer merge time and a similar depth model. Thus, in some embodiments, the image processing device 110 is configured to perform (optional) step S108.
[0037] S108: The image processing device 110 merges the new background layer with the background layers that were ordered after the new background layer when the merge time associated with the background layers that were ordered after the new background layer expires.
[0038] This process keeps the background model organized and reduces redundant data.
[0039] As disclosed above, the foreground 210 and background layers can be encoded into the encoded video stream of scene 120. Thus, in some embodiments, the image processing device 110 is configured to perform (optional) step S110.
[0040] Therefore, in some embodiments, the image processing device 110 is configured to perform (optional) step S110.
[0041] S110: The image processing device 110 encodes the foreground 210 and background layers into an encoded video stream of the scene 120.
[0042] Next, with reference to Figures 4 and 5, two exemplary examples are disclosed in which a change occurs within an image portion of a single background layer. Both figures show image frames 400a, 400b, 500a, and 500b depicting a scene with an office building and parked vehicles, some of which are stationary, while one moves between different image frames.
[0043] Figures 4(a) and 4(b) illustrate a scenario in which vehicle 420 leaves the scene. First, we assume that the image frame has a foreground background representation with several background image models, as shown in Figure 4(a) and image frame 400a. As mentioned above, each background layer is associated with its own merge time. For the sake of ease of explanation, let's assume there are three background layers 410a, 410b, and 410c. The office building (as well as the part of the scene located behind the parking lot and parked vehicles) is assumed to have been stationary for the longest time and therefore belongs to background layer 410a, which has the longest merge time. Vehicle 420 has been stationary for the shortest time and therefore belongs to background layer 410c, which has the shortest merge time. The remaining vehicles have been stationary longer than vehicle 420 and therefore belong to background layer 410b, which has an intermediate merge time. Next, we assume that vehicle 420 has left the scene, as shown in Figure 4(b) and image frame 400b. Therefore, the image processing device 110 detects a change in the image portion of the background layer 410c (which appears as the foreground layer 410c' in Figure 4(b)) as in step S102, and then proceeds to step S104. The foreground layer 410' is merged into the selected background layer 410a after the shortest background merge time (i.e., the time at which the foreground must be merged into either the selected background layer as in step S106a or the newly created background layer as in step S106b). In this example, this is the background merge time for the new background layer created in step S106b. In Figure 4(b), the background layer 410a appears as background layer 410a'. In this respect, there is no need to update the background layer 410a, because the background layer 410a always also contains information about the scene behind the vehicle 420, and this information is used during the comparison before step S106a.
[0044] Figures 5(a) and 5(b) illustrate a scenario in which vehicle 520 enters the scene and comes to a stop. First, we assume that the image frame has a foreground-background representation with several background image models, as shown in Figure 5(a) and image frame 500a. As described above, each background layer is associated with its own merge time. For the sake of clarity, let's assume there are two background layers 410a and 410b. The office building (as well as the portion of the scene located behind the parking lot and parked vehicles) is assumed to have been stationary for the longest time and therefore belongs to background layer 510a, which has the longest merge time. The vehicle has been stationary for the shortest time and therefore belongs to background layer 510b, which has a shorter merge time than background layer 510a. Next, we assume that vehicle 520 enters the scene and comes to a stop, as shown in Figure 5(b) and image frame 500b. Thus, the image processing device 110 detects a change in the image portion of background layer 510a, as in step S102, and then proceeds to step S104. Since the image portion of background layer 510a on which vehicle 520 is located does not correspond to the image portions of other background layers, a new background layer 510c may be created, as in step S106b. This new background layer 510c has a shorter merge time than background layer 510b.
[0045] The described method ensures that the depth map is updated with minimal latency, allowing the image processing device 110 to respond quickly to changes in the scene. As described below, this enables the image processing device 110 to adjust focus, optimize stitching distance, refine analysis rules, and improve geospatial awareness. Regarding focusing, the image processing device 110 can dynamically adjust focus based on the updated depth information to keep the target object in focus. Regarding optimizing stitching distance, in panoramic or multi-camera scenarios, accurate depth information improves stitching between camera feeds. Regarding refining analysis rules, the updated depth map improves object detection and classification, thereby reducing false alarms. Regarding improving geospatial awareness, the image processing device 110 can integrate a Global Positioning System (GPS) overlay with the updated depth map to improve camera positioning and scene understanding. In this regard, images captured by camera 112 may be geotagged. For example, location information may be provided by metadata tags in the Exif (Exchangeable Image File Format) data. In this way, the camera 112 can provide not only the pixel positions of detected objects but also the GPS coordinates of detected objects as output. By managing a cascaded background depth model, this method supports more efficient real-time image processing in dynamic scenes.
[0046] Figure 6 schematically shows the components of the image processing device 600 according to an embodiment in terms of several structural units. The processing circuit 610 is provided using one or more combinations of suitable central processing units (CPUs), multiprocessors, microcontrollers, digital signal processors (DSPs), etc., which are capable of executing software instructions stored in a computer program product 710 (as shown in Figure 7) in the form of a storage medium 630. The processing circuit 610 may further be provided as at least one application-specific integrated circuit (ASIC) or field-programmable gate array (FPGA).
[0047] In particular, the processing circuit 610 is configured to cause the image processing device 600 to perform a set of operations or steps as disclosed above. For example, the storage medium 630 can store a set of operations, and the processing circuit 610 may be configured to retrieve the set of operations from the storage medium 630 and cause the image processing device 600 to perform the set of operations. The set of operations may be provided as a set of executable instructions.
[0048] Accordingly, the processing circuit 610 is configured to perform the methods disclosed herein. The storage medium 630 may also comprise a persistent storage device which may be, for example, one or a combination thereof, of magnetic memory, optical memory, solid-state memory, or even remotely implemented memory. The image processing device 600 may further comprise a communication (comm.) interface 620 configured for communication with at least other entities, functions, nodes, and devices. Accordingly, the communication interface 620 may comprise one or more transmitters and receivers comprising analog and digital components. The processing circuit 610 controls the overall operation of the image processing device 600, for example, by transmitting data and control signals to the communication interface 620 and the storage medium 630, by receiving data and reports from the communication interface 620, and by retrieving data and instructions from the storage medium 630. Other components of the image processing device 600 and related functions are omitted in order not to obscure the concepts presented herein.
[0049] The image processing devices 110, 600 may be provided as standalone devices or as part of at least one further device. This means that a first portion of the instructions performed by the image processing devices 110, 600 may be executed on a first device, a second portion of the instructions performed by the image processing devices 110, 600 may be executed on a second device, and the embodiments disclosed herein are not limited to a specific number of devices on which the instructions performed by the image processing devices 110, 600 can be executed. Therefore, the methods according to the embodiments disclosed herein are suitable for implementation by image processing devices 110, 600 residing in a cloud computing environment. Thus, although a single processing circuit 610 is shown in Figure 6, the processing circuit 610 may be distributed across multiple devices or nodes. The same applies to the computer program 720 in Figure 7.
[0050] Figure 7 shows an example of a computer program product 710 comprising a computer-readable storage medium 730. A computer program 720 may be stored in this computer-readable storage medium 730, which can cause a processing circuit 610, and entities and devices operably coupled to the processing circuit 610, such as a communication interface 620 and the storage medium 630, to perform the methods according to the embodiments described herein. Thus, the computer program 720 and / or the computer program product 710 can provide means for carrying out the steps disclosed herein.
[0051] In the example in Figure 7, the computer program product 710 is shown as an optical disc such as a CD (Compact Disc), DVD (Digital Multipurpose Disc), or Blu-ray disc. The computer program product 710 can also be embodied as a non-volatile storage medium in an external memory device, such as a random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or electrically erasable programmable read-only memory (EEPROM), or more specifically, a USB (Universal Serial Bus) memory or flash memory such as CompactFlash memory. Thus, although the computer program 720 is schematically shown here as a track on the optical disc depicted, the computer program 720 can be stored in any way suitable for the computer program product 710.
[0052] The concept of the present invention has been described above, primarily with reference to several embodiments. However, as will be readily apparent to those skilled in the art, embodiments other than those disclosed above are equally possible within the scope of the concept of the present invention as defined by the appended claims.
Claims
1. A method for updating background layers during scene encoding, performed by an image processing device, wherein the scene is encoded based on classifying objects depicted in the scene as either foreground or background, the background is an ordered background layer, each background layer is divided into ordered background layers associated with its respective depth model (dm), and the method Detecting changes within a portion of an image in a single background layer, The difference between the aforementioned image portion and the corresponding image portion of the background layer ordered after the one background layer is calculated, When the difference is smaller than the threshold, the background layer ordered after the one background layer is selected to represent the image portion. When the difference is not less than the threshold, a new background layer is created to represent the image portion. A method that includes [a certain feature].
2. The method described above is Merging the new background layer with the background layer that was ordered after the new background layer upon the expiration of the merge time (mt) associated with the background layer that was ordered after the new background layer. The method according to claim 1, further comprising:
3. The method according to claim 1, wherein each background layer is associated with its respective merge time (mt).
4. The method according to claim 3, wherein the representation of an object in the scene is merged into the given background layer when the object remains stationary in the scene for longer than the merge time (mt) of the given background layer.
5. The method of a combination of claims 2 and 3, wherein a given background layer is merged with a background layer ordered after it has been present for a longer time than the merge time (mt) of the background layer ordered after it.
6. The method according to claim 1, wherein the background layers are ordered by merge time (mt), and the background layer having the shortest merge time (mt) is ordered closest to the foreground.
7. The method according to claim 1, wherein the image portion has a depth value given by the depth model (dm) of the background layer representing the image portion.
8. The method described above is Encoding the foreground and background layers into the encoded video stream of the scene. The method according to claim 1, further comprising:
9. An image processing device for updating background layers during scene encoding, wherein the scene is encoded based on classifying objects depicted in the scene as either foreground or background, the background is an ordered background layer, each background layer is divided into ordered background layers associated with a respective depth model (dm), and the image processing device comprises a processing circuit, the processing circuit provides to the image processing device, Detecting changes within a portion of an image in a single background layer, The difference between the aforementioned image portion and the corresponding image portion of the background layer ordered after the one background layer is calculated, When the difference is smaller than the threshold, the background layer ordered after the one background layer is selected to represent the image portion. When the difference is not less than the threshold, a new background layer is created to represent the image portion. An image processing device configured to perform the following actions.
10. The image processing device according to claim 10, further configured to carry out the method described in any one of claims 2 to 8.
11. A computer program for updating background layers during scene encoding, wherein the scene is encoded based on classifying objects depicted in the scene as either foreground or background, the background is an ordered background layer, each background layer is divided into ordered background layers associated with a respective depth model (dm), and when the computer program is executed on the processing circuit of an image processing device, the image processing device... Detecting changes within a portion of an image in a single background layer, The difference between the aforementioned image portion and the corresponding image portion of the background layer ordered after the one background layer is calculated, When the difference is smaller than the threshold, the background layer ordered after the one background layer is selected to represent the image portion. When the difference is not less than the threshold, a new background layer is created to represent the image portion. A computer program that includes computer code to perform a certain action.
12. A computer program product comprising a computer program according to claim 11 and a computer-readable storage medium in which the computer program is stored.