Encoding and decoding of video comprising multiple switchable overlays

By adding metadata of overlay image frames to the video stream, the challenge of switchable overlay in videos to bandwidth and synchronization is solved, efficient encoding and decoding is achieved, and streaming media quality and synchronization consistency is improved.

CN120111230APending Publication Date: 2025-06-06AXIS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411737681.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-05
Filing Date
2024-11-29
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Switchable overlays in videos have challenges in bandwidth and streaming quality as well as content synchronization, especially when additional data is required to be transmitted and ensure seamless synchronization.

Method used

By receiving multiple image frames and determining the superimposed image frame, metadata, including the superimposed position, size and identifier, is added to the header of the image frame, and then encode and decode the video stream to achieve synchronous and efficient processing of the superimposed.

Benefits of technology

This approach reduces bandwidth requirements, improves streaming quality, ensures perfect synchronization of overlays with video content, and simplifies decoding and error handling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120111230A_ABST
    Figure CN120111230A_ABST
Patent Text Reader

Abstract

The invention relates to encoding and decoding of video comprising multiple switchable overlays. In particular, methods of encoding and decoding for switchable overlay in video are disclosed. The encoding method comprises receiving (S702) a plurality of image frames; one or more overlay image frames are determined (S704), each overlay image frame comprising a plurality of switchable overlays, each switchable overlay being associated with an identifier. The image frame is associated with a superimposed image frame of the one or more superimposed image frames (S706). Metadata is added to a header of the image frame (S708), where, for each of the plurality of switchable overlays, metadata is added to a header of the image frame (S708). The metadata includes location data identifying a location of the switchable overlay in the overlay image frame, size data identifying a size of the switchable overlay in the overlay image frame, and identification data corresponding to an identifier of the switchable overlay.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to video overlays, and in particular to encoding and decoding methods for switchable overlays in video. Background Art

[0002] An overlay on a video is a feature used to add a layer of content such as text, graphics, or another video on top of the main material. There are several types of overlays, each used for a specific purpose. For example, text overlays include titles, subtitles, and any text information superimposed on the video, usually for subtitles or narratives. Graphic overlays range from simple shapes to complex illustrations, enhancing visual appeal or highlighting key information. In some cases, the examples of overlays mentioned above can be collectively referred to as graphic overlays.

[0003] Overlays are diverse not only in form, but also in application across a variety of fields. For example, in broadcasting, they can be used to display logos, news tickers, or sports scores, while in marketing and advertising, overlays can make promotional videos more engaging with branding elements or calls to action. In surveillance systems, overlays can be a key feature to enhance effectiveness and functionality. For example, overlays can display information such as timestamps and camera locations directly on the footage. This added layer of data can be used to contextualize the video, making it easier to pinpoint when and where specific events occurred. Other examples of overlays in surveillance applications include privacy masks or bounding boxes.

[0004] Switchable overlays in video content represent a cutting-edge way of viewer interaction, allowing users to interact with additional layers of information or graphics based on their preferences. However, this innovative feature introduces some technical challenges, particularly in terms of bandwidth and streaming quality, as well as content synchronization.

[0005] One of the main concerns with switchable overlays is their impact on bandwidth and streaming quality. Typically, these overlays require additional data to be transmitted along with the main video stream, especially when the video may be covered by overlays in different configurations and combinations. This additional data burden can strain bandwidth, potentially affecting streaming quality. Moreover, ensuring seamless content synchronization can pose another major challenge. Overlays, especially those that are interactive or provide supplementary information, need to be perfectly synchronized with the video content. Synchronization can be critical to maintain continuity and relevance. Achieving this level of synchronization can be technically demanding, especially when considering the variety of devices used by viewers.

[0006] Therefore, improvements are needed in this regard. Summary of the invention

[0007] In view of the above, it would be advantageous to solve or at least reduce one or several of the above mentioned disadvantages as set out in the accompanying independent patent claims.

[0008] According to a first aspect of the present invention, a method for encoding one or more video streams is provided, the method comprising the following steps: receiving a plurality of image frames; and determining one or more overlay image frames, each overlay image frame comprising a plurality of switchable overlays, each switchable overlay being associated with an identifier, wherein the overlay image frame has the same size as an image frame in the plurality of image frames.

[0009] The method further includes, for each image frame of at least some of the multiple image frames: associating the image frame with an overlay image frame of one or more overlay image frames; and adding metadata to a header of the image frame, wherein, for each switchable overlay of the multiple switchable overlays, the metadata includes: position data identifying a position of the switchable overlay in the overlay image frame, size data identifying a size of the switchable overlay in the overlay image frame, and identification data corresponding to an identifier of the switchable overlay.

[0010] The method further includes encoding one or more video streams, wherein encoding includes: encoding a plurality of image frames into a first plurality of encoded image frames using a first group of picture (GOP) structure; and encoding one or more superimposed image frames.

[0011] Overlay image frames may be determined from a variety of information sources. For example, a monitoring system may benefit from the integration of analytical metadata that can be used to determine the most relevant overlays for a given situation. The metadata may include details such as the classification, location, and number of objects in multiple image frames. By analyzing the metadata, the system may intelligently decide which types of overlays are best suited to be included in the overlay image frames. Overlay image frames may be further determined by incorporating status information from a camera that captures multiple image frames or from other sensors, such as door sensors that indicate an open or closed state. This approach may allow overlay image frames to include specific icons or text based on the real-time status of various sensors, thereby providing a more dynamic and responsive way to determine one or more overlay image frames.

[0012] The overlay in the overlay image frame can be visualized in different ways. For example, the information (e.g., video analysis metadata, sensor information, etc.) can be rendered as text as a text overlay, or can be rendered as a line graphic or other type of graphic.

[0013] In some cases, the overlay may be determined by converting from an image file such as JPEG, GIF, PNG, etc., further enhancing the usefulness of the encoding method.

[0014] At least some of the one or more image frames are associated with overlay image frames. For example, not all image frames may have an associated overlay that is visualized at the decoder side, e.g., depending on the lack of analytical metadata for the image content of these image frames. In other examples, only image frames that are later encoded into delta frames may be associated with overlay image frames, as described below.

[0015] For each image frame associated with an overlay image frame, metadata representing a plurality of switchable overlays in the overlay image frame is added to the header of the image frame. Advantageously, this may allow simplified decoding, synchronized rendering, resource optimization, improved error handling, and the like.

[0016] For example, by including overlay information in a frame header, a decoder can more efficiently process and render overlays. This data organization can ensure that overlays are displayed accurately and in a timely manner, thereby enhancing the overall viewing experience. By including identification data in the header, identification of the position and size of an overlay that is switched on or off in an associated overlay image frame can be simplified at the decoder side, thereby providing a low-complexity way of identifying associated overlay image data in an overlay image frame.

[0017] Furthermore, embedding the overlay details in the frame header ensures that the overlay is perfectly synchronized with the corresponding frame. This can be critical to maintain continuity, especially in videos where timing and accuracy are critical, such as surveillance footage or live broadcasts.

[0018] In addition, integrating the overlay information into the frame header can reduce processing overhead compared to including the overlay information in a separate data container in the encoded video stream. Including multiple switchable overlays in one overlay image frame can reduce the complexity of including the overlay in one or more video streams compared to including the multiple switchable overlays in separate overlay image frames.

[0019] Furthermore, error handling may be improved because built-in error handling at the decoder side may be used if a difference or corruption in the superimposed data in the header is detected when receiving a corrupted coded image frame.

[0020] Multiple image frames and one or more superimposed frames may be encoded into one or more video streams (bitstreams) according to a video coding format and / or standard, such as, for example, H.261, H.262, H.263, H.264 / AVC, H.265 / HEVC, VP8, VP9, ​​or AV1.

[0021] In some examples, encoding includes using a block-based codec that supports skip blocks, wherein encoded image frames and encoded overlay frames are included in a first video stream, wherein the method further includes setting each of the plurality of image frames as a non-display frame; wherein, for each delta-encoded image frame in the first plurality of encoded image frames, encoding includes: determining an overlay image frame from one or more overlay image frames; setting the overlay image frame as a display frame; encoding the overlay image frame delta frame into an encoded overlay frame referenced to the delta-encoded image frame, wherein each pixel block of the encoded overlay frame that does not correspond to any of the plurality of switchable overlays is set as a skip block; and including the encoded overlay frame as auxiliary data associated with the delta-encoded image frame in the first video stream.

[0022] Determining an overlaid frame of the one or more overlaid image frames includes using an overlaid image frame associated with an image frame encoded into the delta-encoded image frame.

[0023] In this method, the encoded superimposed frame (i.e., the macroblock of pixels encoded from the superimposed image frame) may, for example, be encoded as a further macroblock included in the same "image container" as the encoded image frame. All image data of the image container is therefore encoded into incrementally encoded frames (B frames or P frames). The header of the image container indicates to the decoder which part of the encoded image content in the image container is associated with the image frame. The header may therefore include information for distinguishing which parts of the encoded content are associated with the image frame and which parts are associated with the superimposed image frame. This information may include identifiers or flags indicating the beginning and end of the superimposed image frame and / or image frame within the container. The header may further include reference frame data identifying details about the reference frame used to encode the difference for both the encoded image frame and the encoded superimposed image frame.

[0024] The use of image containers to transmit both image data and overlays, particularly when the overlays are encoded as additional macroblocks, can provide a simplified and efficient method of video content management. An image container, in essence, is a digital file format that encapsulates various data types, main video and overlay data, into a unified structure, depending on the file format used (e.g., MP4, MOV, MKV or any other appropriate format). This not only simplifies processing and management, but also reduces the complexity of transmitting multiple data types. By including the encoded overlay image frame as an additional macroblock in the same container as the (main) image frame, the overlay image frame is directly integrated into the video stream, but as a unique component that can be independently identified (e.g., through information in the header) and processed.

[0025] One advantage of this approach is that it facilitates easier synchronization between the video content and the overlay. Because both types of data are stored in the same container, they are inherently synchronized. This reduces the risk of timing mismatches during playback, ensuring that the overlay appears at the correct moment relative to the video frame. It also simplifies the decoding process, because the decoder can process the video and overlay data simultaneously in a coordinated manner.

[0026] Furthermore, the use of image containers conforms to standard transport protocols, enhancing compatibility with a variety of distribution networks and playback devices. This standardization facilitates broad accessibility and ease of integration into existing video distribution infrastructure. Thus, video content along with its overlays can be transmitted over public networks and, as further described below, can be viewed on conventional devices.

[0027] By encoding the overlay image in such a manner, for example, by encoding an overlay image frame delta frame (having the same size as the associated image frame) into an encoded overlay frame that references the incrementally encoded image frame, and wherein pixel blocks of the encoded overlay frame that do not correspond to any of a plurality of switchable overlays are set as skip blocks, advantageously, a conventional decoding client can decode a video stream with all visible overlays by decoding and displaying only the incrementally encoded overlay image frame, i.e., without using the information provided in the header and without including functionality for individually switching the overlay on or off.

[0028] In video coding, the use of skip blocks is a technique used to increase coding efficiency, particularly in sequences where parts of video frames remain unchanged over several frames. When a block of pixels does not change significantly from one frame to the next, it is marked as a "skip block". Instead of re-encoding this unchanged block for each subsequent frame, the encoder simply references the block in the previous frame to indicate that it should be "skipped" or copied as is. This technique can be used in the context of encoding an overlay image frame that references an encoded image frame, because the pixel blocks between the overlays in the overlay image frame do not include any content in the overlay image frame, and can therefore be skipped, meaning that the corresponding pixel blocks in the reference image frame will be displayed instead.

[0029] In some embodiments, a first plurality of encoded image frames are included in a first video stream, and the encoded one or more overlay frames are included in a second video stream, wherein the step of associating the image frame with an overlay image frame in the one or more overlay image frames includes: including first synchronization data in the first video stream to associate the image frame with an overlay image frame in the one or more overlay image frames.

[0030] Such synchronization data may, for example, include an indication of an index of an overlaid image frame in one or more overlaid frames associated with a particular image frame.In some cases, the synchronization data includes a table mapping an index of an image frame to an index of an associated overlaid image frame.

[0031] By encoding the image frames and superimposed image frames in different video streams, increased flexibility can be achieved on the encoder side and the decoder side. For example, the step of encoding one or more superimposed image frames may include using a second GOP structure different from the first GOP structure. Moreover, the superimposed image frames encoded in the second video stream can be associated with other encoded video streams, and therefore the bit rate required for supplementing multiple encoded video streams with superimposed images can be reduced.

[0032] Accordingly, in some embodiments, the method may further include the steps of: encoding the plurality of image frames into a second plurality of encoded image frames using a third GOP structure; including the second plurality of encoded image frames in a third video stream; and including synchronization data in the third video stream. Advantageously, the present embodiment scales well with further encoding of the plurality of image frames. Thus, in some embodiments, the plurality of image frames may be encoded into even further video streams, i.e. a fourth video stream, a fifth video stream, etc.

[0033] In some embodiments, encoding the plurality of image frames into a first plurality of encoded image frames differs from encoding the plurality of image frames into a second plurality of encoded image frames in at least one of the following aspects: encoding quality, frame rate, GOP structure (i.e., the first GOP structure differs from the third GOP structure), codec, and resolution. Advantageously, the present embodiment allows increased flexibility in the applied encoding method or its attributes, for example to facilitate different client functionalities.

[0034] In some examples, the step of encoding one or more superimposed image frames includes using a scalable video coding (SVC) codec. SVC allows structured information to be transmitted in a layered manner to allow portions of the bitstream to be extracted at a lower bit rate than the complete sequence to be able to decode pictures with multiple image structures (for sequences encoded with spatial scalability), pictures with multiple picture rates (for sequences encoded with temporal scalability), and / or pictures with multiple image quality levels (for sequences encoded with quality scalability (such as signal-to-noise ratio SNR)). Thus, the superimposed image frames encoded in the second video stream can be decoded taking into account the different configurations of the first video stream and the third video stream, so that the superimposed image frames can be decoded according to a configuration that matches the configuration of the encoding of the relevant video stream.

[0035] According to a second aspect of the present invention, the above-mentioned purpose is achieved by a system for encoding one or more video streams, the system comprising: one or more processors; and one or more non-transitory computer-readable media storing instructions executable by the one or more processors, wherein the instructions, when executed, cause the system to perform operations comprising the following steps: receiving a plurality of image frames; determining one or more overlay image frames, each overlay image frame comprising a plurality of switchable overlays, each switchable overlay being associated with an identifier, wherein the overlay image frame has the same size as an image frame in the plurality of image frames; for each image frame in at least some of the plurality of image frames: associating the image frame with an overlay image frame in the one or more overlay image frames; adding metadata to a header of the image frame, wherein, for each switchable overlay in the plurality of switchable overlays, the metadata comprises: position data identifying a position of the switchable overlay in the overlay image frame, size data identifying a size of the switchable overlay in the overlay image frame, and identification data corresponding to the identifier of the switchable overlay; encoding the one or more video streams, wherein the encoding comprises: encoding the plurality of image frames into a first plurality of encoded image frames using a first group of images (GOP) structure; encoding the one or more overlay image frames.

[0036] According to a third aspect of the present invention, the above object is achieved by a non-transitory computer-readable storage medium having instructions stored thereon, which are used to implement the method according to the first aspect when executed on a device with processing capabilities.

[0037] The second and third aspects may generally have the same features and advantages as the first aspect.

[0038] According to a fourth aspect of the present invention, the above-mentioned purpose is achieved by a method for decoding a video stream, the method comprising the following steps: receiving one or more encoded overlay image frames, each encoded overlay image frame including multiple switchable overlays, each switchable overlay being associated with an identifier; receiving multiple encoded image frames in a first encoded video stream; receiving first metadata, the first metadata indicating identifiers of one or more switchable overlays in the multiple switchable overlays visible in the decoded video stream; decoding the multiple encoded image frames into a first plurality of image frames using a GOP structure; and decoding the one or more encoded overlay image frames into one or more overlay image frames, wherein each overlay image frame has the same size as an image frame in the first plurality of image frames.

[0039] The method further includes, when determining that an image frame among the first plurality of image frames is associated with an overlay frame among one or more overlay image frames: extracting second metadata from a header of the image frame, wherein, for each switchable overlay among a plurality of switchable overlays, the second metadata includes: position data identifying a position of the switchable overlay in the overlay image frame, size data identifying a size of the switchable overlay in the overlay image frame, and identification data corresponding to an identifier of the switchable overlay; when the first metadata includes an identifier of the switchable overlay among the plurality of overlays, determining an image frame with overlay by including image data from the image frame and image data from an overlay frame corresponding to the switchable overlay using the position data and size data from the first metadata, and including the image frame with overlay in a decoded video stream.

[0040] Advantageously, a decoder implementing the method may use second metadata from a header of a currently decoded image frame to determine the image frame having these overlays visible, depending on which overlays are to be switched on as indicated in the first metadata. For example, the first metadata may be received from a user of a client (e.g., a smartphone, computer, television, tablet, etc.) that displays the decoded video stream. In other embodiments, the first metadata may be received from a content provider or other separate entity (separate from the user of the display that displays the decoded video stream) that controls which overlays should be visible. For example, the first metadata may be managed by a monitoring system operator who turns over or off overlays on multiple displays depending on factors such as content within the encoded image frame, security clearance of users of the displays, etc.

[0041] By receiving second metadata in the header (identifying the size and position of the corresponding switchable overlay), a low-complexity method for determining which overlays are to be included in the image frame with overlays and identifying which image data should be combined with the visible overlays to form the image frame with overlays as described above can be implemented.

[0042] In some examples, the step of receiving one or more encoded overlay image frames includes: receiving a plurality of incrementally encoded overlay frames, each incrementally encoded overlay frame being associated with an encoded image frame among a plurality of encoded image frames, wherein each pixel block of the incrementally encoded overlay frame that does not include image data of any one of a plurality of switchable overlays is set as a skip block.

[0043] The step of determining an image frame with overlays includes performing incremental frame decoding on an overlay image frame associated with the image frame using the image frame as a reference; wherein the incremental frame decoding includes: when the overlay data does not include an identifier of a switchable overlay among multiple overlays, setting a pixel block corresponding to a switchable graphic overlay as a skip block using position data and size data from second metadata; wherein the image frame with overlays is a display frame.

[0044] Furthermore, each incrementally encoded superimposed image frame is included as auxiliary data in the first encoded video stream, and each of the plurality of image frames is a non-display frame.

[0045] A standard decoder may not forward or transfer an image frame that is designated, marked or labeled as a non-display image frame to an output video stream, such as for display, analysis or storage. Conversely, a standard decoder may forward or transfer an image frame that is designated, marked or labeled as a display image frame to an output video stream, such as for display, analysis or storage. Use of the ND flag (as defined in video coding standards such as H.265, H.266 / VVC, EVC and AV1) may therefore facilitate the use of standard functions of the decoder. Because the image frame decoded from the encoded video stream is an ND frame and the image frame with superposition is a display frame, a decoded video stream including superposition can be achieved. Due to the encoding strategy, in which the superposition image frame is encoded as an incremental frame with reference to the associated image frame, the decoded video stream may include image data from the image frame supplemented with superposition from the superposition image frame. Using the skip block function, only the relevant superposition (e.g., as indicated in the first metadata) is visible in the decoded video stream, where the remaining image data in the decoded video stream will be obtained from the decoded image frame.

[0046] Typically, image frames in video are divided into macroblocks, which are the basic units processed during video compression and encoding. In this approach, instead of integrating multiple overlays (such as text, graphics, or auxiliary images) directly into the image frame itself, they can be encoded as additional macroblocks (auxiliary data) in the same image container as the image data of the image frame.

[0047] This means that the overlay data is processed separately from the main content of the image frame. By encoding the overlay image frame into a further macroblock in the same image container as the image frame, the overlay image frame can be processed independently while still being associated with the corresponding image frame. As described above, using this technique, synchronization and simultaneous processing of the overlay image frame and the associated image frame is facilitated.

[0048] In some examples, one or more encoded overlay image frames are received in a second encoded video stream that is different from the first encoded video stream, wherein the first encoded video stream includes synchronization data indicating an association of each of the one or more overlay image frames with an image frame in the first plurality of image frames.

[0049] As described above, using two different video streams, i.e., the encoded overlay image frames and the encoded image frames (encoded video), can achieve increased flexibility when it comes to encoding and decoding functions and processes. For example, if the overlay stream (second video stream) is simpler or changes less frequently than the main video (first video stream), the overlay stream (second video stream) may require less processing power. This separation can therefore allow the decoder to optimize processing based on the complexity of each stream.

[0050] In some examples, the step of determining an image frame with overlay includes: extracting first image data from an overlay frame corresponding to a switchable overlay; identifying second image data in the image frame to be replaced by the first image data using position data and size data from second metadata; and determining the image frame with overlay by replacing the second pixel block with the first pixel block in the image frame.

[0051] Advantageously, the decoder can efficiently combine the separate video stream and overlay stream into a single composite output. This approach allows the overlay to be dynamically manipulated, being switched on or off without changing the underlying primary video content. It provides a flexible and low complexity approach to video rendering, using the header of the image frame to identify the relevant overlay image data from the overlay image frame, and is particularly useful in scenes where the overlay may be frequently switched on or off.

[0052] According to a fifth aspect of the present invention, the above-mentioned purpose is achieved by a system for decoding a video stream, the system comprising: one or more processors; and a non-temporary computer-readable medium storing one or more instructions executable by the one or more processors, wherein the instructions, when executed, cause the system to perform operations comprising the following steps: receiving one or more encoded overlay image frames, each encoded overlay image frame comprising multiple switchable overlays, each switchable overlay being associated with an identifier; receiving multiple encoded image frames in a first encoded video stream; receiving first metadata, the first metadata indicating identifiers of one or more switchable overlays visible in the decoded video stream among the multiple switchable overlays; decoding the multiple encoded image frames into a first plurality of image frames using a GOP structure; decoding the one or more encoded overlay image frames into one or more overlays; An image frame, wherein each overlay image frame has the same size as an image frame in a first plurality of image frames; when determining that an image frame in the first plurality of image frames is associated with an overlay frame in one or more overlay image frames: extracting second metadata from a header of the image frame, for each switchable overlay in a plurality of switchable overlays, the second metadata including: position data identifying a position of the switchable overlay in the overlay frame, size data identifying a size of the switchable overlay in the overlay frame, and identification data corresponding to an identifier of the switchable overlay; when the first metadata includes an identifier of the switchable overlay in the plurality of overlays, determining an image frame with overlays by including image data from the image frame and image data from an overlay frame corresponding to the switchable overlay using the position data and size data from the second metadata, and including the image frame with overlays in a decoded video stream.

[0053] According to a sixth aspect of the present invention, the above-mentioned purpose is achieved by a non-transitory computer-readable storage medium, on which instructions are stored, and the instructions are used to implement the method according to the fourth aspect when executed on a device with processing capabilities.

[0054] The fifth and sixth aspects may generally have the same features and advantages as the fourth aspect.

[0055] It is further noted that, unless explicitly stated otherwise, the present disclosure relates to all possible combinations of features. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] The above and further objects, features and advantages of the present invention will be better understood through the following illustrative and non-limiting detailed description of embodiments of the present disclosure with reference to the accompanying drawings, in which like reference numerals will be used for similar elements, and in which:

[0057] Figure 1 showing an image frame including an overlay according to an embodiment;

[0058] Figure 2 shows a representation of an image container including a header, image frame data, and overlaid image frame data as auxiliary data according to an embodiment;

[0059] Figure 3 shows a coding scheme for an image frame and a superimposed image frame according to a first embodiment;

[0060] Figure 4 A first image container including a header and image frame data is shown according to an embodiment, the first image container being associated with a second image container including overlay image frame data;

[0061] Figure 5 shows a coding scheme for an image frame and a superimposed image frame according to a second embodiment;

[0062] Figure 6 shows a coding scheme for an image frame and a superimposed image frame according to a third embodiment;

[0063] Figure 7 A flowchart illustrating a method of encoding one or more video streams according to an embodiment;

[0064] Figure 8 A flow chart showing a method of decoding a video stream according to an embodiment. DETAILED DESCRIPTION

[0065] In surveillance applications, overlays such as bounding boxes, privacy masks, and text overlays indicating date and time can provide significant enhancements. Bounding boxes can be beneficial for surveillance applications to track and highlight the movement of objects or individuals within the camera's field of view, facilitating real-time security analysis. Privacy masks can be helpful to blur sensitive areas or maintain the anonymity of individuals, ensuring compliance with privacy laws and ethical standards. Text overlays showing date and time can be helpful to contextualize footage, making it easier to catalog and reference specific events during reviews and investigations. These overlays, as well as other appropriate overlays, can not only enhance the functionality of surveillance systems, but can also contribute to more effective and responsible surveillance practices.

[0066] Figure 1 An image frame 100 depicting a scene is shown by way of example. The image frame 100 includes a plurality of overlays 102, 104, 106. The plurality of overlays in this example include two bounding box overlays, each of which includes a respective bounding box 104 indicating the location and size of a respective vehicle in the scene. The plurality of overlays further include a text overlay 102 indicating the capture time of the image frame 100. The plurality of overlays further include a privacy mask overlay 106 obscuring the identity of pedestrians in the scene.

[0067] The ability to individually toggle these overlays on and off provides additional flexibility and control in surveillance situations. For example, security personnel may choose to activate bounding boxes for enhanced tracking during high-security scenarios, while disabling them to reduce screen clutter during normal operations. Privacy masks may be enabled or disabled based on the viewer's security clearance level. Higher clearance levels may allow individuals to view unmasked footage, revealing sensitive areas or information, while lower clearance levels will cause those areas to be blurred. During live surveillance, date and time overlays may be turned off for a clearer view, but enabled during playback for precise event tracking. This customization allows the surveillance system to adapt to different requirements, balancing the need for detailed information with the clarity and simplicity of the video feed. It also enables users to focus on specific aspects of the video data as needed, optimizing the utilization and efficiency of the video processing system.

[0068] Implementing the ability to switch overlays in video encoding and decoding systems introduces significant challenges, particularly in terms of synchronization and bandwidth management. Synchronization is critical to ensure that overlays correspond accurately to the relevant frames and are properly aligned, especially in real-time or rapidly changing video feeds. Any lag or misalignment between overlays and video frames can lead to confusion or misinterpretation. Additionally, when implementing solutions that include the possibility of individually switching multiple overlays on and off, it is critical to effectively manage bandwidth, as the different possible configurations can significantly increase the data load. This is particularly important in systems where bandwidth limitations or streaming quality are a concern.

[0069] Next, combine Figures 2 to 8 Implementations are discussed that may eliminate or reduce the above concerns in encoders and decoders.

[0070] Figure 7 An example of a method 700 for encoding one or more video streams is illustrated. The method includes receiving S702 a plurality of image frames. These image frames may originate from various sources. For example, they may be captured directly from a camera recording a scene (such as a surveillance camera monitoring a public space or a traffic camera monitoring a busy intersection). In some settings, the encoding device or encoder is integrated into the camera itself, ensuring that the captured frames are processed immediately. Alternatively, in a scenario where the encoder and the camera are separate units, the encoder may receive these frames wirelessly and perform real-time or near real-time encoding using streaming media technology. Another possible source of these image frames is a storage reservoir in which previously recorded or archived video material is stored, ready for encoding.

[0071] The method 700 includes determining S704 one or more superimposed image frames, wherein each superimposed image frame includes a plurality of switchable overlays. Some overlays of the superimposed image frames may be determined by analyzing the image frames and determining the overlays using the analysis metadata. For example, the image frames may be analyzed to detect objects in the image frames, and it may be determined that the overlays include information about the objects, such as object type, bounding box, speed, etc.

[0072] In addition to analytics-based overlays, some overlays can be determined by receiving state information from sensors located in or associated with the scene. These sensors can provide a range of data, like the time and date, temperature readings in industrial settings, the open / closed state of doors in buildings, or even environmental conditions like smoke or gas levels. Incorporating this sensor data into the overlay enriches the video stream with valuable contextual information, enhancing the overall understanding of the monitored scene.

[0073] In some examples, the overlay of overlaid image frames is created by converting a received image file into an overlay, for example, decoded from a JPEG, GIF, or any other suitable image format.

[0074] Each overlay will be associated with an identifier that is used when referencing the overlay. The identifier can take a variety of forms, depending on the requirements of the system and the specific implementation strategy. A straightforward approach is to use a counter that increments with each new overlay in the overlay image frame. For example, the first overlay added to the overlay image frame is assigned an identifier of "1", the next overlay is assigned an identifier of "2", and so on. Another approach may involve more complex identifiers, such as a hash value calculated based on the overlay's data. This approach provides each overlay with a unique fingerprint that is derived from its visual content or size and position data, for example. For example, an overlay containing a bounding box around a vehicle may generate a specific hash value based on the size, position, and / or other characteristics of the bounding box. Such an identifier facilitates more complex management, such as quickly checking for duplicate overlays or retrieving a specific overlay based on its visual content.

[0075] For at least some of the plurality of image frames, the method continues by associating the image frame with an overlay image frame of the one or more overlay image frames S706. In some cases, all image frames are associated with corresponding overlay image frames. In some cases, only the incrementally encoded image frames are associated with overlay image frames, as described below in conjunction with Figure 3 It should be noted that more than one image frame may be associated with the same overlay image frame, for example if the overlay of the overlay image frame does not change during the time period of the captured video stream.

[0076] When the image frame has been associated with an overlay image frame, information about each of the plurality of overlays is added to a header of the image frame. The method 700 comprises adding S708 metadata to a header of the image frame, wherein, for each of the plurality of switchable overlays, the metadata comprises: position data identifying a position of the switchable overlay in the overlay image frame, size data identifying a size of the switchable overlay in the overlay image frame, and identification data corresponding to an identifier of the switchable overlay.

[0077] The method further comprises checking S710 whether all image frames that should be associated with the superimposed image frame have been processed. Otherwise, steps S706 and S708 are performed again for the unprocessed image frames.

[0078] Figure 2 and Figure 4 Two different ways of associating an image frame with an overlay image frame and adding metadata to the header of the image frame are shown.

[0079] Figure 2 The representation of the image container 200 is shown by way of example. The image container includes a Figure 1 The image container 200 includes data 202, 204, 206 of the image frame 100 in order to facilitate the switchable overlay 102, 104, 106. Figure 1 The image container further includes auxiliary data 206 corresponding to the overlay image frame 206 (which has the same size as the image frame 204). The overlay image frame 206 includes a plurality of switchable overlays 102, 104, 106 located at positions in the overlay image frame 106 corresponding to positions in the image frame 204, such that superimposing the overlay image frame 206 on the image frame 204 will result in Figure 1 Image frame 100 in.

[0080] Figure 2 Each switchable overlay 102, 104, 106 in the image container 200 is associated with an identifier (not shown). The image container 200 further includes a header 202. The header 202 includes metadata (not shown), and for each switchable overlay 102, 104, 106 in the plurality of switchable overlays, the metadata includes: position data identifying the position of the switchable overlay in the overlay image frame, size data identifying the size of the switchable overlay in the overlay image frame, and identification data corresponding to the identifier of the switchable overlay. For example, the header may include the following data represented in an appropriate manner:

[0081] ID Position (x,y) Dimensions (x,y) 1 20,10 80,20 2 15,150 60,50 3 110,150 60,50 4 300,100 20,20

[0082] Table 1

[0083] In Table 1, ID 1 corresponds to text overlay 102 , ID 2 corresponds to left bounding box overlay 104 , ID 3 corresponds to right bounding box overlay 104 , and ID 4 corresponds to privacy mask overlay 106 .

[0084] The header 202 may further include specific information to distinguish the overlay image frame data 206 from the main video content (image frame 204). This may be in the form of an identifier or marker marking the beginning and end of the overlay image frame data 206 and / or image frame data 204 within the image container 200.

[0085] Figure 4 The representation of the first image container 402 and the second image container 404 is shown by way of example. The first image container 402 and the second image container 404 in combination include a Figure 1 The data 202, 204, 206 of the image frame 100 in the embodiment of the present invention are used to facilitate the switchable overlay 102, 104, 106. However, unlike Figure 2 Different, in Figure 4 In the embodiment, the image frame 204 and the header 202 are provided in a first image container 402, and the overlay image frame 206 is provided in a second image container 404. However, the first image container 402 is associated 406 with the second image container 404. This association 406 can be implemented in any suitable manner, such as associating the identifiers of the image container 402 and the image container 404 in a suitable data structure, and combining the image container 402 and the identifier of the image container 404. Figure 5 and Figure 6 Describe the example.

[0086] Return to Figure 7 When it is determined S710 that all image frames that should be associated with the overlay image frame have been processed, encoding of one or more video streams is performed. The encoding includes encoding S714 a plurality of image frames into a first plurality of encoded image frames using a first group of images (GOP) structure, and encoding S714 the one or more overlay image frames.

[0087] Figure 3 , Figure 5 and Figure 6 An encoding scheme for encoding a plurality of image frames and one or more superimposed image frames according to an embodiment is shown by way of example.

[0088] Figure 3 By way of example, for example Figure 2 The encoding scheme for the image frames and the superimposed image frames is explained in . Figure 3 , the encoding scheme results in a single video stream 300 corresponding to both the encoded image frames 302a . . . 302n and the encoded overlay image frames 304a . . . 304n.

[0089] A plurality of image frames (eg, including Figure 2 The image frame 204 in the image is encoded into a first plurality of encoded image frames 302a ... 302n. The plurality of encoded image frames 302a ... 302n are Figure 3 Each of the plurality of image frames is set as a non-display frame.

[0090] Examples include Figure 2 One or more of the superimposed image frames 206 in the image data are encoded into one or more encoded superimposed image frames 304a ... 304n. The one or more encoded superimposed image frames are encoded in Figure 3 The overlay image frame is set as the display frame.

[0091] exist Figure 3 In an embodiment, for each incrementally encoded image frame (denoted as "P") in a plurality of encoded image frames 302a...302n, an overlay frame associated with the incrementally encoded image frame is determined and encoded using incremental frame encoding (P / B frame encoding) into an encoded overlay frame that references the incrementally encoded image frame. The incremental frame encoding includes setting each pixel block of the encoded overlay frame that does not correspond to any of a plurality of switchable overlays to a skip block. Using this technique, a conventional decoder that does not implement the functionality of reading a header to understand which portions of the overlay image frame are associated with which overlay can nonetheless provide a decoded image frame that includes all of the plurality of overlays.

[0092] In some embodiments, incrementally encoding the superimposed image frame includes using only the incrementally encoded image frame as a reference, i.e., encoding the superimposed image frame as a P frame. In some examples, incrementally encoding the superimposed image frame may also include referencing an earlier encoded superimposed image frame, such as Figure 3 Using previously encoded overlay image frames as references can increase compression, especially for static overlays.

[0093] exist Figure 3 In an embodiment, each encoded overlay frame 304a ... 304n is included in the first video stream as auxiliary data associated with the incrementally encoded image frame to which it refers. Figure 2 As described, the same image container may be used for both the encoded image frame data and the encoded overlay image frame data. Figure 2 The header 202 of the image container described in may include reference frame data indicating a specific frame from which changes to the overlaid image frames 304a . . . 304n are calculated.

[0094] In some examples, the encoded image frames 302a ... 302n and the encoded overlay image frames 304a ... 304n are provided in separate video streams. Figure 5 Such an embodiment is shown.

[0095] exist Figure 5 In the embodiment, a plurality of encoded image frames 302a ... 302n are included in a first video stream 502, and one or more encoded superimposed frames 304a ... 304n are included in a second video stream 504. In this embodiment, the step of associating an image frame with a superimposed image frame in one or more superimposed image frames includes including synchronization data 506 in the first video stream 502 to associate the image frame with a superimposed image frame in one or more superimposed image frames. The synchronization data 506 may include any appropriate data structure. One method is to use a timestamp. Each image frame and its corresponding superimposed frame may be marked with a timestamp indicating the exact time when they should be displayed. This ensures that the superimposition is synchronized with the correct frame in the video stream during playback. Another method is to use a frame counter. Each frame in both the video stream and the superimposed stream may be assigned a sequence number. The decoder uses these numbers to match the superimposed frame with the corresponding video frame. Metadata tags within the video stream 502 may be used to transmit the synchronization data 506. These tags may contain information such as a unique identifier or pointer to a specific superimposed image frame.

[0096] The synchronization data may be divided into separate data blocks, for example, a portion of the synchronization data 506 may be included in the first video stream 502 after or in conjunction with each encoded image frame 302a ... 302n. In other examples, the synchronization data 506 may be included in the first video stream in conjunction with each GOP of the encoded image frames 302a ... 302n.

[0097] like Figure 5 As shown, in an example, one encoded superimposed image frame 304a...304n may be associated with more than one encoded image frame 302a...302n. In other examples, a one-to-one correspondence between the encoded superimposed image frames 302a...302n and the encoded image frames 304a...304n is achieved. The superimposed image frame may be encoded using a GOP different from the GOP used when encoding the multiple image frames into the multiple encoded image frames 302a...302n. In some examples, a new encoded superimposed image frame 304a...304n is encoded only when a change in one of the multiple superimposed images in the superimposed image frame is detected.

[0098] Separating the video streams of the coded image frames and the coded overlay image frames provides significant flexibility and efficiency, particularly when sharing the overlay image frames across different video streams. This approach allows the same set of overlays to be used with multiple video content streams, each of which may be encoded differently based on specific requirements or constraints. For example, it may be advantageous to encode the multiple image frames into several video streams, where encoding the multiple image frames into a first plurality of coded image frames differs from encoding the multiple image frames into a second plurality of coded image frames in at least one of the following aspects: encoding quality, frame rate, GOP structure, codec, and resolution.

[0099] For example, consider a scenario where the same video clip needs to be broadcast in two different formats: one stream is a high-definition stream for high-bandwidth environments, and the other is a low-resolution stream for limited bandwidth situations. The high-definition stream may have higher encoding quality, frame rate, and resolution, while the low-resolution stream is optimized to reduce data consumption. By placing the overlay image frames in separate streams, these overlays can be applied to both video streams without having to re-encode each video stream. This not only saves encoding time and resources, but also ensures consistency of overlay content between different versions of the video.

[0100] Furthermore, this separation allows for greater flexibility in modifying the video streams independently. For example, if the codec or GOP structure of one video stream needs to be changed to be compatible with certain playback systems, this can be done without affecting the overlay stream or other video streams. It also facilitates dynamic adaptation of the video streams to different network conditions or device capabilities, because the overlay stream remains constant and compatible during these changes.

[0101] Figure 6 Here, a plurality of image frames are encoded into a second plurality of encoded image frames 604a ... 604n using a third GOP structure. Figure 6 The third GOP structure indicated in is different from the (first) GOP structure used to encode the plurality of image frames into the first plurality of encoded image frames 302a ... 302n, but this is only an example. In other examples, the two GOP structures are the same.

[0102] exist Figure 6, the same synchronization data 506 is included in both the first video stream 502 and the third video stream 602. However, in some examples, the synchronization data included in the third video stream may be adjusted based on the difference in encoding the plurality of image frames into the first plurality of encoded image frames 302a ... 302n and encoding the plurality of image frames into the second plurality of encoded image frames 604a ... 604n. For example, if the low-resolution stream (e.g., video stream 602) has a lower frame rate, the synchronization data included in the video stream 602 (like a timestamp or frame counter) may need to take into account the reduced number of frames. This may involve adjusting the timestamp or frame counter to align with the reduced frame rate.

[0103] In an example, encoding one or more overlay image frames includes using a scalable video coding (SVC) codec. Using SVC to encode overlay image frames is a strategic approach, especially when processing multiple streams 502, 602 of video data encoded at different resolutions. SVC is an extension of the H.264 / AVC standard and is designed to provide video streams that are easily adapted to different bandwidths and display resolutions. When overlay image frames are encoded using SVC, it allows these overlays to adapt to the different resolutions of video streams 502, 602. For example, if there are two streams, one is high definition and the other is standard definition, then SVC-encoded overlay image frames can be appropriately applied to the two streams 502, 602. The overlay will be correctly aligned with the resolution of each stream 502, 602, ensuring that it appears consistently and clearly, regardless of the underlying video resolution.

[0104] It should be noted that in some embodiments, the strategy of sharing overlay image frames across different video streams can also be applied to video streams originating from different groups of images. For example, for certain types of overlays, such as overlays relating to date and time or other general metadata relating to the scene captured by different groups of images (i.e., one group from a first camera and a second group from a second camera), overlays can be shared between video streams originating from different groups of images. In these examples, Figure 7The method described in may be modified to further include: receiving a second plurality of image frames; for each image frame of at least some of the second plurality of image frames: associating the image frame with an overlay image frame of one or more overlay image frames; adding metadata to a header of the image frame, wherein, for each switchable graphic overlay of a plurality of switchable graphic overlays, the metadata includes: position data identifying a position of the switchable graphic overlay in the overlay image frame, size data identifying a size of the switchable graphic overlay in the overlay image frame, and identification data corresponding to an identifier of the switchable graphic overlay. The encoding may then include encoding the second plurality of image frames into a second plurality of encoded image frames using a third GOP structure, and including the second plurality of encoded image frames in a third video stream. The step of associating the image frames of the second plurality of image frames with the overlay image frames of the one or more overlay image frames may then include: including second synchronization data in the third video stream to associate the image frames of the second plurality of image frames with the encoded overlay image frames of the one or more encoded overlay image frames.

[0105] Figure 8 A decoding method 800 for decoding a video stream encoded as described herein is described.

[0106] The method 800 comprises receiving S802 one or more encoded overlay image frames, each encoded overlay image frame comprising a plurality of switchable overlays, each switchable overlay being associated with an identifier.

[0107] The method 800 further comprises receiving S804 a plurality of encoded image frames in a first encoded video stream, and receiving S806 first metadata indicating identifiers of one or more switchable overlays in the plurality of switchable overlays visible in the decoded video stream. Such metadata may be received, for example, from a user of a device displaying the decoded video stream.

[0108] The method 800 further comprises decoding S808 the plurality of encoded image frames into a first plurality of image frames using the GOP structure. The method 800 further comprises decoding S808 one or more encoded overlay image frames into one or more overlay image frames, wherein each overlay image frame has the same size as an image frame in the first plurality of image frames.

[0109] The method 800 further includes decoding S810 one or more encoded overlay image frames.

[0110] Specifically, decoding S810 includes, when determining that an image frame in the first plurality of image frames is associated with an overlay frame in the one or more overlay image frames:

[0111] 1) extracting S812 second metadata from a header of the image frame, the second metadata comprising, for each switchable overlay of the plurality of switchable overlays: position data identifying a position of the switchable overlay in the overlay frame, size data identifying a size of the switchable overlay in the overlay frame, and identification data corresponding to an identifier of the switchable overlay;

[0112] 2) When the first metadata includes an identifier of a switchable overlay among multiple overlays, determine S814 an image frame with overlay by including image data from the image frame and image data from an overlay frame corresponding to the switchable overlay using position data and size data from the second metadata, and include the image frame with overlay in the decoded video stream.

[0113] As long as it is determined S816 that there are more image frames in the plurality of decoded image frames to add overlay, ie more image frames of the first plurality of image frames are associated with overlay frames in the one or more overlay image frames, steps S812 and S814 are performed.

[0114] As above combined Figure 3 and Figure 5 to Figure 6 As described, the plurality of encoded image frames and the one or more encoded superimposed image frames may be received by the decoder in the same video stream or in different video streams. In an embodiment of receiving a single video stream, the step S804 of receiving the one or more encoded superimposed image frames may include receiving a plurality of incrementally encoded superimposed frames, each incrementally encoded superimposed frame being associated with an encoded image frame in the plurality of encoded image frames, wherein each pixel block of the incrementally encoded superimposed frame that does not include image data of any one of the plurality of switchable superimpositions is set as a skip block. Moreover, each of the plurality of decoded image frames is a non-display frame as defined by the encoder, see above.

[0115] In a 1-stream embodiment, wherein each incrementally encoded overlay image frame is included as auxiliary data in the encoded video stream, the step of determining S814 an image frame having an overlay includes performing incremental frame decoding of an overlay image frame associated with the image frame using the image frame as a reference. During incremental frame decoding, when the overlay data does not include an identifier of a switchable overlay of a plurality of overlays, a pixel block corresponding to the switchable overlay is set as a skip block using position data and size data from the second metadata. Thus, the decoded overlay image frame will include one or more overlays that should be visible according to the first metadata, and image data from the reference image frame. The decoded overlay image frame is a display frame as defined by the encoder, see above.

[0116] In a 2-stream implementation, one or more encoded superimposed image frames are received in a second encoded video stream different from the first encoded video stream, wherein the first encoded video stream includes synchronization data indicating an association of each of the one or more superimposed image frames with an image frame in the first plurality of image frames. In this implementation, the step of determining S814 having superimposed image frames includes:

[0117] 1) extracting first image data from an overlay frame corresponding to a switchable overlay;

[0118] 2) identifying second image data in the image frame to be replaced by the first image data using the position data and the size data from the second metadata; and

[0119] 3) Determining an image frame with superposition by replacing the second pixel block with the first pixel block in the image frame.

[0120] In the 2-stream embodiment, similar to the 1-stream embodiment, the decoded image frame is set not to be displayed, and the image frame with superimposition is set to be displayed.

[0121] As above combined Figure 6 As described, the 2-stream implementation may be modified to define an X-stream implementation, where X > 2. The above description of decoding the 2-stream implementation may then apply mutatis mutandis.

[0122] In an example, an encoder (and a similar decoder) implementing the encoding method as described herein may be implemented in a single device. The encoder may be implemented, for example, in a camera. In other examples, some encoding / decoding functions may be implemented in a server, in a cloud, or in a separate device. In general, a device (camera, server, etc.) implementing the encoding / decoding method described herein may include a circuit configured to implement the encoding / decoding functions described herein. The described functions may advantageously be implemented in one or more computer programs executable on a programmable system, the programmable system including at least one programmable processor connected to receive data and instructions from a data storage system, at least one input device (such as a camera) and at least one output device (such as a display), and to transmit data and instructions thereto. Suitable processors for executing instruction programs include, for example, both general-purpose processors and special-purpose microprocessors, and a single processor or multiple processors or one of the cores of any type of computer. The processor may be supplemented or incorporated into an ASIC (application-specific integrated circuit).

[0123] The above embodiments should be understood as illustrative examples of the present invention. Further embodiments of the present invention are contemplated. For example, an image frame capturing a scene may include 2D data or 3D data. It should be understood that any feature described with respect to any one embodiment may be used alone or in combination with other features described, and may also be used in combination with one or more features of any other embodiment or any combination of any other embodiment. In addition, equivalents and modifications not described above may also be adopted without departing from the scope of the present invention as defined in the appended claims.

Claims

1. A method for encoding one or more video streams, the method comprising the following steps: receiving a plurality of image frames; determining one or more overlay image frames, each overlay image frame comprising a plurality of individually switchable overlays, each switchable overlay being associated with an identifier, wherein the overlay image frame has the same size as an image frame of the plurality of image frames; for each image frame of at least some of the plurality of image frames: associating the image frame with an overlay image frame of the one or more overlay image frames; adding metadata to a header of the image frame, wherein for each switchable overlay of the plurality of switchable overlays, the metadata comprises: position data identifying a position of the switchable overlay in the overlay image frame, size data identifying a size of the switchable overlay in the overlay image frame, and identification data corresponding to an identifier of the switchable overlay; The one or more video streams are encoded, wherein the encoding comprises: encoding the plurality of image frames into a first plurality of encoded image frames using a first group of images (GOP) structure; and encoding the one or more superimposed image frames.

2. The method according to claim 1, wherein: The encoded image frame and the encoded superimposed frame are included in a first video stream, wherein the method further comprises setting each of the plurality of image frames as a non-display frame; wherein the encoding comprises: Use a block-based codec that supports skipping blocks; For each incrementally encoded image frame among the first plurality of encoded image frames: determining an overlay image frame among the one or more overlay image frames; setting the overlay image frame as a display frame; encoding the overlay image frame incremental frame into an encoded overlay frame with reference to the incrementally encoded image frame, wherein each pixel block of the encoded overlay frame that does not correspond to any of the plurality of switchable overlays is set as a skip block; and including the encoded overlay frame in the first video stream as auxiliary data associated with the incrementally encoded image frame.

3. The method according to claim 1, wherein: The first plurality of encoded image frames are included in a first video stream, and the encoded one or more overlay frames are included in a second video stream, wherein the step of associating the image frame with an overlay image frame of the one or more overlay image frames comprises: Synchronization data is included in the first video stream to associate the image frame with an overlaid image frame of the one or more overlaid image frames.

4. The method according to claim 3, wherein: The step of encoding one or more superimposed image frames includes using a second GOP structure different from the first GOP structure.

5. The method according to claim 3, further comprising the steps of: encoding the plurality of image frames into a second plurality of encoded image frames using a third GOP structure; including the second plurality of encoded image frames in a third video stream; as well as The synchronization data is included in the third video stream.

6. The method according to claim 5, wherein: Encoding the plurality of image frames into the first plurality of coded image frames differs from encoding the plurality of image frames into the second plurality of coded image frames in at least one of the following aspects: Encoding quality, frame rate, GOP structure, codec and resolution.

7. The method according to claim 6, wherein: The step of encoding the one or more overlay image frames includes using a Scalable Video Coding (SVC) codec.

8. A system for encoding one or more video streams, the system comprising: one or more processors; and One or more non-transitory computer-readable media storing instructions executable by the one or more processors, wherein the instructions, when executed, cause the system to perform operations comprising: receiving a plurality of image frames; determining one or more overlay image frames, each overlay image frame comprising a plurality of individually switchable overlays, each switchable overlay being associated with an identifier, wherein the overlay image frame has the same size as an image frame of the plurality of image frames; for each image frame of at least some of the plurality of image frames: associating the image frame with an overlay image frame of the one or more overlay image frames; adding metadata to a header of the image frame, wherein for each switchable overlay of the plurality of switchable overlays, the metadata comprises: position data identifying a position of the switchable overlay in the overlay image frame, size data identifying a size of the switchable overlay in the overlay image frame, and identification data corresponding to an identifier of the switchable overlay; The one or more video streams are encoded, wherein the encoding comprises: encoding the plurality of image frames into a first plurality of encoded image frames using a first group of images (GOP) structure; and encoding the one or more superimposed image frames.

9. A non-transitory computer-readable storage medium having stored thereon instructions for implementing the method of claim 1 when executed on one or more devices having processing capabilities.

10. A method for decoding a video stream, comprising the following steps: receiving one or more encoded overlay image frames, each encoded overlay image frame comprising a plurality of individual switchable overlays, each switchable overlay being associated with an identifier; receiving a plurality of encoded image frames in a first encoded video stream; receiving first metadata indicating identifiers of one or more switchable overlays of the plurality of switchable overlays visible in a decoded video stream; decoding the plurality of encoded image frames into a first plurality of image frames using a GOP structure; decoding the one or more encoded overlay image frames into one or more overlay image frames, wherein each overlay image frame has the same size as an image frame in the first plurality of image frames; Upon determining that an image frame of the first plurality of image frames is associated with an overlay frame of the one or more overlay image frames: extracting second metadata from a header of the image frame, the second metadata comprising, for each switchable overlay of the plurality of switchable overlays: position data identifying a position of the switchable overlay in the overlay frame, size data identifying a size of the switchable overlay in the overlay frame, and identification data corresponding to an identifier of the switchable overlay; When the first metadata includes an identifier of a switchable overlay among the multiple overlays, an image frame with overlay is determined by including image data from the image frame and image data from an overlay frame corresponding to the switchable overlay using the position data and the size data from the second metadata, and the image frame with overlay is included in the decoded video stream.

11. The method according to claim 10, wherein: The step of receiving one or more encoded overlay image frames comprises: receiving a plurality of incrementally encoded overlay frames, each incrementally encoded overlay frame being associated with an encoded image frame of the plurality of encoded image frames, wherein each pixel block of the incrementally encoded overlay frames not including image data of any of the plurality of switchable overlays is set as a skip block; wherein the step of determining the image frame with overlays comprises: performing incremental frame decoding of an overlay image frame associated with the image frame using the image frame as a reference, wherein when the first metadata does not include an identifier of a switchable overlay of the plurality of overlays, the incremental frame decoding comprises setting a pixel block corresponding to the switchable overlay as a skip block using the position data and the size data from the second metadata; wherein the image frame with overlays is a display frame; and Each incrementally encoded superimposed image frame is included as auxiliary data in the first encoded video stream, wherein each of the plurality of image frames is a non-display frame.

12. The method according to claim 10, wherein: The one or more encoded overlay image frames are received in a second encoded video stream that is different from the first encoded video stream, wherein the first encoded video stream includes synchronization data that indicates an association of each of the one or more overlay image frames with an image frame in the first plurality of image frames.

13. The method according to claim 12, wherein: The steps of determining an image frame with an overlay include: extracting first image data from an overlay frame corresponding to the switchable overlay; identifying second image data in the image frame to be replaced by the first image data using the position data and the size data from the second metadata; and The image frame with overlay is determined by replacing the second pixel block with the first pixel block in the image frame.

14. A system for decoding a video stream, the system comprising: one or more processors; and One or more non-transitory computer-readable media storing instructions executable by the one or more processors, wherein the instructions, when executed, cause the system to perform operations comprising: receiving one or more encoded overlay image frames, each encoded overlay image frame comprising a plurality of individual switchable overlays, each switchable overlay being associated with an identifier; receiving a plurality of encoded image frames in a first encoded video stream; receiving first metadata indicating identifiers of one or more switchable overlays of the plurality of switchable overlays visible in the decoded video stream; decoding the plurality of encoded image frames into a first plurality of image frames using a GOP structure; decoding the one or more encoded overlay image frames into one or more overlay image frames, wherein each overlay image frame has the same size as an image frame in the first plurality of image frames; Upon determining that an image frame of the first plurality of image frames is associated with an overlay frame of the one or more overlay image frames: extracting second metadata from a header of the image frame, the second metadata comprising, for each switchable overlay of the plurality of switchable overlays: position data identifying a position of the switchable overlay in the overlay frame, size data identifying a size of the switchable overlay in the overlay frame, and identification data corresponding to an identifier of the switchable overlay; When the first metadata includes an identifier of a switchable overlay among the multiple overlays, an image frame with overlay is determined by including image data from the image frame and image data from an overlay frame corresponding to the switchable overlay using the position data and the size data from the second metadata, and the image frame with overlay is included in the decoded video stream.

15. A non-transitory computer-readable storage medium having stored thereon instructions for implementing the method of claim 10 when executed on one or more devices having processing capabilities.