Language and purpose information for text comments in a video bitstream using supplemental enhancement information messages
By embedding and extracting text data into video bitstreams, this method solves the problem of effectively embedding and extracting text data in existing technologies, thereby improving the diversity of video content applications and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-09
- Publication Date
- 2026-08-04
AI Technical Summary
Existing video encoding schemes struggle to effectively embed and extract text data, impacting the diverse applications of video content and user experience.
Text data embedding and extraction are achieved by packaging text data into Supplemental Enhancement Information (SEI) messages and inserting them into the video bitstream.
It enables the effective embedding and extraction of text data during video encoding and decoding, enhancing the diverse applications of video content and user interaction capabilities.
Smart Images

Figure CN122514965A_ABST
Abstract
Description
[0001] Cross-reference to related applications This application claims priority to European Application No. 24305046.5, filed on 9 January 2024, and European Application No. 24305554.8, filed on 8 April 2024, which are incorporated herein by reference in their entirety. Technical Field
[0002] At least one of these embodiments generally relates to a method and apparatus for embedding text data and language or purpose information into a video bitstream and retrieving text data from the video bitstream, the text data being embedded using supplemental enhancement information messages. Background Technology
[0003] To achieve high compression efficiency, image and video coding schemes typically employ prediction and transform to utilize spatial and temporal redundancy in video content. Generally, intra-frame or inter-frame prediction is used to leverage intra-frame or inter-frame correlations, followed by transform, quantization, and entropy coding of the difference between the original and predicted image blocks (typically represented as prediction error or prediction residual). During encoding, the original image blocks are often segmented / split into sub-blocks, for example, using quadtree segmentation. To reconstruct the video, the compressed data is decoded through the inverse process corresponding to prediction, transform, quantization, and entropy coding. Summary of the Invention
[0004] According to a first aspect of at least one embodiment, a method includes: obtaining encoded video data and text data; packaging the text data and additional data into a supplemental enhancement information message; inserting the supplemental enhancement information message into a bitstream including the encoded video data; and providing the bitstream.
[0005] According to a second aspect of at least one embodiment, a method includes: obtaining a bitstream comprising encoded video data and a supplemental enhancement information message, the supplemental enhancement information message comprising text data and additional data; extracting text data from the supplemental enhancement information message; and providing the extracted text data.
[0006] According to a third aspect of at least one embodiment, an apparatus includes one or more processors configured to: acquire encoded video data and text data; package the text data and additional data into a supplemental enhancement information message; insert the supplemental enhancement information message into a bitstream including the encoded video data; and provide the bitstream.
[0007] According to a fourth aspect of at least one embodiment, an apparatus includes one or more processors configured to: acquire a bitstream including encoded video data and a supplemental enhancement information message, the supplemental enhancement information message including text data and additional data; extract text data from the supplemental enhancement information message; and provide the extracted text data.
[0008] One or more embodiments of this invention also provide a computer-readable storage medium storing instructions for encoding or decoding video data according to at least a portion of any of the methods described above. One or more embodiments also provide a computer-readable storage medium storing a bitstream generated according to the encoding methods described above. One or more embodiments also provide a computer program product including instructions for performing at least a portion of any of the methods described above. Attached Figure Description
[0009] Figure 1 A block diagram illustrating an example of a system in which various aspects and embodiments are implemented is shown.
[0010] Figure 2 Examples of contexts in which the following embodiments can be implemented are described.
[0011] Figure 3 A block diagram illustrating an example of a video encoder is shown.
[0012] Figure 4 A block diagram illustrating an example of a video decoder is shown.
[0013] Figure 5 The illustration shows an example of segmentation performed on images from an original video sequence.
[0014] Figure 6 An example of overlapping text SEI messages according to at least one embodiment is illustrated.
[0015] Figure 7 An example structure of a video bitstream including a text SEI message is illustrated according to at least one embodiment.
[0016] Figure 8A An example process for encoding a bitstream including a text SEI message carrying text data, according to an embodiment, is illustrated.
[0017] Figure 8B The illustration depicts an example process for decoding a bitstream including a text SEI message carrying text data, according to an embodiment. Detailed Implementation
[0018] Various embodiments involve embedding text data into a bitstream that includes encoded video by packaging the text data into an SEI message and inserting it into the bitstream.
[0019] Although the principles are described in the context of the VVC (Video Coding Universal) or HEVC (High-Efficiency Video Coding) specifications, this aspect is not limited to these video coding standards and can be applied to, for example, other standards and recommendations, as well as any extensions of such standards and recommendations (including VVC and HEVC). Unless otherwise indicated or technically excluded, the aspects described in this application may be used individually or in combination.
[0020] Figure 1 The diagram illustrates an example of a system implementing various aspects and embodiments therein. System 1000 can be implemented as a device or apparatus including the various components described below and configured to perform one or more aspects of the aspects described in this document. Examples of such devices include, but are not limited to, various electronic devices such as computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, vehicle entertainment systems, vehicle control systems, drones, video surveillance cameras, and more generally, data servers. Elements of system 1000 can be implemented individually or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 1000 are distributed across multiple ICs and / or discrete components. In various embodiments, system 1000 is communicatively coupled to one or more other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various embodiments, system 1000 is configured to implement one or more aspects of the aspects described in this document.
[0021] System 1000 includes at least one processor 1010 configured to execute instructions loaded therein for implementing various aspects as described in this document. Processor 1010 may include embedded memory, input / output interfaces, and various other circuitry as known in the art. System 1000 includes at least one memory 1020 (e.g., a volatile memory device and / or a non-volatile memory device). System 1000 includes a storage device 1040 that may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash-based memory, disk drives, solid-state drives (SSDs), and / or optical disk drives. As a non-limiting example, storage device 1040 may include internal storage devices, attached storage devices (including removable and non-removable storage devices), and / or network-accessible storage devices (also known as cloud storage devices).
[0022] System 1000 includes an encoder / decoder module 1030 configured to, for example, process data to provide encoded or decoded video, and the encoder / decoder module 1030 may include its own processor and memory. The encoder / decoder module 1030 represents a module that can be included in a device to perform one or more of the encoding and / or decoding functions described further below. It is well known that a device may include one or both encoding and decoding modules. Additionally, the encoder / decoder module 1030 may be implemented as a separate element of system 1000, or it may be incorporated into processor 1010 as a combination of hardware and software as known to those skilled in the art.
[0023] Program code to be loaded onto processor 1010 or encoder / decoder 1030 to execute one or more of the aspects described in this document may be stored in storage device 1040 and subsequently loaded onto memory 1020 for execution by processor 1010. According to various embodiments, one or more of processor 1010, memory 1020, storage device 1040, and encoder / decoder module 1030 may store one or more entries of various entries during the execution of the processes described in this document. Such stored entries may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from processing equations, formulas, operations, and operational logic.
[0024] In some embodiments, the internal memory of processor 1010 and / or encoder / decoder module 1030 is used to store instructions and provide working memory for processing during encoding or decoding. However, in other embodiments, external memory of the processing device (e.g., processor 1010 or encoder / decoder module 1030) is used for one or more of these functions. External memory may be memory 1020 and / or storage device 1040, such as volatile memory and / or non-volatile flash memory. In several embodiments, external non-volatile flash memory is used to store, for example, the operating system of a television. In at least one embodiment, fast external volatile memory (such as RAM) is used as working memory for video encoding and decoding operations, such as for MPEG-2 (MPEG stands for Moving Picture Experts Group; MPEG-2 is also known as ISO / IEC 13818, and 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC stands for High Efficiency Video Coding, also known as H.265), or VVC (Various Video Coding, also known as H.266).
[0025] Inputs to the components of system 1000 can be provided through various input devices as indicated in block 1130. Such input devices include, but are not limited to: a radio frequency (RF) section that receives, for example, RF signals transmitted over the air by a broadcaster; component (COMP) input terminals (or a set of COMP input terminals); universal serial bus (USB) input terminals; and / or high-definition multimedia interface (HDMI) input terminals. Other examples include composite video.
[0026] In various embodiments, the input device of block 1130 has associated corresponding input processing elements as known in the art. For example, the RF section may be associated with elements suitable for: selecting a desired frequency (also referred to as selecting a signal, or limiting a signal band to a band), downconverting the selected signal, further band-limiting it to a narrower band to select (e.g.,) a signal band that may be referred to as a channel in some embodiments), demodulating the downconverted and band-limited signal, performing error correction, and demultiplexing to select a desired data packet stream. The RF section in various embodiments includes one or more elements for performing these functions, such as frequency selectors, signal selectors, band limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF section may include tuners that perform various functions among these functions, including, for example, downconverting a received signal to a lower frequency (e.g., intermediate frequency or near-baseband frequency) or downconverting it to baseband. In one set-top box embodiment, the RF section and its associated input processing elements receive RF signals transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and re-filtering to a desired frequency band. Various embodiments rearrange the order of the components described above (and others), remove some of these components, and / or add other components that perform similar or different functions. Adding components may include inserting components between existing components, such as, for example, inserting amplifiers and analog-to-digital converters. In various embodiments, the RF section includes an antenna. The RF section may conform to standards specifications, such as those published by Digital Video Broadcasting (DVB), the Advanced Television Systems Committee (ATSC), the Radio Industry and Commerce Association (ARIB), or other organizations.
[0027] Additionally, the USB and / or HDMI terminals may include corresponding interface processors for connecting the system 1000 to other electronic devices across USB and / or HDMI connections. It should be understood that various aspects of input processing (e.g., Reed-Solomon error correction) may be implemented as needed, for example, within a separate input processing IC or within the processor 1010. Similarly, various aspects of USB or HDMI interface processing may be implemented as needed, either within a separate interface IC or within the processor 1010. Demodulation, error correction, and demultiplexing streams are provided to various processing elements, including, for example, the processor 1010 and the encoder / decoder 1030, which operate in conjunction with memory and storage elements to process the data streams as needed for presentation on the output device.
[0028] Various components of system 1000 can be provided within an integrated housing, in which various components can be interconnected and transmit data therebetween using a suitable connection arrangement 1140 (e.g., internal buses as known in the art, including inter-IC (I2C) buses, wiring and printed circuit boards).
[0029] System 1000 includes a communication interface 1050 that enables communication with other devices via a communication channel 1060. The communication interface 1050 may include, but is not limited to, a transceiver configured to transmit and receive data via the communication channel 1060. The communication interface 1050 may include, but is not limited to, a modem or network interface card (NIC), and the communication channel 1060 may be implemented, for example, within a wired and / or wireless medium.
[0030] In various embodiments, a wireless network, such as a Wi-Fi network (e.g., IEEE 802.11, where IEEE stands for Institute of Electrical and Electronics Engineers), is used to stream or otherwise provide data to system 1000. In these embodiments, Wi-Fi signals are received via a communication channel 1060 and a communication interface 1050 suitable for Wi-Fi communication. The communication channel 1060 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other over-the-top communications. Other embodiments use a set-top box to provide streaming data to system 1000, delivering data via an HDMI connection to input block 1130. Still other embodiments use an RF connection to input block 1130 to provide streaming data to system 1000. As indicated above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks.
[0031] System 1000 can provide output signals to various output devices, including display 1100, speaker 1110, and other peripheral devices 1120. Display 1100 in various embodiments includes one or more of, for example, a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. Display 1100 can be used in televisions, tablet computers, laptop computers, cellular phones (mobile phones), or other devices. Display 1100 can also be integrated with other components (e.g., as in a smartphone) or separate (e.g., an external monitor for a laptop computer). In various examples of embodiments, other peripheral devices 1120 include one or more of a stand-alone digital video disc (or digital universal disc) (DVR, for both terms), a disk player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 1120 that provide functionality based on the output of system 1000. For example, a disk player performs the function of playing the output of system 1000.
[0032] In various embodiments, control signals are transmitted between system 1000 and display 1100, speaker 1110, or other peripheral devices 1120 using signaling such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols that enable device-to-device control with or without user intervention. Output devices can be communicatively coupled to system 1000 via dedicated connections through corresponding interfaces 1070, 1080, and 1090. Alternatively, output devices can be connected to system 1000 via communication interface 1050 using communication channel 1060. In electronic devices such as, for example, televisions, display 1100 and speaker 1110 can be integrated into a single unit with other components of system 1000. In various embodiments, display interface 1070 includes a display driver, such as, for example, a timing controller chip.
[0033] For example, if the RF portion of input 1130 is part of a separate set-top box, then display 1100 and speaker 1110 can alternatively be separated from one or more other components. In various embodiments where display 1100 and speaker 1110 are external components, output signals can be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs.
[0034] The embodiments may be executed by computer software implemented by processor 1010, or by hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments may be implemented by one or more integrated circuits. As a non-limiting example, memory 1020 may be of any type suitable for the technical environment and may be implemented using any suitable data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. As a non-limiting example, processor 1010 may be of any type suitable for the technical environment and may encompass one or more of microprocessors, general-purpose computers, special-purpose computers, and processors based on multi-core architectures.
[0035] Figure 2 An example of a context in which the following embodiments can be implemented is described. In this context 200, system 210 transmits a video stream to system 230 using communication channel 220. Examples of system 210 include a camera, storage device, computer, drone, video surveillance camera, server, or any device capable of delivering a video stream. The video stream is encoded and transmitted by system 210, or received and / or stored by system 210 and then transmitted. Communication channel 220 is a wired (e.g., Internet, Ethernet, wired network) or wireless (e.g., WiFi, 3G, 4G or 5G, satellite TV, terrestrial TV) network link. System 230 receives and decodes the video stream to generate a sequence of decoded images. An example of system 230 is a set-top box. The obtained sequence of decoded images is then transmitted to display system 250 using communication channel 240, which can be a wired or wireless network as described above. Display system 250 then displays the images. An example of display system 250 is a television or display monitor.
[0036] In this embodiment, system 230 and display system 250 are included in a single device, thus combining the reception, decoding, and display of the video stream. Examples of such devices are televisions, computers, tablet computers, smartphones, head-mounted displays, in-vehicle entertainment systems, and medical devices.
[0037] Figure 3A block diagram of an example video encoder is shown. Variations of this encoder 300 are considered, but for clarity, encoder 300 is described below without describing all anticipated variations. Before being encoded, the video sequence may undergo pre-coding processing 301, such as applying a color transformation to the input color image (e.g., from RGB 4:4:4 to YCbCr 4:2:0), or performing remapping on the input image components to obtain a more compression-resistant signal distribution (e.g., using histogram equalization of one of the color components). For standards that include such mechanisms, metadata may be associated with the pre-processing and attached to the bitstream, for example, in the form of Supplemental Enhancement Information (SEI) messages.
[0038] In the encoder, the image is encoded by encoder elements, as described below. The image to be encoded is segmented (302), for example, as... Figure 5 The process is further described in detail and is performed on a unit basis, such as a coding unit. For example, each unit is encoded using an intra-frame or inter-frame mode. When a unit is encoded in intra-frame mode, it performs intra-frame prediction (360). In inter-frame mode, motion estimation (375) and compensation (370) are performed. The encoder decides (305) which of the intra-frame or inter-frame modes to use for encoding the unit and indicates the intra-frame / inter-frame decision by, for example, a prediction mode flag. For example, the prediction residual is calculated by subtracting (320) the prediction block from the original image block. The prediction residual is then transformed (325) and quantized (330). The quantized transform coefficients, along with the motion vector and other syntax elements, are entropy encoded (345) to output a bitstream. The encoder may skip the transform and apply quantization directly to the untransformed residual signal. The encoder may bypass both the transform and quantization, i.e., encode the residual directly without applying the transform or quantization process.
[0039] The encoder decodes the coded block to provide a reference for further prediction. The quantized transform coefficients are dequantized (340) and inverse transformed (350) to decode the prediction residual. The decoded prediction residual and prediction block are combined (355) to reconstruct the image block. A loop filter (365) is applied to the reconstructed image to perform, for example, deblocking / SAO (Sample Adaptive Shift) and Adaptive Loop Filter (ALF) filtering to reduce coding artifacts. The filtered image is stored in a reference image buffer (380) for further use.
[0040] Encoders typically also perform video decoding as part of the encoding of video data.
[0041] Figure 4A block diagram illustrating an example video decoder is shown. In decoder 400, the bitstream is decoded by decoder elements as described below. Video decoder 400 typically performs a decoding process that is the reverse of the encoding process described in the previous figures. The input to the decoder includes a video bitstream, which can be generated by... Figure 3 The video encoder 300 generates the image. First, the bitstream is entropy decoded (430) to obtain transform coefficients, motion vectors, and other encoded information. Image segmentation information indicates how to segment the image. Therefore, the decoder can segment (435) the image based on the decoded image segmentation information. The transform coefficients are dequantized (440) and inverse transformed (450) to decode the prediction residual. The decoded prediction residual and prediction block are combined (455) to reconstruct the image block. The prediction block (470) can be obtained from intra-frame prediction (460) or motion-compensated prediction (i.e., inter-frame prediction) (475). A loop filter (465) is applied to the reconstructed image. The filtered image is stored at the reference image buffer (480).
[0042] The decoded image can undergo further post-decoding processing (485), such as inverse color transformation (e.g., from YCbCr4:2:0 to RGB 4:4:4) or pre-encoding processing ( Figure 3 The inverse of the remapping process performed in (301). The post-decoding process can use the metadata derived in the precoding process and signaled in the bitstream.
[0043] Figure 5 The illustration shows an example of segmentation performed on images from an original video sequence. The original video sequence 500 includes multiple images 510. Each image comprises multiple pixels, typically arranged in a grid consisting of rows and columns. In this document, a pixel is considered to consist of three components: a luminance component and two chrominance components. However, it is possible to include other types of pixels with fewer or more components, such as including only a luminance component or adding a depth component or adding a transparency component.
[0044] The image is divided into multiple coded entities. First, as indicated by reference numeral 530, the image is divided into a grid of blocks called Coding Tree Units (CTUs). A CTU consists of a block of luminance samples and two corresponding blocks of chrominance samples. The size of such a block is typically N×N, and N is usually a power of two, for example, having a maximum value of "128". Second, the image is divided into groups of one or more CTUs. For example, it can be divided into one or more tile rows and tile columns, where a tile is a sequence of CTUs covering a rectangular area of the image. In some cases, a tile can be divided into one or more blocks, each of which consists of at least one row of CTUs within the block. Above the concepts of tiles and blocks, there exists another coded entity called a slice, which can contain at least one tile of the image or at least one block of a tile. In the example indicated by reference numeral 520, image 21 is divided into three slices S1, S2, and S3 in a raster scan slicing pattern, each slice comprising multiple tiles (not shown), and each tile comprising only one block.
[0045] As indicated by reference numeral 540, a CTU can be partitioned into a hierarchical tree of one or more sub-blocks called coding units (CUs). The CTU is the root (i.e., parent node) of the hierarchical tree and can be partitioned into multiple CUs (i.e., child nodes). If each CU is not further partitioned into smaller CUs, it becomes a leaf of the hierarchical tree; otherwise, it becomes a parent node (i.e., child node) of a smaller CU. During image encoding, the partitioning is adaptive, with each CTU partitioned to optimize compression efficiency.
[0046] For example, CTU 540 is first divided into four square CUs using a quadtree-type partitioning method. The top-left CU 541 is a leaf node in the hierarchical tree because it is not further partitioned; that is, it is not the parent node of any other CU. Using the quadtree-type partitioning method again, the top-right CU is further divided into four smaller square CUs 551, 552, 553, and 554. The bottom-left CU is vertically divided into three rectangular CUs 561, 562, and 563 using a ternary tree-type partitioning method. The bottom-right CU is vertically divided into two rectangular CUs 571 and 572 using a binary tree-type partitioning method.
[0047] In HEVC, the concepts of Prediction Unit (PU) and Transform Unit (TU) emerge. In fact, in HEVC, the coding entities used for prediction (i.e., PU) and transformation (i.e., TU) can be subdivisions of a CU. For example, as shown in the figure, a CU of size 2N×2N can be divided into PUs of size N×2N or PUs of size 580. Furthermore, the CU can be divided into four TUs of size N×N 590 or TUs of size (N / 2)×(N / 2) "16". Other video coding standards also use these concepts. In VVC, except in some special cases, the boundaries of TUs and PUs are aligned with the boundaries of CUs. Therefore, a CU typically includes one TU and one PU.
[0048] In this application, the terms "reconstruction" and "decoding" are used interchangeably, as are the terms "pixel" and "sample," and the terms "image," "picture," "subpicture," "slice," and "frame." Generally, but not necessarily, the term "reconstruction" is used on the encoder side, while "decoding" is used on the decoder side. In this application, the term "block" or "picture block" can refer to any one of CTU, CU, PU, and TU. Furthermore, "block" or "picture block" can refer to macroblocks, partitions, and subblocks as specified in H.264 / AVC or other video coding standards, and more generally, it can also refer to arrays of samples of various sizes.
[0049] SEI messages are syntactic structures defined in various MPEG standards to allow the carrying of metadata. They are specific types of NAL (Network Access Layer) units, which are the basic blocks in the MPEG bitstream format. SEI syntax may vary slightly across different standards, but it typically includes at least the payload type, payload length, and the payload itself. As an example, Table 1 illustrates the SEI syntax defined for VVC. Table 1.
[0050] In addition, specific syntax structures are typically defined for each payload type and instantiated by the sei_payload syntax structure for each payload type.
[0051] The embodiments described below were designed with the foregoing in mind.
[0052] At least one embodiment proposes defining specific SEI messages to carry text data, hereinafter referred to as text SEI messages. Such text SEI messages are inserted into the video bitstream, such that text data related to the video content is carried within the video content itself. Text SEI messages can be used for a variety of applications. One example involves the video production context and includes adding comments or annotations to portions of a video sequence. In such an example, the text data may include production information such as shooting date, scene number, camera operator's name, or technical data such as frame rate or resolution. Another example involves indexing purposes, where text data can be used to add descriptive annotations or keywords. This can improve the use of content in large video sequences where different parts of the sequence may be of particular interest. Other examples of using text data include user-generated comments or descriptions, such as in user-generated content, comments or links related to products displayed in a video sequence, copyright information, etc.
[0053] Another example involves a player for a video sequence that lists available comments and provides the relevant portion of the video to play. For instance, when a comment indicates "my preferred sequence" or other wording that the sequence is preferred, the video sequence including that comment can be played directly.
[0054] The display of this text data can be handled by a device that plays back the video conversion. When a video sequence including a text SEI message is played back by a player device, the text data can be displayed in a specific area of the screen. In multi-window display applications such as computer devices, the text data can be displayed in a window separate from the window displaying the video content.
[0055] Such implementations can be applied in the context of AVC, HEVC, and VVC / VSEI or other video codecs. In at least one embodiment, a new SEI message type is created to embed text data. A new payloadType identifier value is defined for text SEI messages. A specific payloadType value can be selected to identify a text SEI message. Values between 60 and 100 can be used, such as 60 or 61. The sei_payload syntax structure is extended to accommodate this new value and the payload is interpreted as the newly defined “text_comment” syntax.
[0056] In at least one embodiment, the syntax for “text_comment” is defined as illustrated in Table 2. Table 2.
[0057] The semantics of the table elements are as follows: A text_comment_cancel_flag value of 1 indicates that the SEI message cancels the persistence of any previous text comment SEI messages associated with one or more layers to which the text comment SEI message is applied. A text_comment_cancel_flag value of 0 indicates that the text comment message continues.
[0058] `text_comment_id` indicates the identifier of the text comment. The value of `text_comment_id` should be in the range of 0 to 255 (inclusive). In at least one embodiment, a specific value is reserved for a specific type of information. For example, the value "0" may be reserved for copyright information. In another embodiment, a range of values is reserved for a specific type of information. For example, values "0" to "9" may be reserved for a specific type of information, "0" may be reserved for copyright information, "1" may be reserved for parameters of the content, "2" may be reserved for the author's name, etc.
[0059] A `text_comment_id_cancel_flag` value of 1 cancels the duration of the `text_comment_id`-th text comment. A `text_comment_id_cancel_flag` value of 0 indicates that the text comment should continue.
[0060] `text_comment_id_persistence_flag` specifies the persistence of the `text_comment_id`-th text comment in the current layer: `text_comment_id_persistence_flag` equal to 0 indicates that the text comment applies only to the currently decoded picture. `text_comment_id_persistence_flag` equal to 1 indicates that the text comment applies to the currently decoded picture and persists for all subsequent pictures in the current layer in output order until one or more of the following conditions are true: a new CLVS begins in the current layer, or the bitstream ends, or pictures in the current layer associated with a text comment SEI message with the same `text_comment_id` value are output in output order following the current picture.
[0061] The text_comment_alignment_zero_bit should be equal to 0. This element is padded for byte alignment purposes.
[0062] text_comment[text_comment_id][i] contains the i-th byte of the text comment.
[0063] In at least one embodiment, the persistence time is specified based on the number of POCs rather than a flag, especially for consistency with other SEI messages in the AVC context. In that case, text_comment_id_persistence_flag will be replaced by text_comment_id_repetition_period: a value of 0 indicates no persistence, a value of 1 indicates persistence until cancellation (as with the flag), and a value greater than 1 will indicate the number of POCs for which the message is valid during its duration.
[0064] In at least one embodiment illustrated in the syntax of Table 3, reserving the identifier value of 0 allows all previous text SEI messages to be cancelled. In other words, when text_comment_id is zero, all previous text SEI messages are cancelled regardless of their text_comment_id values. Table 3.
[0065] In a variant embodiment, persistence is signaled as a list of identifiers to be cancelled. In another variant embodiment, retained identifier values indicate the cancellation of all identifiers.
[0066] In cases where text comments conflict, such as due to persistence, the comment with the smallest text_comment_id should be considered. The importance of a text comment is in reverse order of its text_comment_id.
[0067] In at least one embodiment illustrated in Table 4, additional grammatical elements related to language information are added to the text comment SEI. This language information allows specifying the language used to signal the comment. Several applications are foreseeable, such as, for example, automatic translation, or selecting comments with appropriate language fields from multiple versions of the same comment. Table 4.
[0068] In the example grammar in Table 4, language information is enabled by the `text_comment_language_flag`. When this flag is equal to 1, it specifies that `text_comment_language` is signaled and a `text_comment_language` grammar element is inserted. In at least one embodiment, the `text_comment_language_flag` grammar element is not signaled and is inferred to be 1. In other words, the `text_comment_language` grammar element is always present in the grammar.
[0069] The `text_comment_language` syntax element specifies the language of the upcoming text comment. In at least one embodiment, this field signals the message in plaintext (e.g., a byte string). In at least one embodiment, this field uses a two-letter ISO 639-1 standard format. In at least one embodiment, this field uses a three-letter ISO 639-2 standard format.
[0070] If text_comment_language does not exist, or if it is a special value such as an empty string or ID 0, it is inferred to be an "undefined value".
[0071] In at least one embodiment, multiple languages can be supported by text messages by using the same text_comment_id for multiple text messages with different text_comment_language values. In other words, text comment SEIs referring to the same image unit (i.e., for the same image) with the same text_comment_id must have different text_comment_language values. The receiver device (or display device) selects the appropriate text comment from the multiple text messages to display. This selection can be made using a selection in the user interface or configuration parameters of the device.
[0072] In at least one embodiment, the text_comment_ids do not need to be strictly identical, but they should have the same modulus value. For example, the language can be specified using the n=8 least significant bits of the text_comment_id (using 8 bits, giving 255 possible values including "undefined" values), while the remaining most significant bits identify the message. Thus, text messages with the same most significant bits refer to the same message that supports different languages. Note that in this embodiment, the language can be directly identified by the n least significant bits of the text_comment_id, and the text_comment_language value is not signaled as before, but rather inferred from the least significant bits of the text_comment_id.
[0073] For example, a message SEI with text_comment_id 25 and 112 refers to the same message with id 0 in the language identified by 25 and 112; a message SEI with text_comment_id 281 (=256+25) and 368 (=256+112) refers to another message with ID 1, where the language is also identified by 25 and 112.
[0074] In at least one embodiment illustrated in Table 5, additional grammatical elements related to the purpose information are added to the text comment SEI. This linguistic information allows the purpose (i.e., meaning) of the text comment to be specified. This can affect how the text comment is processed and how it is presented to the user. For example, if the purpose of the text comment is copyright, it may not be displayed in the same way that the purpose is indexed. The receiver device (or display device) selects the appropriate presentation of the text comment to display based on the purpose. Table 5.
[0075] In the example grammar in Table 5, language information is enabled by the `text_comment_purpose_flag`. When this flag is equal to 1, it specifies that the purpose of the text comment is signaled and the `text_comment_purpose` grammar element is inserted. In at least one embodiment, the `text_comment_purpose_flag` grammar element is not signaled and is inferred to be 1. In other words, the `text_comment_purpose` grammar element is always present in the grammar.
[0076] The syntax element `text_comment_purpose_id` specifies the purpose of the upcoming text comment. If `text_comment_purpose_id` does not exist, or if it is a special value such as purpose ID 0, it is inferred as an "undefined value". Table 6 lists some examples of `text_comment_purpose_id` and corresponding purposes. ID value explain 0 Undefined purpose General text messages 1 copyright Copyright information can define what is permitted on the video. 2 Source (producer, library, everything about acquisition, post-processing, and archiving information; this information is not suitable for EXIF / TIFF data) Information about the video's source can be used in a video library. 3 summary Video summary 4 Change tracking For editing videos, the history of edits already made or new editing suggestions. 5 Explanatory information To understand the video, you may need additional information, such as information about the teaching content. 6 index Keywords are plain text that search engines can use to index videos. 7 Meeting (participants, speakers, screen sharing source, date, etc.) Supplemental information for the recorded meeting, providing information about the conditions for a live meeting. 8 Related links, information, and products (for advertising purposes) Additional metadata on a specific image or shot, which consists of plain text, such as name, biography of a movie or news figure, product URL, etc. 9 Chapter / Narrative Focus The chapter information, shot names, or events in the video can be used for navigation within the video. 10..15 reserve For future use Table 6.
[0077] In at least one embodiment, multiple purposes can coexist and be combined and signaled. In this case, the semantic modification is as follows. The `text_comment_purpose_flag` syntax element is changed to `text_comment_purpose_num`, specifying the number of purposes to be signaled. If it is 0, it acts as if `text_comment_purpose_flag` is false. Furthermore, the `text_comment_purpose_id` single syntax element is changed to, for example, a table identified as `text_comment_purpose_ids[i]`, and for `i` from 0 to `text_comment_purpose_num-1`, includes the purposes to be signaled.
[0078] In at least one embodiment illustrated in Table 7, both additional grammatical elements related to language information and additional grammatical elements related to purpose information are added to the text comment SEI. The corresponding grammatical elements are the same as those described in Tables 4, 5, and 6. All variations described in the context of these tables also apply to this embodiment. Table 7.
[0079] Figure 6 An example of overlapping text SEI messages according to at least one embodiment is illustrated. A text SEI message may include some information related to the persistence of given text data. This allows control over the validity time range. Persistence can be controlled for identifiers, thereby allowing time overlap between messages with different identifiers. The validity time range of a text SEI message is defined by a start point and a stop point. The start point is indicated by a flag indicating that the text data is persistent (in other words, it is valid for the current frame and future frames), and the stop point is indicated by a cancellation flag indicating where the information is no longer valid, or the time when the video sequence ends, or the time when a new SEI message of the same type is received.
[0080] Typically, all messages of the same SEI payload type are canceled by the cancellation flag: this is the case with `text_comment_cancel_flag` here. However, to allow for time overlap, in embodiments that allow for finer control over each text comment ID, the persistence flag is ID-specific, and an ID-specific cancellation flag (`text_comment_id_cancel_flag`) can be used to cancel the validity of one or more previous text SEI messages with that specific ID, while text SEI messages with other IDs can remain valid.
[0081] This allows for temporal overlap of text SEI messages, such as Figure 6 As illustrated in the diagram: Because persistence (which begins with a persistence flag and ends with a cancel flag) is specified for an id, text comments with different ids can coexist independently. For example, when element "Text B" is canceled, element "Text A," which started before "Text B," remains active. Canceling element "Text B" does not cancel all active elements.
[0082] Figure 7An example structure of a video bitstream including a text SEI message is illustrated according to at least one embodiment. In this bitstream example 700, the text SEI is typically inserted before the video data of the picture units (704, 706) (703). It can also be inserted before, after, or in the middle of other parameter sets or other prefix information (such as SPS (701) and PPS (702)) (705).
[0083] Figure 8A The illustration depicts an example process for encoding a bitstream including a text SEI message carrying text data, according to an embodiment. Process 810 is, for example, performed by... Figure 1 Device 1000 processor 1010, Figure 2 System 210 or Figure 3 The encoder 300 is used for implementation. In step 812, the processor obtains video data, text data, and additional information. The video data is encoded in a bitstream, for example, and the text data is to be inserted into the video data. In step 814, the processor packages the text data and additional data into a text SEI message, for example, based on the syntax of Tables 1 to 7. As described in Table 2, the SEI message also includes an identifier (text_comment_id) associated with the text data. In step 816, the processor inserts the SEI message into a bitstream that includes the encoded video data. In step 818, the processor provides a bitstream that includes the encoded video data and text data packaged into the SEI message. The additional data may be related to the purpose of the language and / or text data and may be used accordingly to display the text data.
[0084] Figure 8B The illustration depicts an example process for decoding a bitstream including a text SEI message carrying text data, according to an embodiment. Process 850 is, for example, performed by... Figure 1 Equipment 1000 Figure 2 System 230 or Figure 4 The decoder 400 is used to implement this. In step 852, the processor obtains a bitstream including encoded video data, a text SEI message, and additional data, for example, based on the syntax of Tables 1 and 2. In step 854, the processor extracts the text data and additional data included in the SEI message. In step 856, the processor provides the text data extracted from the bitstream. As described in Table 2, the SEI message also includes an identifier (text_comment_id) associated with the text data. This identifier is used, for example, to handle the persistence of the text data, as referenced above. Figure 6 The additional data may be relevant to the purpose of the language and / or text data, and may be used accordingly to display the text data.
[0085] This application describes various aspects, including tools, features, embodiments, models, methods, etc. Many of these aspects are described in detail, and are generally described in a manner that may sound limiting, at least for the purpose of illustrating the various characteristics. However, this is for the purpose of clarity and does not limit the application or scope of those aspects. In fact, all the different aspects can be combined and interchanged to provide further aspects. Furthermore, the aspects described can also be combined and interchanged with aspects described in earlier filings.
[0086] The aspects described and considered in this application can be implemented in many different forms. The accompanying drawings provide some embodiments, but other embodiments are also contemplated, and the discussion of these drawings does not limit the breadth of implementations. At least one aspect generally relates to video encoding and decoding, and at least one other aspect generally relates to the transmission of a generated or encoded bitstream. These and other aspects can be implemented as methods, apparatus, computer-readable storage media having instructions stored thereon for encoding or decoding video data according to any of the methods, and / or computer-readable storage media having a bitstream generated according to any of the methods stored thereon.
[0087] Various methods are described herein, and each of these methods includes one or more steps or actions for implementing the method. Unless a specific order of steps or actions is required for proper operation of the method, the order and / or use of specific steps and / or actions may be modified or combined.
[0088] The various methods and other aspects described in this application can be used to modify the module, for example, as follows: Figure 3 and Figure 4 The motion compensation and motion estimation modules of the video encoder 300 and decoder 400 shown are illustrated.
[0089] Various numerical values are used in this application, for example, 128 for block size. Specific values are for illustrative purposes, and the aspects described are not limited to these specific values.
[0090] Various implementations involve decoding. As used herein, “decoding” can encompass all or part of a process, such as performing a received encoded sequence to produce a final output suitable for display. In various embodiments, such a process includes one or more processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such a process also includes, or alternatively includes, processes performed by the decoder of the various implementations described herein, such as an adaptive illumination compensation process.
[0091] As a further example, in one embodiment, "decoding" refers only to entropy decoding; in another embodiment, "decoding" refers only to differential decoding; and in yet another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" is intended to specifically refer to a subset of operations or generally to a broader decoding process will be clear based on the specific context of the description and is considered well understood by those skilled in the art.
[0092] Various implementations involve encoding. In a manner similar to the discussion above regarding “decoding,” the term “encoding,” as used herein, can encompass all or part of a process performed on an input video sequence to produce an encoded bitstream. In various embodiments, such a process includes one or more processes typically performed by an encoder, such as segmentation, differential coding, transform, quantization, and entropy coding. In various embodiments, such a process also includes, or alternatively includes, processes performed by the encoders of the various implementations described herein.
[0093] As a further example, in one embodiment, “encoding” refers only to entropy encoding; in another embodiment, “encoding” refers only to differential encoding; and in yet another embodiment, “encoding” refers to a combination of differential encoding and entropy encoding. Whether the phrase “encoding process” is intended to specifically refer to a subset of operations or generally to a broader encoding process will be clear based on the specific context of the description and is considered well understood by those skilled in the art.
[0094] Note that the grammatical elements used in this article are descriptive terms. Therefore, they do not preclude the use of other grammatical element names.
[0095] When the accompanying drawings are presented as flowcharts, it should be understood that block diagrams of the corresponding devices are also provided. Similarly, when the accompanying drawings are presented as block diagrams, it should be understood that flowcharts of the corresponding methods / processes are also provided.
[0096] Various embodiments involve rate-distortion optimization. Specifically, during the encoding process, a balance or trade-off between rate and distortion is typically considered, often taking into account computational complexity constraints. Rate-distortion optimization is generally expressed as minimizing a rate-distortion function, which is a weighted sum of rate and distortion. Different approaches exist to address the rate-distortion optimization problem. For example, such methods can be based on an extended test of all encoding options, including all considered modes or encoding parameter values, with a complete evaluation of their encoding costs and the associated distortion of the reconstructed signal after encoding and decoding. Faster methods can also be used to avoid (save) encoding complexity, particularly the computation of approximate distortion based on prediction or prediction of the residual signal rather than the reconstructed residual signal. A hybrid of these two approaches can also be used, such as by using approximate distortion for only some of the possible encoding options and full distortion for others. Other methods evaluate only a subset of the possible encoding options. More generally, many methods employ any of a variety of techniques to perform optimization, but optimization is not necessarily a complete evaluation of both encoding costs and associated distortion.
[0097] This application describes various aspects, including tools, features, embodiments, models, methods, etc. Many of these aspects are described in detail, and are generally described in a manner that may sound limiting, at least for the purpose of illustrating the various characteristics. However, this is for the purpose of clarity and does not limit the application or scope of those aspects. In fact, all the different aspects can be combined and interchanged to provide further aspects. Furthermore, the aspects described can also be combined and interchanged with aspects described in earlier filings.
[0098] The implementations and aspects described herein can be implemented, for example, in a method or process, apparatus, software program, data stream, or signal. Even if discussed only in the context of a single form of implementation (e.g., discussed only as a method), the implementation of the discussed features can also be implemented in other forms (e.g., apparatus or program). An apparatus can be implemented, for example, in suitable hardware, software, and firmware. A method can be implemented, for example, in a processor, which generally refers to a processing device, including, for example, a computer, microprocessor, integrated circuit, or programmable logic device. Processors also include communication devices, such as, for example, computers, tablet computers, smartphones, cellular phones, portable / personal digital assistants, and other devices that facilitate the transfer of information between end users.
[0099] References to “an embodiment” or “an embodiment” or “an implementation” or “implementation”, and other variations thereof, mean that the specific features, structures, characteristics, etc., described in connection with the embodiment are included in at least one embodiment. Therefore, the phrases “in an embodiment” or “in an embodiment” or “in an implementation” or “in an implementation” appearing throughout this application, and any other variations thereof, do not necessarily refer to the same embodiment.
[0100] Additionally, this application may relate to "determining" fragments of various information. Determining information may include one or more of, for example, estimation information, calculation information, prediction information, or information retrieved from memory.
[0101] Furthermore, this application may relate to “accessing” fragments of various information. Accessing information may include one or more of the following: receiving information, retrieving information (e.g., retrieving information from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.
[0102] Additionally, this application may relate to "receiving" fragments of various information. As with "access," receiving is intended to be a broad term. Receiving information may include one or more of, for example, accessing information or retrieving information (e.g., retrieving information from memory). Furthermore, "receiving" is generally referred to in one or more ways during operations such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0103] In this application, the terms "reconstruction" and "decoding" are used interchangeably, as are the terms "pixel" and "sample," and the terms "image," "picture," "frame," "slice," and "tile." Generally, but not necessarily, the term "reconstruction" is used on the encoder side, while "decoding" is used on the decoder side.
[0104] To be understood, for example, in the cases of “A / B,” “A and / or B,” and “at least one of A and B,” the use of any of the following “ / ,” “and / or,” and “at least one of…” is intended to cover selecting only the first listed option (A), or only the second listed option (B), or both options (A and B). As a further example, in the cases of “A, B, and / or C” and “at least one of A, B, and C,” such wording is intended to cover selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or all three options (A, B, and C). As will be clear to those skilled in the art and related fields, this can be extended to as many entries as possible listed.
[0105] Furthermore, among other things, as used herein, the term "signaling" also refers to instructing the corresponding decoder to do something. For example, in some embodiments, the encoder signals a specific one of the illumination compensation parameters. Thus, in embodiments, the same parameter is used on both the encoder and decoder sides. Therefore, for example, the encoder can transmit (explicitly signal) the specific parameter to the decoder so that the decoder can use the same specific parameter. Conversely, if the decoder already has the specific parameter as well as other parameters, signaling can be used without transmission (implicitly signaling) to allow only the decoder to know and select the specific parameter. Bit savings are achieved in various embodiments by avoiding the transmission of any actual functionality. It should be understood that signaling can be implemented in a variety of ways. For example, in various embodiments, information is signaled to the corresponding decoder using one or more syntax elements, flags, etc. Although the verb form of the term "signaling" has been referred to above, the word "signal" can also be used as a noun herein.
[0106] As will be apparent to those skilled in the art, implementations can generate various signals that are formatted to carry, for example, information that can be stored or transmitted. The information may include, for example, instructions for performing a method or data generated by one of the described implementations. For example, the signal may be formatted to carry a bitstream of the described embodiment. Such a signal may be formatted as, for example, electromagnetic waves (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting may include, for example, encoding the data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. It is well known that signals can be transmitted via a variety of different wired or wireless links. The signal may be stored on a processor-readable medium.
[0107] We have described several embodiments. The features of these embodiments may be provided individually or in any combination across various claim classes and types.
Claims
1. A method comprising: - Obtain encoded video and text data; - Package text data into supplementary and enhanced information messages; - Package additional data representing the language and / or purpose of the text data into a supplemental enhancement message; - Insert supplemental enhancement information messages into the bitstream that includes encoded video data; as well as - Provides bitstream.
2. A method comprising: - Obtain a bitstream including encoded video data and supplemental enhancement information messages, the supplemental enhancement information messages including text data and additional data representing the language and / or purpose of the text data; - Extract text data and additional data representing the language or purpose of the text data from supplemental and enhanced information messages; as well as - Provide extracted text data based on additional data representing the language or purpose.
3. The method according to any one of claims 1 or 2, wherein, Identifiers are associated with text data.
4. The method according to any one of claims 1 to 3, wherein, The supplemental enhanced information message further includes information indicating whether the text data is applied to a single image or to subsequent images.
5. The method according to claim 4, wherein, The persistent information is associated with an identifier.
6. The method according to any one of claims 1 to 5, wherein, The supplemental and enhanced information message conforms to ITU-TH.
274.
7. An apparatus comprising one or more processors, said one or more processors being configured to: - Obtain encoded video and text data; - Package text data into supplementary and enhanced information messages; - Package additional data representing the language and / or purpose of the text data into a supplemental enhancement message; - Insert supplemental enhancement information messages into the bitstream that includes encoded video data; and - Provides bitstream.
8. An apparatus comprising one or more processors, said one or more processors being configured to: - Obtain a bitstream including encoded video data and supplemental enhancement information messages, the supplemental enhancement information messages including text data and additional data representing the language and / or purpose of the text data; - Extract text data and additional data representing the language or purpose of the text data from supplementary and enhanced information messages; and - Provide extracted text data based on additional data representing the language or purpose.
9. The apparatus according to any one of claims 7 or 8, wherein, Identifiers are associated with text data.
10. The apparatus according to claims 7 to 9, wherein, The supplemental enhanced information message further includes information indicating whether the text data is applied to a single image or to subsequent images.
11. The apparatus according to claim 10, wherein, The persistent information is associated with an identifier.
12. The apparatus according to any one of claims 7 to 11, wherein, The supplemental and enhanced information message conforms to ITU-TH.
274.
13. A non-transitory computer-readable medium comprising data content generated by the method according to any one of claims 1 to 6.
14. A computer program product comprising instructions, which, when executed by one or more processors, are used to perform the method of any one of claims 1 to 6.