Single-channel encoding to a multi-channel container and subsequent image compression
By converting single-channel data into n-dimensional values and compressing using n-dimensional curves or spaces, the method addresses inefficiencies in existing pipelines, enhancing data capacity utilization and reducing artifacts.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-09-14
- Publication Date
- 2026-03-30
AI Technical Summary
Existing video/image distribution pipelines struggle with efficient encoding and decoding of single-channel data streams, such as depth data, vertex data, and index data, due to the lack of support for single-channel transmission in conventional multi-channel containers, leading to inefficient use of data capacity and the presence of compression artifacts.
A method and apparatus that convert single-channel data into n-dimensional values using a mapper, assigning these values to pixels in a virtual image frame, and compressing the frame according to the container type, utilizing n-dimensional curves or spaces to map scalar values to n-dimensional values based on way tree partitions, allowing compatibility with existing hardware without modifying infrastructure.
This approach enables efficient use of multi-channel container data capacity, preserves spatial and temporal coherence, and suppresses compression artifacts, while maintaining compatibility with existing hardware and infrastructure.
Smart Images

Figure 0007837471000002 
Figure 0007837471000003 
Figure 0007837471000004
Abstract
Description
[Technical Field]
[0001] 1. Cross-reference to related applications This application claims the benefit of priority under U.S. Provisional Application No. 63 / 407,885, filed on 19 September 2022, and European Application No. 23151686.5, filed on 16 January 2023, and these applications in their entirety are incorporated herein by reference.
[0002] 2. Disclosure Areas Various exemplary embodiments generally relate to video / image compression, and more specifically to video / image encoding and decoding, but are not limited thereto. [Background technology]
[0003] 3.Background Compression reduces the memory capacity required to store videos and images, as well as the bandwidth needed to transmit them. The Motion Picture Experts Group (MPEG) codec, based on the H.264 compression standard, is a prime example of a codec used for this purpose. Other codecs are also available on the market. Many codecs are designed to be compatible with a variety of professional, consumer, and mobile phone cameras. [Overview of the Initiative] [Problems that the invention aims to solve]
[0004] Outline of a specific embodiment Disclosed herein are various methods and apparatus for packing single-channel data into multi-channel containers, such as MP4, TIFF, or JPEG containers, in order to achieve good utilization of the container's data capacity. In various embodiments, packing is performed using a mapping relationship represented by an n-dimensional curve, or multiple 2-dimensional spaces in an n-dimensional space. n way(2 n 2 nThe process is executed based on mapping relationships represented by street-tree (e.g., an octave tree of n=3) partitions. At least some embodiments are compatible with existing hardware in legacy video / image distribution pipelines; that is, they do not inherently depend on hardware / infrastructure modifications. Rather, such embodiments can be advantageously implemented as software and / or firmware, for example, by interfacing a corresponding add-on data processing module with an existing codec without changing the codec's native container format(s). [Means for solving the problem]
[0005] According to an exemplary embodiment, an encoding method is provided which includes: using a processor to convert a plurality of scalar values of a received data stream into corresponding plurality of n-dimensional values, the conversion being performed using a mapper; using the processor to assign each of the n-dimensional values as a pixel value to each pixel of a virtual image frame, where n is an integer greater than 1; and using the processor to compress the virtual image frame according to the type of image data container, wherein the mapper is a plurality of n-dimensional curves or n-dimensional spaces. n Based on the relationships represented by the way tree partition, scalar values are configured to map to corresponding n-dimensional values.
[0006] According to another exemplary embodiment, a non-temporary computer-readable medium is provided which, when executed by an electronic processor, stores instructions causing the electronic processor to perform an operation including the method described above.
[0007] In yet another exemplary embodiment, an apparatus for encoding image data is provided, comprising at least one processor and at least one memory containing program code, wherein the at least one memory and the program code are configured to cause the apparatus to perform at least the following using the at least one processor: convert a plurality of scalar values of a received data stream into corresponding plurality of n-dimensional values using an electronic mapper; assign each of the n-dimensional values as a pixel value to each pixel of a virtual image frame, where n is an integer greater than 1; and compress the virtual image frame according to the type of container for the image data, wherein the electronic mapper is a plurality of n-dimensional curves or n-dimensional spaces. n Based on the relationships represented by the way tree partition, scalar values are configured to map to corresponding n-dimensional values. [Brief explanation of the drawing]
[0008] Other aspects, features, and advantages of various disclosed embodiments will become more fully apparent, for example, from the following detailed description and accompanying drawings:
[0009] Figure 1 shows an example of a video / image distribution pipeline;
[0010] Figure 2 is a flowchart of the encoding method usable in the video / image distribution pipeline of Figure 1 according to the embodiment;
[0011] Figure 3 illustrates the operating principle of the 1D→3D mapper usable with the encoding method shown in Figure 2 according to the embodiment;
[0012] Figure 4 illustrates the operating principle of a 1D to 3D mapper usable with the encoding method shown in Figure 2 according to another embodiment;
[0013] Figures 5A and 5B illustrate the operating principle of a 1D to 3D mapper that can be used with the encoding method shown in Figure 2 according to yet another embodiment;
[0014] FIG. 6 is a flowchart of a decoding method that can be used in the video / image delivery pipeline of FIG. 1 according to an embodiment.
[0015] FIG. 7 is a block diagram showing a computing device according to an embodiment. DETAILED DESCRIPTION
[0016] Detailed Description The present disclosure and aspects thereof can be embodied in various forms including computer-implemented methods, computer program products, computer systems and networks, user interfaces, and hardware, devices or circuits controlled by application programming interfaces, as well as hardware-implemented methods, signal processing circuits, memory arrays, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), etc. The above is only intended to give a general idea of the various aspects of the present disclosure and is not intended to limit the scope of the present disclosure in any way.
[0017] FIG. 1 is a diagram showing a processing example of a video / image delivery pipeline 100 that shows various stages from video / image capture to video / image content display according to an embodiment. A series of video / image frames 102 can be captured or generated using an image generation block 105. The frames 102 may be digitally captured (e.g., by a digital camera) or generated by a computer (e.g., using computer animation) to provide video and / or image data 107. Alternatively, the frames 102 may be captured on film by a silver halide camera. Then, the video / image data 107 may be provided by scanning the film and converting it into a digital format.
[0018] In production stage 110, data 107 may be edited to provide a video / image production stream 112. The data of the video / image production stream 112 may be provided to a processor (e.g., one or more processors such as a central processing unit, CPU, etc.) for post-production editing in post-production block 115. Post-production editing in block 115 may include, for example, adjusting or modifying the color and brightness of specific areas of an image to improve image quality or to achieve the visual appearance of the video according to the filmmaker's intentions. This part of post-production editing is sometimes called "color timing" or "color grading". Other edits (e.g., scene selection and sequencing, image cropping, addition of computer-generated visual special effects, artifact removal, etc.) may be performed in block 115 to bring about a "final" version 117 of the production for distribution. During post-production editing 115, the video and / or images may be optimized for viewing on a reference display 125.
[0019] Following post-production 115, the data of the final version 117 may be delivered to an encoding block 120 for further delivery to decoding devices and playback devices such as downstream television sets, set-top boxes, movie theaters, etc. In some embodiments, the encoding block 120 may include audio encoders and video encoders such as those defined by ATSC, DVB, DVD, Blu-Ray, and other distribution formats to generate an encoded bitstream 122. In a receiver, the encoded bitstream 122 is decoded by a decoding unit 130 to generate a corresponding decoded signal 132 that represents a copy or an exact approximation of the signal 117. The receiver may be attached to a target display 140 that may have somewhat or completely different characteristics from the reference display 125. In such a case, the decoded signal 132 can be mapped to the characteristics of the target display 140 by generating a display mapping signal 137 using a display management (DM) block 135. Depending on the embodiment, the decoding unit 130 and the display management block 135 may include separate processors or may be based on a single integrated processing unit. The encoding block 120 and / or the decoding unit 130 can be implemented using various embodiments disclosed below.
[0020] The codec used in the encoding block 120 and / or the decoding unit 130 enables video / image data processing and compression / decompression. Compression is used in the encoding block 120 to make the corresponding file(s) smaller. The decoding process performed by the decoding unit 130 typically includes decompressing the received video / image data file(s) into a form usable for playback and / or further editing. Examples of codecs that can be used in the encoding block 120 and the decoding unit 130 include, but are not limited to, the XviD / DivX codec, the MPEG codec, and the H.264 codec.
[0021] A container is a digital file that packages video / audio data and corresponding metadata into a single file. Metadata may include subtitles, resolution information, creation date, device type, language information, etc. A container file interleaves different data types in a way that allows the decoding unit 130 to easily access its components. Different types of containers are typically identified by their respective file extensions, including MP4, WAV, AIFF, AVI, MOV, WMV, MKV, TIFF, JPEG, HEVC, FLV, F4V, SWF, and others.
[0022] For example, the JPEG image file format is a common choice for storing and transmitting photographic images, such as still images and individual video frames. Many operating systems have viewers that support the visualization of JPEG image files saved with the JPG or JPEG extension. Many web browsers also support the visualization of JPEG image files. JPEG encoding typically involves the following operations: (1) Conversion: A color image is converted from the RGB (Red, Green, Blue) color space to the luminance / chrominance space. (2) Downsampling: Downsampling is typically performed on the chrominance component but not the luminance component. For example, an image frame can be downsampled at a ratio of 2:1 horizontally and 1:1 vertically (2h1v). (3) Grouping: Pixels for each color component are often grouped into 8x8 pixel groups called "data units". If the number of rows is not an integer multiple of 8, the bottom row is duplicated one or more times to make it the required multiple. A similar duplication may be applied to the rightmost column. (4) Discrete Cosine Transformation (DCT): The DCT is applied to each data unit to create an 8x8 map of the transformed components. The DCT typically involves information loss due to the limitations of the precision of machine-based calculations. (5) Quantization: Each of the 64 transformed components in the data unit is divided by another number called the Quantization Coefficient (QC) and rounded to an integer. This operation generally results in further information loss. Larger QC values tend to result in greater loss. Many encoders rely on the QC table recommended by the JPEG standard for quantization. (6) Encoding: The 64 quantized transformation coefficients (which are integers) of each data unit are encoded using a combination of run-length encoding (RLE) and Huffman coding. (7) Header: The final operation adds a header listing all relevant JPEG parameters. The corresponding JPEG decoder uses inverse operations to generate an image that closely resembles the original encoded image.
[0023] As another example, a TIFF image frame consists of a rectangular grid of pixels. The two axes of this geometry are called horizontal (or X or width) and vertical (or Y or length). The horizontal and vertical resolutions do not need to be equal. A baseline TIFF image divides the vertical range of the image into one or more strips, which are encoded and compressed individually (separately). The TIFF format is an alternative to tiled image formats, where both the horizontal and vertical ranges of the image are divided into smaller units. The data for one pixel contains one or more samples. For example, an RGB image typically has one red sample, one green sample, and one blue sample per pixel, while a grayscale image has only one sample per pixel. The TIFF format can be used with both additive color models (e.g., RGB) and subtractive color models (e.g., cyan, magenta, yellow, black, i.e., CMYK). In at least some examples, the interpretation of channel data is performed outside the TIFF container. Interpretation can be assisted by metadata, such as an International Color Consortium (ICC) profile. The TIFF format does not restrict the number of samples per pixel, nor does it restrict how many bits are encoded for each sample. For example, 3 samples per pixel is the lower limit for multispectral imaging supported by TIFF, while hyperspectral imaging (which TIFF also supports) may use 100 or more samples per pixel. The support for custom selection of the number of samples per pixel in the TIFF format is utilized in at least some embodiments disclosed below herein. TIFF images may be uncompressed, compressed using lossless compression, or compressed using lossy compression. An example of a lossless compression method compatible with the TIFF format is LZW (Lempel-Ziv-Welch) compression.
[0024] Augmented reality (AR) applications, virtual reality (VR) applications, and various applications involving the rendering of three-dimensional (3D) scenes typically generate additional image data streams, each of which may take the form of a corresponding data sequence transmitted via a corresponding dedicated data channel. Examples of such data streams include, but are not limited to, depth data, vertex data, and index data. However, many of the containers (file formats) of the types described above, and the corresponding conventional hardware for the encoding block 120 and decoding unit 130, inherently do not support single-channel transmission, which unfortunately makes efficient encoding / decoding and compression / decompression of the additional data streams described above difficult.
[0025] The various embodiments disclosed herein address at least some of the problems pointed out above in the art by providing various methods for packing single-channel data into a multi-channel container in order to achieve one or more of the following: (i) efficient use of the data capacity of a multi-channel container, (ii) preservation of the spatial and / or temporal coherence of a single-channel data stream in its packed form, and (iii) suppression or avoidance of excessive compression artifacts in the expanded single-channel data in the decoder unit 130. At least some embodiments are fully compatible with existing hardware of the corresponding video / image distribution pipeline (100, Figure 1, etc.), i.e., they do not inherently depend on modifications to that hardware / infrastructure. Rather, such embodiments can be advantageously implemented in software and / or firmware by, for example, interfacing a corresponding relatively small data processing module with the existing codec without changing the native container format(s) of the codec.
[0026] Figure 2 is a flowchart of the encoding method 200 available in the encoding block 120 according to the embodiment. The encoding method 200 enables encoding a single-channel data stream (e.g., a data sequence) into an n-channel container, where n is an integer greater than 1. The single-channel data stream can be, for example, a depth data stream corresponding to a 3D scene. The n-channel container can be, for example, one of the multi-channel containers described above.
[0027] The encoding method 200 includes receiving the next value of a single-channel data stream (in block 202). In the first instance of block 202, the next value received is the first value of the stream. In subsequent instances of block 202, the next value received is the value of the stream following the previously received value.
[0028] Furthermore, the encoding method 200 includes selecting the next pixel of the virtual image frame (in block 204). Here, the term “virtual image frame” refers to an image frame similar to a conventional image frame. However, unlike the latter, the virtual image frame does not represent a conventional image. Rather, different pixels of the virtual image frame can be assigned any desired pixel value, for example, a value generated by a suitable mapper (see, for example, block 206 and Figures 3-5). The pixels of such a virtual image frame can be spatially arranged, for example, in a rectangular array having the pixels in rows and columns, similar to a conventional image frame.
[0029] In some examples, the top-left corner pixel of the frame is selected in the first instance of block 204. In any subsequent instance of block 204, the selection process may follow a raster pattern or any other suitable predetermined pattern. The processing of the entire virtual image frame in encoding method 200 is typically completed after each pixel of the frame has been selected once. For example, in the case of a raster pattern, the pixel selected in the last instance of block 204 may be the bottom-right corner pixel of the frame.
[0030] In other embodiments, pixel processing orders other than the sequential order described above are also used. In one embodiment, random pixel selection is performed. In various additional embodiments, pixels of the virtual image frame are processed in parallel within single-instruction, multiple data (SIMD) vectorization; multiple-instruction, multiple data (MIMD) multithreading; or a combination thereof. Such alternatives can also be implemented in various embodiments of the decoding method 600 (see Figure 6).
[0031] The encoding method 200 also includes converting the one-dimensional (one-dimensional, scalar) values received in block 202 to corresponding n-dimensional (nD) values (in block 206) using a selected 1D→nD mapper. In some specific examples, the nD values can be represented by vectors in nD space. In Cartesian coordinates, an origin-based vector v in nD space is a sequence of values (x1, x2, ..., x n ) is expressed by, where x in is the length of the projection of vector v onto the coordinate axis corresponding to the i-th dimension in nD space. The number n is an algorithm parameter that depends on the container used in encoding method 200. For example, for JPEG and MP4 containers, n=3. For TIFF containers, n is the number of samples per pixel, and as described above, n can be > 3 in at least some examples. Several non-restrictive examples of 1D→nD mappers that can be used in block 206 are described in detail below with reference to Figures 3-5.
[0032] The encoding method 200 also includes assigning the nD values generated in block 206 to the pixels selected in block 204 (in block 208). For example, for a container supporting an RGB color scheme with n=3, the x1, x2, and x3 components of the corresponding 3D vector are assigned to the selected pixels as the R, G, and B values of those pixels, respectively (in block 208). As another example, for a TIFF container with n=16, the x1, x2, ..., x of the corresponding 16D vector are assigned. 16 The components are assigned to pixels as 16 samples, each corresponding to a multispectral image (in block 208). Other suitable assignment schemes can be used in other examples in block 208.
[0033] The encoding method 200 also includes determining (in the decision block 210) whether or not the end of the corresponding virtual image frame has been reached. If it is determined that the end of the frame has not been reached ("No" in the decision block 210), the operation of the encoding method 200 loops back to block 202. Otherwise ("Yes" in the decision block 210), the virtual image frame (with all pixels allocated) is compressed in the conventional way according to the container format (in block 212). The container can then be directed from the encoding block 120 to the decoding unit 130, as described earlier (see also Figure 1). Once the operation of block 212 is complete, the encoding method 200 terminates.
[0034] Figures 3 to 5 show some non-limiting examples of 1D→nD mappers that can be used in block 206 of encoding method 200 according to various embodiments. For the sake of clarity and not as a limitation, the illustrated examples correspond to the number n=3. Based on the provided description, those skilled in the art will be able to construct and use various 1D→nD mappers corresponding to other values of n, e.g., n=2 and n>3, without excessive experimentation. In at least some examples, the corresponding 1D→nD mappers are implemented using one or more lookup tables (LUTs).
[0035] Figure 3 illustrates the operating principle of a 1D to 3D mapper that can be used in block 206 of the encoding method 200 according to one embodiment. More specifically, Figure 3 shows a three-dimensional logarithmic spiral 302 in a Cartesian coordinate system where the coordinate axes are labeled X1, X2, and X3, respectively. This spiral 302 can be used to map a scalar value d in the range [0, D] to a 3D vector (Y, Cb, Cr), where Y, Cb, and Cr represent luminance, blue difference chroma component, and red difference chroma component, respectively.
[0036] In some embodiments, a scalar value d (D≧d≧0) is mapped onto the helix 302 by finding a point on the helix at a distance d from the helix's origin O along the helix. Then, the Cartesian coordinates (x3, x2, x1) of the found point are used to determine the corresponding values of Y, Cb, and Cr, respectively. In some specific examples, in block 206 of encoding method 200, the following equations (1) to (3) are used to program the processor of encoding block 120 to perform the corresponding operations on the fly: JPEG0007837471000001.jpg24147 Here, a and b are parameters that determine how the radius of the spiral 302 increases, Cr0 and Cb0 are (offset) constants, and F(d) is a function of d that determines the distance scale along the Y (luminance) axis. In some specific examples, the function F(d) is a quadratic function of d. In some other specific examples, equations (1) to (3) are used to pre-calculate the corresponding LUT, which is then accessed by the processor to perform the corresponding operation in block 206.
[0037] If the mapping performed in block 206 is implemented based on the spiral 302, such a mapping is spatially and temporally coherent. In particular, such a mapping is distance-preserving. Furthermore, the mapping is unique for any two points located on the spiral 302. Moreover, for any three points located on the spiral 302, the distance relationship in 3D space is the same as the distance relationship along the spiral 302 (i.e., in 1D space). For example, if point B is located between points A and C on the spiral 302, the mapping is such that in 3D space, the distance between points A and C is greater than the distance between points A and B, and greater than the distance between points B and C. These properties, arising from the spatial and temporal coherence of the mapping, are typically useful for achieving efficient compression in formats such as JPEG, HEVC, MP4, and many others, and for suppressing the undesirable manifestation of compression artifacts in the decompressed single-channel data calculated by the decoder unit 130.
[0038] Figure 4 illustrates the operating principle of a 1D to 3D mapper that can be used in block 206 of encoding method 200 according to another embodiment. More specifically, Figure 4 shows a 3D Peano curve 402 in a Cartesian coordinate system with axes labeled X1, X2, and X3, respectively. The granularity of the Peano curve 402 is 2 bits per dimension. In additional embodiments, Peano curves of other granularities can be used similarly. In other additional embodiments, various nD Peano curves can be used to implement corresponding 1D to nD mappers of various granularities (where n=2 or n>3).
[0039] The Peano curve 402 is used to map scalar values d in the range [0, D] to 64 different 3D vectors (Y, Cb, Cr) or (R, G, B). In some examples, a scalar value d (D ≥ d ≥ 0) is mapped to the curve 402 by finding a point on the curve where the distance along the curve from the curve's origin O rounds to d. Then, using the Cartesian coordinates (x3, x2, x1) of the found point, the corresponding values for Y, Cb, and Cr, or the corresponding values for R, G, and B, are determined for the pixel in question.
[0040] Peano curves, such as Peano curve 402, are examples of space-filling curves (with endpoints) whose range covers the entire range of the corresponding n-dimensional hypercube. Other examples of space-filling curves include, but are not limited to, Hilbert curves and Morton curves. Space-filling curves are a special case of fractal curves. Various 1D→nD mappers suitable for implementing block 206 of coding method 200 can be constructed using various suitable space-filling curves and / or fractal curves (with endpoints) in the same manner as described above with reference to Peano curve 402 and Figure 4.
[0041] Figures 5A to 5B schematically illustrate the operating principle of a 1D→3D mapper that can be used in block 206 of encoding method 200 according to yet another embodiment. More specifically, FIG. 5A shows an octree-partitioned 3D cube 500. FIG. 5B shows an octree 510 used for partitioning the 3D cube 500.
[0042] In general, an octree, such as octree 510 in FIG. 5B, is a tree data structure in which each internal node has exactly eight children. An octree can be used to partition a bounded three-dimensional space (such as 3D cube 500) by recursively subdividing the bounded space into eight octants. In a point-region (PR) octree, a node stores an explicit three-dimensional point that is the “center” of the subdivision of that node. This center point defines one of the corners of each of the eight children. In a matrix-based (MX) octree, the subdivision point implicitly becomes the center of the space represented by that node. The root node of a PR octree can represent infinite space. The root node of an MX octree represents a finite bounded space, and the implicit center is clearly defined. For example, the root node R of octree 510 (see FIG. 5B) represents 3D cube 500. Octree 510 is an example of a 2 n way tree with n = 3. Those skilled in the art in the relevant field will be able to easily understand how to construct various 2 n way trees for other values of n without performing excessive experiments. In additional examples, the value of n is 2, 4, 5, etc.
[0043] The eight children of the root node R of octree 510 are nodes 0 through 7 (see Figure 5B). The subcubes (octants) of 3D cube 500 corresponding to child nodes 0 through 7 are similarly labeled 0 through 7 in Figure 5A. Subcube 7 is not directly visible in the diagram shown in Figure 5A. For illustrative purposes, only the children of child node 1 are explicitly shown in Figure 5B. In Figure 5B, these grandchild nodes are labeled 10 through 17. The subcubes (octants) of subcube 1 corresponding to grandchild nodes 10 through 17 are similarly labeled 10 through 17 in Figure 5A. Subcube 17 is not directly visible in the diagram shown in Figure 5A. For illustrative purposes, only the children of grandchild node 11 are explicitly shown in Figure 5B. These great-grandchild nodes are labeled 110 through 117 in Figure 5B. Similarly, the subcubes (octants) of subcube 11 corresponding to great-grandchild nodes 110-117 are also labeled 110-117 in Figure 5A. Subcube 117 is not directly visible in the diagram shown in Figure 5A. A person skilled in the art will readily understand that each of child nodes 0 and 2-7 similarly has grandchild nodes that also have great-grandchild nodes (not explicitly shown in Figure 5A). Furthermore, a person skilled in the art will also understand that each of grandchild nodes 10 and 12-17 similarly has great-grandchild nodes (not explicitly shown in Figure 5A). There are a total of 256 great-grandchild nodes in the octree 510. There are also a total of 256 corresponding great-grandchild subcubes in the 3D cube 500. The 256 great-grandchild subcubes partition the 3D cube 500 into 256 non-overlapping parts, each having the shape of a cube.
[0044] In some specific examples, to implement a 1D to 3D mapper for block 206 of encoding method 200, the three dimensions X1, X2, X of a 3D cube 500 (3)These are assigned to represent R, G, B pixel values or Y, Cb, Cr pixel values, respectively. The range [0,D] of the scalar d value is divided into 256 intervals. Each interval is assigned to one of 256 great-grandchild subcubes of the 3D cube 500. The R, G, B (or Y, Cb, Cr) pixel values of the mapped scalar d value are then determined using the Cartesian coordinates (x3, x2, x1) of the center of the corresponding great-grandchild subcube. In some specific examples, the relationship described above between the scalar d value and the coordinates (x3, x2, x1) of the center of the subcube is pre-calculated, compiled into a corresponding LUT, and then accessed by the encoder to perform the appropriate calculation in block 206.
[0045] In various embodiments, the encoding granularity of the octree, i.e., the number of subcubes in the 3D cube 500, is selected based on the desired accuracy and compression ratio intended for the encoder. The finer the subdivision of the cube, the more likely the lossy compression performed in block 212 of the encoding method 200 is to result in errors during the corresponding decoding performed in the decoding unit 130 (see also Figure 6). Therefore, the selection of the encoding granularity of the octree may also need to be based on the amount of errors introduced by the lossy compression.
[0046] Figure 6 is a flowchart of the decoding method 600 usable in the decoding unit 130 according to the embodiment. The decoding method 600 and the encoding method 200 are compatible with each other. Thus, the decoding method 600 enables the substantial recovery of a single-channel data stream encoded for transmission using a selected multi-channel container format and the encoding method 200.
[0047] The decoding method 600 includes decompressing the virtual image frame, which has been compressed according to the container format, (in block 602). The decompression performed in block 602 is the inverse operation of the compression performed in block 212 of the encoding method 200.
[0048] Furthermore, the decoding method 600 includes selecting the next pixel in the decompressed virtual image frame and reading the nD pixel value of the selected pixel (in block 604). In a typical example, the selection operation in block 604 of the decoding method 600 is implemented in the same way as the selection operation in block 204 of the encoding method 200 described above. Therefore, readers should refer to the above description in block 204 for the relevant details of the pixel selection process.
[0049] The decoding method 600 also includes converting the nD pixel values read in block 604 to corresponding scalar values using a suitably selected nD→1D demapper (in block 606). The selected nD→1D demapper is one in which the demapping performed by it is the reverse of the mapping performed by the 1D→nD mapper used in block 206 of the encoding method 200. In at least some examples, the nD→1D demapper used in block 606 of the decoding method 600 and the 1D→nD mapper used in block 206 of the encoding method 200 are implemented based on the same LUT. More specifically, the nD→1D demapper used in block 606 is configured to refer to a scalar value in its LUT based on a given nD value, while the 1D→nD mapper used in block 206 of the encoding method 200 is configured to refer to an nD value in the same LUT based on a given scalar value. In various implementations, the nD→1D demapper operates to account for errors introduced by lossy compression. For example, if the virtual frame pixel RGB values (100, 200, 30) decode as (94, 203, 31) due to errors introduced by lossy compression, the nD→1D demapper operates to correct the error to return the original (100, 200, 30) RGB values. Such error correction is achieved, for example, by selecting the granularity of the 1D→nD mapping such that the typical scattering of decoded points around the original constellation points is within the range closest to the original constellation points (e.g., in terms of Euclidean distance) rather than other constellation points. This feature of the demapper is often called maximum likelihood detection.
[0050] The decoding method 600 also includes outputting the scalar value determined in block 606 as the next value in the corresponding data stream (data sequence) (in block 608). If there is no loss and / or significant compression artifacts, the latter data stream is a copy or approximate copy of the data stream received in block 202 of the encoding method 200.
[0051] The decoding method 600 also includes determining (in the determination block 610) whether or not the end of the corresponding virtual image frame has been reached. If it is determined that the end of the frame has not been reached ("No" in the determination block 610), the operation of the decoding method 600 loops back to block 604. Otherwise ("Yes" in the determination block 610), the decoding method 600 terminates.
[0052] Figure 7 is a block diagram showing a computing device 700 according to an embodiment. The device 700 can be used, for example, in the coding block 120. A computing device similar to the device 700 can also be used in the decoding unit 130. Based on the following description of the device 700, those skilled in the art will readily understand how to manufacture, configure, and use a similar computing device for the decoding unit 130.
[0053] Device 700 comprises an input / output (I / O) device 710, an encoding engine 720, and a memory 730. The I / O device 710 can be used to enable device 700 to receive at least a portion of the video / image stream 117 and output at least a portion of the encoded bitstream 122. The memory 730 may have, for example, a buffer for receiving encoded and compressed image data via the video / image stream 117. The received image data may, in particular, include the single-channel data stream described above. Once the data is buffered, the memory 730 can provide a portion of the data to the encoding engine 720 for processing there. The encoding engine 720 includes a processor 722 and a memory 724. The memory 724 may store program code therein, which, when executed by the processor 722, enables the encoding engine 720 to perform a variety of encoding operations, including, but not limited to, the various encoding operations described above with reference to some or all of Figures 2 to 5. The memory 724 may also store the LUTs described above, which can be accessed by the processor 722 as needed.
[0054] Various aspects of the present invention can be further understood from the following enumerated example embodiments (EEE).
[0055] EEE(1): An encoding method comprising: using a processor to convert a plurality of scalar values of a received data stream into corresponding plurality of n-dimensional values, wherein the conversion is performed using a mapper; using the processor to assign each of the n-dimensional values as a pixel value to each pixel of a virtual image frame, where n is an integer greater than 1; and using the processor to compress the virtual image frame according to the type of image data container, wherein the mapper is a plurality of n-dimensional curves or n-dimensional spaces. nThe way tree is configured to map scalar values to corresponding n-dimensional values based on the relationships represented by the tree partitions. Here, an n-dimensional straight line is not an example of the "n-dimensional curve" described above. For example, a "curve" contains at least two parts that are not collinear (colinear) with each other in the corresponding n-dimensional space.
[0056] EEE(2): n is greater than 3, as described in EEE(1).
[0057] EEE(3): The method according to EEE(1) or EEE(2), wherein the corresponding n-dimensional values are a set including red, green, and blue values, or a set including cyan, magenta, and yellow values, or a set including luminance values, blue difference chroma values, and red difference chroma values.
[0058] EEE(4): The method according to any one of EEE(1) to EEE(3), wherein the plurality of scalar values are depth data, vertex data, or index data representing a 3D scene.
[0059] EEE(5): The method according to EEE(1), EEE(3), or EEE(4), wherein the n-dimensional curve is a helix and n=2 or n=3.
[0060] EEE(6): The method according to any one of EEE(1) to (5), wherein the n-dimensional curve is a space-filling curve having two endpoints.
[0061] EEE(7): The method according to EEE(6), wherein the space-filling curve is selected from the group consisting of Peano curves, Hilbert curves, Morton curves, and fractal curves.
[0062] EEE(8): The mapper is configured to determine the corresponding n-dimensional value by any one of EEE(1) to EEE(6): finding a position on the n-dimensional curve that represents the scalar value, having a distance measured along the n-dimensional curve from its endpoints; and representing each component of the corresponding n-dimensional value in terms of a set of coordinates of the position in the n-dimensional space.
[0063] EEE(9): The mapper is the plurality of 2 n A method according to any one of EEE(1) to EEE(5) and EEE(8), wherein the method is configured to identify one partition representing the scalar value from among the way tree partitions, and to determine the corresponding n-dimensional value by representing each component of the corresponding n-dimensional value with the coordinate set of the one partition in the n-dimensional space.
[0064] EEE(10): The mapper is the n-dimensional curve or the n-dimensional space of the multiple 2 n A method according to any one of EEE(1) to EEE(9), configured to use a lookup table pre-calculated based on a way tree partition.
[0065] EEE(11): The method according to any one of EEE(1) to EEE(10), further comprising: generating an uncompressed image frame by decompressing a compressed image frame using the processor or another processor, wherein the decompression is performed according to the type of container and the compressed image frame is generated by compression; and converting a plurality of n-dimensional pixel values of the uncompressed image frame to a plurality of other scalar values using the processor or another processor, wherein the conversion is performed using a demapper, wherein the demapper is configured to perform a demapping operation that is the inverse of the corresponding mapping operation of the mapper.
[0066] EEE(12): Both the mapper and the demapper are multiple n-dimensional curves or n-dimensional spaces.n The method described in EEE(11), which is configured to use the same lookup table pre-calculated based on way tree partitioning.
[0067] EEE(13): A non-temporary computer-readable medium that stores instructions that cause an electronic processor to perform an operation including one of EEE(1) to EEE(12) when executed by the electronic processor.
[0068] EEE(14): A device for encoding image data, comprising at least one processor and at least one memory containing program code, wherein the at least one memory and the program code are configured to cause the device to perform at least the following using the at least one processor: converting a plurality of scalar values of a received data stream into corresponding plurality of n-dimensional values using an electronic mapper; assigning each of the n-dimensional values as a pixel value to each pixel of a virtual image frame, where n is an integer greater than 1; and compressing the virtual image frame according to the type of image data container, wherein the electronic mapper is a plurality of n-dimensional curves or n-dimensional spaces. n Based on the relationships represented by the way tree partition, scalar values are configured to map to corresponding n-dimensional values.
[0069] EEE(15): The apparatus according to EEE(14), wherein the electronic mapper is configured to determine a position on the n-dimensional curve having a certain distance measured along the n-dimensional curve from its endpoints, which represents the scalar value, and to represent each component of the corresponding n-dimensional value in terms of a set of coordinates of the position in the n-dimensional space.
[0070] EEE(16): The electronic mapper is the plurality of 2 nThe apparatus according to EEE(14), which is configured to identify one partition representing the scalar value from among the way tree partitions, and to represent each component of the corresponding n-dimensional value by the coordinate set of the one partition in the n-dimensional space.
[0071] EEE(17): The electronic mapper is the n-dimensional curve or the n-dimensional space of the multiple 2 n A device described in any one of EEE(14) to EEE(16), configured to use a lookup table pre-calculated based on a way tree partition.
[0072] EEE(18): The apparatus according to any one of EEE(14) to EEE(17), wherein the at least one memory and the program code cause the apparatus to further generate an uncompressed image frame by uncompressing a compressed image frame according to the type of container, and to convert a plurality of n-dimensional pixel values of the uncompressed image frame into another plurality of scalar values using an electronic demapper, wherein the electronic demapper is configured to perform a demapping operation that is the inverse of the corresponding mapping operation of the electronic mapper.
[0073] EEE(19): Both the electronic mapper and the electronic demapper are multiple 2 of the n-dimensional curve or the n-dimensional space. n The apparatus described in EEE(18), configured to use a common lookup table (e.g., each copy of the same thing) precalculated based on a way tree partition.
[0074] EEE(20): The apparatus described in any one of EEE(14) to EEE(19), wherein the plurality of scalar values are depth data, vertex data, or index data representing a three-dimensional scene.
[0075] With regard to the processes, systems, methods, heuristics, etc., described herein, the steps of such processes, etc., have been described as occurring in a certain ordered sequence. However, it goes without saying that such processes may be performed in an order other than that described herein. Furthermore, it goes without saying that certain steps may be performed simultaneously, other steps may be added, or certain steps described herein may be omitted. In other words, the descriptions of processes herein are provided for the purpose of illustrating specific embodiments and should not be construed as limiting the scope of the claims.
[0076] Accordingly, it goes without saying that the above description is illustrative and not limiting. Reading the above description will reveal many embodiments and uses beyond the examples provided. The scope of the claims should not be determined by reference to the above description, but rather by reference to the appended claims, in accordance with the full scope of equivalents to which such claims are entitled. It is anticipated and intended that future developments will occur in the technology discussed herein, and that the disclosed systems and methods will be incorporated into such future embodiments. In other words, it goes without saying that this application is subject to modification and alteration.
[0077] All terms used in the claims are intended to be given the broadest reasonable interpretation and the ordinary meaning as understood by those skilled in the art described herein, unless expressly otherwise provided herein. In particular, singular articles such as “a,” “the,” and “said” mean that there is one or more elements described, unless expressly otherwise provided herein.
[0078] The disclosure summary is provided to enable readers to quickly grasp the nature of the technical disclosure. It is submitted with the understanding that it is not intended to be used to interpret or limit the claims or their meaning. Furthermore, as can be seen in the preceding detailed description, various features are grouped together in various embodiments for clarity. This method of disclosure does not reflect the intention that the embodiments described in the claims contain more features than are explicitly described in each claim. Rather, as reflected in the following claims, the inventive subject matter lies in fewer features than all the features of a single disclosed embodiment combined. Therefore, the following claims are incorporated into the detailed description herein, and each claim stands alone as the subject matter claimed individually.
[0079] This disclosure includes references to exemplary embodiments, but this specification is not intended to be constrained. Various modifications of the embodiments described, as well as other embodiments within the scope of this disclosure that are apparent to those skilled in the art to which this disclosure relates, are deemed to be within the principles and scope of this disclosure, as expressed, for example, in the following claims.
[0080] Some embodiments can be implemented as circuit-based processes, including possible implementations on a single integrated circuit.
[0081] Some embodiments can be embodied in the form of methods and apparatus for carrying out those methods. Some embodiments can also be embodied in the form of program code recorded on a tangible medium such as a magnetic recording medium, an optical recording medium, a solid-state memory, a floppy disk, a CD-ROM, a hard drive, or any other non-temporary machine-readable storage medium, and when the program code is loaded into and executed by a machine such as a computer, that machine becomes an apparatus for carrying out the various embodiments described herein. Some embodiments can also be embodied in the form of program code stored on a non-temporary machine-readable storage medium, including being loaded into and / or executed by a machine, and when the program code is loaded into and executed by a machine such as a computer or processor, that machine becomes an apparatus for carrying out the various embodiments described herein. When implemented on a general-purpose processor, the program code segment combines with the processor to provide a unique device that operates similarly to a particular logic circuit.
[0082] Unless otherwise specified, each number and range should be interpreted as an approximation, as if preceded by the words "about" or "approximately."
[0083] The use of figure numbers and / or figure reference labels in the claims is intended to identify one or more possible embodiments of the claimed subject matter in order to facilitate the interpretation of the claims. Such use is not necessarily construed as limiting those claims to the embodiments shown in the corresponding figures.
[0084] Where elements described in the following method claims are listed in a specific order with corresponding designations, these elements are not necessarily intended to be limited to being carried out in that specific order unless the claims imply a specific order for carrying out some or all of these elements.
[0085] Any reference in this specification to “one embodiment” or “a particular embodiment” means that certain features, structures, or characteristics described in relation to that embodiment may be included in at least one embodiment of this disclosure. While the phrase “in one embodiment” appears in various places in this specification, not all of them necessarily refer to the same embodiment, nor are separate or additional embodiments necessarily mutually exclusive with other embodiments. The same applies to the term “implementation.”
[0086] In this specification, unless otherwise specified, the use of sequential adjectives such as “first,” “second,” “third,” etc., to refer to one of several similar objects merely indicates that different examples of such similar objects are being referred to, and does not imply that the similar objects referred to in this manner must be in a corresponding order or sequence, whether temporal, spatial, ranking, or otherwise.
[0087] Unless otherwise specified herein, in addition to its plain meaning, the conjunction "if" can also be interpreted as "when," "upon," "in response to determining," or "in response to detecting," and the interpretation may depend on the corresponding specific context. For example, the phrases "if it is determined" or "if [a stated condition] is detected" can be interpreted as "upon determining" or "in response to determining," or "upon detecting [the stated condition or event]" or "in response to detecting [the stated condition or event]."
[0088] Furthermore, in this specification, the terms “couple,” “coupling,” and “connect” refer to any implementation known or to be developed in the art in which energy is transferred between two or more elements, and the intervention of one or more additional elements is intended, though not required. Conversely, terms such as “directly coupled” and “directly connected” mean the absence of such additional elements.
[0089] Where used herein in relation to elements and standards, the terms "compatible" and "in accordance with" mean that the element communicates with other elements in the manner specified by the standard, either entirely or partially, and that other elements recognize it as being sufficiently capable of communicating with other elements in the manner specified by the standard. A compatible element does not need to operate internally in the manner specified by the standard.
[0090] The functionality of the various elements shown in the diagram, including the functional blocks labeled “Processor” and / or “Controller,” can be provided not only by the use of dedicated hardware but also by the use of software-executable hardware in conjunction with appropriate software. Where provided by a processor, the functionality may be provided by a single dedicated processor, a single shared processor, or multiple individual processors, some of which may be shared. Furthermore, the explicit use of the terms “Processor” or “Controller” should not be interpreted as referring only to, but not limited to, hardware capable of executing software, implicitly including, but not limited to, digital signal processor (DSP) hardware, network processors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), read-only memory (ROM), random-access memory (RAM), and non-volatile storage for storing software. Other conventional and / or custom hardware may also be included. Similarly, the switches shown in the diagram are conceptual only. These functions may be executed through the operation of program logic, through dedicated logic, through the interaction of program control and dedicated logic, or manually, and specific techniques can be selected by the implementer so that they are understood more concretely from the context.
[0091] As used in this application, the term “circuit, circuitry” can mean one or more or all of the following: (a) a hardware-only circuit implementation (such as an implementation of analog and / or digital circuits only); (b) a combination of hardware circuits and software, for example (where applicable): (i) a combination of analog and / or digital hardware circuits (one or more) and software / firmware; (ii) any part of a hardware processor (one or more), software and memory (one or more) having software (including digital signal processors (one or more)) that works together to enable a device such as a mobile phone or server to perform various functions; and (c) a hardware circuit (one or more) and / or processor (one or more), such as a microprocessor (one or more) or a part of a microprocessor (one or more), which requires software (e.g., firmware) to operate, but which may not be present when not required for operation. This definition of circuit applies to all use of the term in this application, including in the claims. As a further example, as used in this application, the term "circuit" also includes, but is not limited to, a hardware circuit or processor (or more processors) or an implementation of a part of a hardware circuit or processor and the software and / or firmware associated with it (or them). Furthermore, as an example and applicable to a particular claim element, the term "circuit" also includes a baseband integrated circuit or processor integrated circuit for a mobile device, or a similar integrated circuit in a server, cellular network device, or other computing or network device.
[0092] Those skilled in the art will understand that any block diagrams and flowcharts in this specification represent conceptual diagrams of exemplary circuits embodying the principles of this disclosure. Similarly, any flowcharts, flow diagrams, state transition diagrams, pseudocode, etc., will be understood to represent various processes that are substantially represented in a computer-readable medium and can be executed by a computer or processor, whether such a computer or processor is explicitly indicated or not.
[0093] The “BRIEF SUMMARY OF SOME SPECIFIC EMBODIMENTS” in this specification are intended to present several exemplary embodiments, and additional embodiments are described with reference to the “DETAILED DESCRIPTION” and / or one or more drawings. The “BRIEF SUMMARY OF SOME SPECIFIC EMBODIMENTS” are not intended to identify essential elements or features of the claimed subject matter, nor are they intended to limit the scope of the claimed subject matter.
Claims
1. An encoding method for encoding a single-channel data stream into an n-channel image container, The process involves using a processor to convert multiple scalar values in a received single-channel data stream into corresponding multiple n-dimensional values, where n is an integer greater than 1 and depends on the format of the n-channel image container used for encoding, and the conversion is performed using a mapper. Using the aforementioned processor, each of the n-dimensional values is assigned as a pixel value to each pixel of the virtual image frame, This includes using the processor to compress the virtual image frame according to an n-channel image container format, The aforementioned mapper is a plurality of n-dimensional curves or n-dimensional spaces. n Based on a mapping relationship represented by a way tree partition, each of the scalar values is configured to map to a corresponding n-dimensional value, wherein for each of the scalar values, the mapping relationship is a position on the n-dimensional curve that represents the scalar value, having a distance measured along the n-dimensional curve from its endpoint or one of the multiple 2 n This includes determining a partition among the way tree partitions and representing each component of the corresponding n-dimensional value in terms of the position or the coordinate set of the partition in the n-dimensional space. method.
2. The method according to claim 1, wherein n is greater than 3.
3. The method according to claim 1 or 2, wherein the corresponding n-dimensional values are a set including red values, green values, and blue values, or a set including cyan values, magenta values, and yellow values, or a set including luminance values, blue difference chroma values, and red difference chroma values.
4. The method according to claim 1 or 2, wherein the plurality of scalar values are depth data, vertex data, or index data representing a three-dimensional scene.
5. The method according to claim 1, wherein the n-dimensional curve is a helix and n=2 or n=3.
6. The method according to claim 1 or 2, wherein the n-dimensional curve is a space-filling curve having two endpoints.
7. The method according to claim 6, wherein the space-filling curve is selected from the group consisting of Peano curves, Hilbert curves, Morton curves, and fractal curves.
8. The mapper is the n-dimensional curve or the plurality of 2 in the n-dimensional space. n The method according to claim 1 or 2, configured to use a lookup table pre-calculated based on a way tree partition.
9. The process involves generating a decompressed image frame by decompressing a compressed image frame using the aforementioned processor or another processor, wherein the decompression is performed according to the container format, and the compressed image frame was generated by compression. The process further includes converting a plurality of n-dimensional pixel values of the decompressed image frame into a plurality of other scalar values using the aforementioned processor or another processor, wherein the conversion is performed using a demapper. Here, the demapper is configured to perform a demapping operation that is the reverse of the corresponding mapping operation of the mapper. The method according to claim 1 or 2.
10. Both the mapper and the demapper are multiple 2 of the n-dimensional curve or the n-dimensional space. n The method according to claim 9, configured to use a common lookup table calculated in advance based on a way tree partition.
11. A non-temporary computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform an operation including the method according to claim 1 or 2.
12. A device for encoding image data, At least one processor, It comprises at least one memory containing program code, The at least one memory and the program code are configured to cause the device to perform at least the following using the at least one processor: The method involves using an electronic mapper to convert multiple scalar values in a received single-channel data stream into corresponding multiple n-dimensional values, wherein n is an integer greater than 1 and depends on the format of the n-channel image container used for encoding. Each of the aforementioned n-dimensional values is assigned as a pixel value to each pixel of the virtual image frame, The virtual image frame is compressed according to an n-channel image container format, Here, the electronic mapper is a plurality of n-dimensional curves or n-dimensional spaces. n Based on a mapping relationship represented by a way tree partition, each of the scalar values is configured to map to a corresponding n-dimensional value, wherein for each of the scalar values, the mapping relationship is a position on the n-dimensional curve having a certain distance measured along the n-dimensional curve from its endpoints that represents the scalar value, or the plurality of two n This includes determining a partition among the way tree partitions and representing each component of the corresponding n-dimensional value in terms of the position or the coordinate set of the partition in the n-dimensional space.
13. The electronic mapper is the n-dimensional curve or the n-dimensional space of the plurality of 2 n The apparatus according to claim 12, configured to use a lookup table calculated in advance based on a way tree partition.
14. The at least one memory and the program code are further processed by the at least one processor in the device. Decompressed image frames are decompressed according to the container format to generate decompressed image frames, Using an electronic demapper, the system converts multiple n-dimensional pixel values of the expanded image frame into multiple other scalar values. The apparatus according to claim 12 or 13, wherein the electronic demapper is configured to perform a demapping operation that is the reverse of the corresponding mapping operation of the electronic mapper.
15. Both the electronic mapper and the electronic demapper are multiple 2 of the n-dimensional curve or the n-dimensional space. n The apparatus according to claim 14, configured to use a common lookup table calculated in advance based on a way tree partition.
16. The apparatus according to claim 12 or 13, wherein the plurality of scalar values are depth data, vertex data, or index data representing a three-dimensional scene.
Citation Information
Patent Citations
Quality-driven streaming
JP2015531186A
High dynamic range video format with low dynamic range compatibility
JP2025536502A
Gray Tracking Across Dynamically Changing Display Characteristics
US20200105179A1
Point cloud compression using a space filling curve for level of detail generation
US20200217937A1
Wrapped reshaping for codeword augmentation with neighborhood consistency
WO2022103902A1