Truncation error signaling and adaptive jitter for lossy bandwidth compression
By implementing truncation error signaling and adaptive dithering in data processing, the method addresses the issue of visual artifacts in compression, improving image reconstruction accuracy and reducing errors.
Patent Information
- Application Number
- CN202380083694.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-14
- Filing Date
- 2023-11-29
- Publication Date
- 2025-07-15
AI Technical Summary
Existing compression techniques cannot effectively alleviate visual artifacts, especially error problems arising from lossy bandwidth compression.
Through truncation error signaling and adaptive jitter technology, the error value set and residual sample set of truncation data are calculated, bit streams are generated and analytical reconstruction is performed to reduce visual artifacts.
Improve the accuracy of image reproduction, reduce the occurrence of visual artifacts, and optimize the performance of lossy bandwidth compression.
Smart Images

Figure CN120322802A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims the benefit of U.S. Non - Provisional Patent Application Serial No. 18 / 066,087, entitled “TRUNCATION ERROR SIGNALING AND ADAPTIVE DITHER FOR LOSSY BANDWIDTH COMPRESSION”, filed on December 14, 2022, which is hereby incorporated by reference in its entirety. Technical Field
[0003] The present disclosure generally relates to processing systems, and more particularly, to one or more techniques for data processing. Background Art
[0004] Computing devices generally perform graphics and / or display processing (e.g., using a Graphics Processing Unit (GPU), Central Processing Unit (CPU), display processor, etc.) to render and display visual content. Such computing devices can include, for example, computer workstations, mobile phones (such as smartphones), embedded systems, personal computers, tablet computers, and video game consoles. The GPU is configured to execute a graphics processing pipeline that includes one or more processing stages that operate together to execute graphics processing commands and output frames. The Central Processing Unit (CPU) can control the operation of the GPU by issuing one or more graphics processing commands to the GPU. Modern CPUs are generally capable of concurrently executing multiple applications, and each of the multiple applications may need to utilize the GPU during execution. A display processor can be configured to convert digital information received from the CPU into analog values and can issue commands to a display panel to display visual content. Devices that provide content for visual presentation on a display can utilize a CPU, GPU, and / or display processor.
[0005] Current compression techniques may not effectively mitigate visual artifacts that arise as a by - product of the compression / decompression process. Improved techniques for compression / decompression of images and other types of data are needed. Summary of the Invention
[0006] A simplified summary of one or more aspects is presented below in order to provide a basic understanding of such aspects. This summary is not an extensive overview of all contemplated aspects, and is neither intended to identify key or critical elements of all aspects nor to describe the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simplified form as a prelude to the more detailed description that is presented later.
[0007] In one aspect of the present disclosure, a method, a computer-readable medium, and an apparatus for data processing are provided. The apparatus includes a memory and at least one processor coupled to the memory, and at least partially based on information stored in the memory, the at least one processor is configured to: perform a truncation process on data, where the data is associated with display processing, image processing, or the data processing, and where the truncation process on the data generates truncated data; calculate a set of truncation error values associated with the truncation process on the truncated data; generate a set of residual samples for the truncated data; and generate a bitstream based on the set of residual samples for the truncated data and the set of truncation error values associated with the truncation process.
[0008] In one aspect of the present disclosure, a method, a computer-readable medium, and an apparatus for data processing are provided. The apparatus includes a memory and at least one processor coupled to the memory, and at least partially based on information stored in the memory, the at least one processor is configured to: obtain a bitstream associated with a set of residual samples for truncated data and a set of truncation error values for a truncation process, where the bitstream corresponds to data associated with display processing, image processing, or the data processing; parse the set of residual samples for the truncated data and the set of truncation error values from the bitstream to obtain a parsed set of residual samples for the truncated data; and reconstruct the truncated data based on the parsed set of residual samples and the set of truncation error values, where the reconstruction of the truncated data generates untruncated data.
[0009] To achieve the foregoing and related purposes, one or more aspects include the features described comprehensively below and particularly pointed out in the claims. The following description and the drawings set forth in detail some illustrative features of one or more aspects. However, these features are only indicative of some of the various ways in which the principles of the various aspects may be employed, and this specification is intended to include all such aspects and their equivalents. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 is a block diagram illustrating an example content generation system in accordance with one or more techniques of the present disclosure.
[0011] Figure 2 illustrates an example GPU in accordance with one or more techniques of the present disclosure.
[0012] Figure 3 illustrates an example image or surface in accordance with one or more techniques of the present disclosure.
[0013] Figure 4 is a diagram illustrating an example of an encoder and a decoder.
[0014] Figure 5A diagram that illustrates an example of a truncated error signaling technique.
[0015] Figure 6 A diagram that illustrates example aspects of adaptive dithering based on a truncated error.
[0016] Figure 7 A call flow diagram that illustrates an example communication between an encoder and a decoder according to one or more techniques of the present disclosure.
[0017] Figure 8 A flowchart of an example method of data processing according to one or more techniques of the present disclosure.
[0018] Figure 9 A flowchart of an example method of data processing according to one or more techniques of the present disclosure.
[0019] Figure 10 A flowchart of an example method of data processing according to one or more techniques of the present disclosure.
[0020] Figure 11 A flowchart of an example method of data processing according to one or more techniques of the present disclosure. Detailed Description
[0021] Various aspects of the system, apparatus, computer program product, and method will be described more fully hereinafter with reference to the accompanying drawings. However, the present disclosure may be embodied in many different forms and should not be construed as limited to any specific structure or function presented throughout the present disclosure. Rather, these aspects are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the present disclosure to those skilled in the art. Based on the teachings herein, those skilled in the art should understand that the scope of the present disclosure is intended to cover any aspect of the systems, apparatuses, computer program products, and methods disclosed herein, whether implemented independently of other aspects of the present disclosure or in combination with other aspects of the present disclosure. For example, any number of the aspects set forth herein may be used to implement an apparatus or practice a method. In addition, the scope of the present disclosure is intended to cover such apparatuses or methods implemented using other structures, functionality, or a combination of structures and functions in addition to or different from the various aspects of the disclosure set forth herein. Any aspect disclosed herein may be embodied by one or more elements of a claim.
[0022] Although various aspects are described herein, many variations and permutations of these aspects fall within the scope of the disclosure. Although some potential benefits and advantages of the aspects of the disclosure are mentioned, the scope of the disclosure is not intended to be limited to particular benefits, uses, or objectives. Instead, the aspects of the disclosure are intended to apply broadly to different wireless technologies, system configurations, processing systems, networks, and transmission protocols, some of which are illustrated by way of example in the figures and the following description. The detailed description and the figures merely illustrate the disclosure and do not limit the disclosure, and the scope of the disclosure is defined by the appended claims and their equivalent technical solutions.
[0023] Several aspects are presented with reference to various devices and methods. These devices and methods are described in the following detailed description and illustrated in the figures by various blocks, components, circuits, processes, algorithms, etc. (collectively referred to as "elements"). These elements can be implemented using electronic hardware, computer software, or any combination thereof. Whether such elements are implemented as hardware or software depends on the particular application and the design constraints imposed on the overall system.
[0024] For example, an element or any part of an element or any combination of elements can be implemented as a "processing system" that includes one or more processors (which may also be referred to as processing units). Examples of processors include microprocessors, microcontrollers, graphics processing units (GPUs), general-purpose GPUs (GPGPUs), central processing units (CPUs), application processors, digital signal processors (DSPs), reduced instruction set computing (RISC) processors, system-on-chips (SOCs), baseband processors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), programmable logic devices (PLDs), state machines, gated logic components, discrete hardware circuits, and other suitable hardware configured to perform the various functions described in the present disclosure. One or more processors in the processing system can execute software. Whether referred to as software, firmware, middleware, microcode, hardware description language, or other names, software can be broadly understood to mean instructions, instruction sets, code, code segments, program code, programs, subroutines, software components, applications, software applications, software packages, routines, subroutines, objects, executable files, execution threads, processes, functions, etc.
[0025] The term "application" can refer to software. As described herein, one or more techniques can refer to an application (e.g., software) configured to perform one or more functions. In such examples, the application can be stored in a memory (e.g., on-chip memory of a processor, system memory, or any other memory). The hardware (such as a processor) described herein can be configured to execute the application. For example, the application can be described as including code that, when executed by the hardware, causes the hardware to perform one or more of the techniques described herein. As an example, the hardware can access the code from the memory and execute the code accessed from the memory to perform one or more of the techniques described herein. In some examples, components are identified in the present disclosure. In such examples, a component can be hardware, software, or a combination thereof. Components can be separate components or sub-components of a single component.
[0026] In one or more examples described herein, the described functionality can be implemented in hardware, software, or any combination thereof. If implemented in software, the functionality can be stored or encoded on a computer-readable medium as one or more instructions or code. Computer-readable media includes computer storage media. Storage media can be any available media that can be accessed by a computer. By way of example and not limitation, such computer-readable media can include random access memory (RAM), read only memory (ROM), electrically erasable programmable ROM (EEPROM), optical disk storage, magnetic disk storage, other magnetic storage devices, combinations of the aforementioned types of computer-readable media, or any other media that can be used to store computer-executable code in the form of instructions or data structures that can be accessed by a computer.
[0027] As used herein, instances of the term "content" can refer to "graphics content", "image", etc., regardless of whether the term is used as an adjective, a noun, or other part of speech. In some examples, as used herein, the term "graphics content" can refer to content generated by one or more processes of a graphics processing pipeline. In additional examples, as used herein, the term "graphics content" can refer to content generated by a processing unit configured to perform graphics processing. In yet additional examples, as used herein, the term "graphics content" can refer to content generated by a graphics processing unit.
[0028] Many devices (e.g., CPUs, GPUs, DPUs, etc.) can perform compression / decompression processes on data (e.g., image data) to save memory bandwidth when the data is written to or read from memory. For example, a GPU can compress a surface (i.e., data) and send the compressed surface to a DPU, and then the DPU can decompress the surface and display the surface on a display. Some compression processes (e.g., lossy compression) may introduce errors (e.g., visual artifacts) into the data when the data is reconstructed (i.e., decompressed).
[0029] This disclosure relates to various techniques related to truncation error signaling and adaptive dithering for lossy bandwidth compression. In one example, a device (i.e., an encoder) performs a truncation process on data, where the data is associated with display processing, image processing, or data processing, and where the truncation process for the data produces truncated data. The device calculates a set of truncation error values (e.g., an average set of truncation error values) associated with the truncation process for the truncated data. The device generates a set of residual samples for the truncated data. The device generates a bitstream based on the set of residual samples for the truncated data and the set of truncation error values associated with the truncation process. In another example, the device (or another device, i.e., a decoder) obtains a bitstream associated with a set of residual samples for truncated data and a set of truncation error values for a truncation process, where the bitstream corresponds to data associated with display processing, image processing, or data processing. The device parses the set of residual samples for the truncated data and the set of truncation error values from the bitstream to obtain a parsed set of residual samples for the truncated data. The device reconstructs the truncated data based on the parsed set of residual samples and the set of truncation error values, where the reconstruction of the truncated data produces untruncated data.
[0030] Compared to including a set of truncation error values in a bitstream, the untruncated data resulting from the reconstruction of truncated data can have less error (e.g., visual artifacts) compared to the untruncated data resulting from a reconstruction process that is not based on the set of truncation error values included in the bitstream. Additionally, by utilizing the set of truncation error values as part of a dithering process, the foregoing techniques can also reduce the occurrence of errors compared to other compression / decompression techniques. Therefore, compared to other compression / decompression techniques, the foregoing techniques can improve the accuracy of image reproduction.
[0031] Figure 1FIG. 0 is a block diagram illustrating an example content generation system 100 configured to implement one or more techniques of the present disclosure. The content generation system 100 includes a device 104. The device 104 may include one or more components or circuits for performing the various functions described herein. In some examples, one or more components of the device 104 may be components of a SOC. The device 104 may include one or more components configured to perform one or more techniques of the present disclosure. In the illustrated example, the device 104 may include a processing unit 120, a content encoder / decoder 122, and a system memory 124. In some aspects, the device 104 may include a plurality of components (e.g., a communication interface 126, a transceiver 132, a receiver 128, a transmitter 130, a display processor 127, and one or more displays 131). The display 131 may refer to one or more displays 131. For example, the display 131 may include a single display or multiple displays, and the multiple displays may include a first display and a second display. The first display may be a left-eye display, and the second display may be a right-eye display. In some examples, the first display and the second display may receive different frames for presentation thereon. In other examples, the first display and the second display may receive the same frames for presentation thereon. In additional examples, the results of graphics processing may not be displayed on the device. For example, the first display and the second display may not receive any frames for presentation thereon. Instead, the frames or the results of graphics processing may be passed to another device. In some aspects, this is referred to as split rendering.
[0032] The processing unit 120 may include an internal memory 121. The processing unit 120 may be configured to perform graphics processing using the graphics processing pipeline 107. The content encoder / decoder 122 may include an internal memory 123. In some examples, the device 104 may include a processor that may be configured to perform one or more display processing techniques on one or more frames generated by the processing unit 120 and then display the frames via one or more displays 131. Although the processor in the example content generation system 100 is configured as the display processor 127, it should be understood that the display processor 127 is an example of a processor, and other types of processors, controllers, etc. may be used in place of the display processor 127. The display processor 127 may be configured to perform display processing. For example, the display processor 127 may be configured to perform one or more display processing techniques on one or more frames generated by the processing unit 120. One or more displays 131 may be configured to display or otherwise present the frames processed by the display processor 127. In some examples, one or more displays 131 may include one or more of the following: a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, a projection display device, an augmented reality display device, a virtual reality display device, a head-mounted display, or any other type of display device.
[0033] Memory external to the processing unit 120 and the content encoder / decoder 122, such as system memory 124, may be accessible to the processing unit 120 and the content encoder / decoder 122. For example, the processing unit 120 and the content encoder / decoder 122 may be configured to read from and / or write to the external memory, such as system memory 124. The processing unit 120 may be communicatively coupled to the system memory 124 via a bus. In some examples, the processing unit 120 and the content encoder / decoder 122 may be communicatively coupled to the internal memory 121 via the bus or via a different connection.
[0034] The content encoder / decoder 122 may be configured to receive graphics content from any source, such as the system memory 124 and / or the communication interface 126. The system memory 124 may be configured to store the received encoded or decoded graphics content. The content encoder / decoder 122 may be configured to receive the encoded or decoded graphics content in the form of encoded pixel data, for example, from the system memory 124 and / or the communication interface 126. The content encoder / decoder 122 may be configured to encode or decode any graphics content.
[0035] The internal memory 121 or the system memory 124 may include one or more volatile or non-volatile memories or storage devices. In some examples, the internal memory 121 or the system memory 124 may include RAM, static random access memory (SRAM), dynamic random access memory (DRAM), erasable programmable ROM (EPROM), EEPROM, flash memory, magnetic data media, or optical storage media, or any other type of memory. According to some examples, the internal memory 121 or the system memory 124 may be a non-transitory storage medium. The term "non-transitory" may indicate that the storage medium is not embodied in a carrier wave or a propagated signal. However, the term "non-transitory" should not be construed to mean that the internal memory 121 or the system memory 124 is immovable or that its contents are static. For example, the system memory 124 may be removed from the device 104 and moved to another device. As another example, the system memory 124 may not be removable from the device 104.
[0036] The processing unit 120 may be a CPU, GPU, GPGPU, or any other processing unit configurable to perform graphics processing. In some examples, the processing unit 120 may be integrated into the motherboard of the device 104. In additional examples, the processing unit 120 may be present on a graphics card installed in a port of the motherboard of the device 104, or may otherwise be incorporated into a peripheral device configured to interoperate with the device 104. The processing unit 120 may include one or more processors, such as one or more microprocessors, GPUs, ASICs, FPGAs, arithmetic logic units (ALUs), DSPs, discrete logic components, software, hardware, firmware, other equivalent integrated or discrete logic circuits, or any combination thereof. If the technology is implemented partially in software, the processing unit 120 may store instructions for the software in a suitable non-transitory computer-readable storage medium (e.g., the internal memory 121), and may execute the instructions in hardware using one or more processors to perform the techniques of the present disclosure. Any of the above (including hardware, software, combinations of hardware and software, etc.) may be considered one or more processors.
[0037] The content encoder / decoder 122 can be any processing unit configured to perform content decoding. In some examples, the content encoder / decoder 122 can be integrated into the motherboard of the device 104. The content encoder / decoder 122 can include one or more processors, such as one or more microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), arithmetic logic units (ALUs), digital signal processors (DSPs), video processors, discrete logic components, software, hardware, firmware, other equivalent integrated or discrete logic circuits, or any combination thereof. If the technology is implemented partially in software, the content encoder / decoder 122 can store instructions for the software in a suitable non-transitory computer-readable storage medium (e.g., internal memory 123), and can execute the instructions in hardware using one or more processors to perform the techniques of this disclosure. Any of the foregoing (including hardware, software, combinations of hardware and software, etc.) can be considered one or more processors.
[0038] In some aspects, the content generation system 100 can include a communication interface 126. The communication interface 126 can include a receiver 128 and a transmitter 130. The receiver 128 can be configured to perform any of the receiving functions described herein with respect to the device 104. Additionally, the receiver 128 can be configured to receive information from another device, such as eye or head positioning information, rendering commands, and / or location information. The transmitter 130 can be configured to perform any of the sending functions described herein with respect to the device 104. For example, the transmitter 130 can be configured to send information to another device, which can include a request for content. The receiver 128 and the transmitter 130 can be combined into a transceiver 132. In such examples, the transceiver 132 can be configured to perform any of the receiving functions and / or sending functions described herein with respect to the device 104.
[0039] Refer again to Figure 1, in some aspects, the processing unit 120 and / or the display processor 127 may include a compressor / decompressor 198 configured to: perform a truncation process on data associated with display processing, image processing, or data processing, wherein the truncation process on the data produces truncated data; calculate a set of truncation error values associated with the truncation process on the truncated data; generate a set of residual samples for the truncated data; and generate a bitstream based on the set of residual samples for the truncated data and the set of truncation error values associated with the truncation process. In some aspects, the compressor / decompressor 198 is configured to: obtain a bitstream associated with a set of residual samples for truncated data and a set of truncation error values for a truncation process, wherein the bitstream corresponds to data associated with display processing, image processing, or data processing; parse the set of residual samples for the truncated data and the set of truncation error values from the bitstream to obtain a parsed set of residual samples for the truncated data; and reconstruct the truncated data based on the parsed set of residual samples and the set of truncation error values, wherein the reconstruction of the truncated data produces untruncated data. Although the following description may focus on data processing, the concepts described herein may be applicable to other similar processing techniques, such as graphics processing or display processing.
[0040] A device such as device 104 may refer to any device, apparatus, or system configured to perform one or more of the techniques described herein. For example, the device may be a server, a base station, a user equipment, a client device, a station, an access point, a computer (such as a personal computer, a desktop computer, a laptop computer, a tablet computer, a computer workstation, or a mainframe), a final product, a device, a telephone, a smartphone, a server, a video game platform or console, a handheld device (such as a portable video game device or a personal digital assistant (PDA)), a wearable computing device (such as a smartwatch, an augmented reality device, or a virtual reality device), a non-wearable device, a display or display device, a television, a set-top box, an intermediate network device, a digital media player, a video streaming device, a content streaming device, an in-vehicle computer, any mobile device, any device configured to generate graphical content, or any device configured to perform one or more of the techniques described herein. The processes herein may be described as being performed by particular components (e.g., a GPU), but in other embodiments, other components (e.g., a CPU) compliant with the disclosed embodiments may be used to perform them.
[0041] The GPU can process various types of data or data groups in the GPU pipeline. For example, in some aspects, the GPU can process two types of data or data groups, such as context register groups and draw call data. The context register group can be a set of global state information, such as information about global registers, shader programs, or constant data, which can adjust how the graphics context will be processed. For example, the context register group can include information about the color format. In some aspects of the context register group, there can be bits indicating which workload belongs to the context register. Additionally, multiple functions or programs can be run simultaneously and / or in parallel. For example, a function or program can describe an operation, such as a color mode or color format. Thus, the context register can define various states of the GPU.
[0042] The context state can be used to determine how a single processing unit (e.g., vertex fetcher (VFD), vertex shader (VS), shader processor, or geometry processor) operates and / or in which mode the processing unit operates. To do so, the GPU can use context registers and programming data. In some aspects, the GPU can generate workloads in the pipeline, such as vertex or pixel workloads, based on the context register definition of the mode or state. Certain processing units (e.g., VFD) can use these states to determine certain functions, such as how to aggregate vertices. Since these modes or states may change, the GPU may need to change the corresponding context. Additionally, the workload corresponding to the mode or state can follow the changed mode or state.
[0043] Figure 2 An example GPU 200 is illustrated in accordance with one or more techniques of the present disclosure. As Figure 2 shown, the GPU 200 includes a command processor (CP) 210, a draw call group 212, a VFD 220, a VS 222, a vertex cache (VPC) 224, a triangle setup engine (TSE) 226, a rasterizer (RAS) 228, a Z - process engine (ZPE) 230, a pixel interpolator (PI) 232, a fragment shader (FS) 234, a render backend (RB) 236, an L2 cache (UCHE) 238, and a system memory 240. Although Figure 2 shown that the GPU 200 includes processing units 220 to 238, the GPU 200 can include multiple additional processing units. Additionally, the processing units 220 to 238 are merely examples, and the GPU can use any combination or order of processing units according to the present disclosure. The GPU 200 also includes a command buffer 250, a context register group 260, and a context state 261.
[0044] As Figure 2As shown, the GPU can utilize the CP (such as CP 210) or a hardware accelerator to parse a command buffer into context register groups (such as context register group 260) and / or draw call data groups (such as draw call group 212). Subsequently, the CP 210 can transmit the context register group 260 or the draw call group 212 to a processing unit or block in the GPU via separate paths. Further, the command buffer 250 can alternate different states of context registers and draw calls. For example, the command buffer can be constructed in the following manner: context registers of context N, draw calls of context N, context registers of context N+1, and draw calls of context N+1.
[0045] The GPU can render images in a variety of different ways. In some instances, the GPU can use rasterization and / or tile-based rasterization to render images. In a tile-based rasterization GPU, an image can be divided or separated into different parts or tiles. After dividing the image, each part or tile can be rasterized individually. A tile-based rasterization GPU can divide a computer graphics image into a grid format such that each part of the grid (i.e., a tile) is rasterized individually. In some aspects, during a binning pass, the image can be divided into different bins or tiles. In some aspects, during a binning pass, a visibility stream can be constructed where visible primitives or draw calls can be identified. Contrary to tile-based rasterization, direct rasterization does not divide a frame into smaller bins or tiles. Instead, in direct rasterization, the entire frame is rasterized at once. Additionally, some types of GPUs may allow both tile-based rasterization and direct rasterization (e.g., flex rasterization).
[0046] In some aspects, a GPU can apply a drawing or rendering process to different bins or tiles. For example, the GPU can render for a bin and perform all the drawing for the primitives or pixels within the bin. During the process of rendering for a bin, the render target can be located in the GPU internal memory (GMEM). In some instances, after rendering for a bin, the content of the render target can be moved to the system memory and the GMEM can be freed to render the next bin. Additionally, the GPU can render for another bin and perform the drawing for the primitives or pixels within that bin. Thus, in some aspects, there may be a small number of bins, for example, four bins, that cover all the drawing in a surface. Further, the GPU can loop through all the drawing in a bin but perform the drawing calls that are visible, i.e., the drawing calls that contain visible geometry. In some aspects, a visibility stream can be generated, for example, during the binning process, to determine the visibility information of each primitive in an image or scene. For example, such a visibility stream can identify whether a particular primitive is visible. In some aspects, this information can be used to remove invisible primitives, for example, during the rendering process. Additionally, at least some of the primitives identified as visible can be rendered during the rendering process.
[0047] In some aspects of tile rendering, there can be multiple processing stages or processes. For example, the rendering can be performed in two passes, such as a visibility or bin-visibility pass and a rendering or bin-rendering pass. During the visibility pass, the GPU can input the rendering workload, record the positions of the primitives or triangles, and then determine which primitives or triangles fall into which bin or region. In some aspects of the visibility pass, the GPU can also identify or mark the visibility of each primitive or triangle in the visibility stream. During the rendering pass, the GPU can input the visibility stream and process one bin or region at a time. In some aspects, the visibility stream can be analyzed to determine which primitives or primitive vertices are visible or invisible. Thus, the visible primitives or primitive vertices can be processed. By doing so, the GPU can reduce the unnecessary workload of processing or rendering invisible primitives or triangles.
[0048] In some aspects, during a visibility process, certain types of primitive geometries can be processed, e.g., geometries that are only positioned. Additionally, based on the positioning or location of the primitives or triangles, the primitives can be classified into different bins or regions. In some instances, classifying the primitives or triangles into different bins can be performed by determining visibility information for these primitives or triangles. For example, the GPU can determine or write the visibility information for each primitive in each bin or region to, e.g., system memory. This visibility information can be used to determine or generate a visibility stream. During a rendering process, the primitives in each bin can be rendered individually. In these cases, the visibility stream can be fetched from memory for discarding primitives that are not visible for that bin.
[0049] Some aspects of a GPU or GPU architecture can provide multiple different options for rendering (e.g., software rendering and hardware rendering). In software rendering, the driver or CPU can replicate the entire frame geometry by processing each view Figure 1 pass. Additionally, some different states can change based on the view. Thus, in software rendering, the software can replicate the entire workload by changing some states that can be used to render for each viewpoint in an image. In some aspects, there can be increased overhead since the GPU may submit the same workload multiple times for each viewpoint in an image. In hardware rendering, the hardware or GPU may be responsible for replicating or processing the geometry for each viewpoint in an image. Thus, the hardware can manage the replication or processing of primitives or triangles for each viewpoint in an image.
[0050] Figure 3 An image or surface 300 is illustrated in accordance with one or more techniques of the present disclosure, including a plurality of primitives divided into a plurality of bins. As Figure 3 shown, the image or surface 300 includes a region 302 that includes primitives 321, 322, 323, and 324. The primitives 321, 322, 323, and 324 are divided or placed into different bins, e.g., bins 310, 311, 312, 313, 314, and 315. Figure 3 An example of tile-based rendering for primitives 321 - 324 using multiple viewpoints is illustrated. For example, the primitives 321 - 324 are in a first viewpoint 350 and a second viewpoint 351. Thus, the GPU processing or rendering the image or surface 300 including the region 302 can utilize multi-viewpoint or multi-view rendering.
[0051] As indicated herein, a GPU or Graphics Processing Unit may use a tile rendering architecture to reduce power consumption or save memory bandwidth. As further stated above, such a rendering method may divide a scene into multiple bins, as well as include visibility passes that identify the visible triangles in each bin. Thus, in tile rendering, a full screen may be divided into multiple bins or tiles. Then, the scene may be rendered multiple times, e.g., once or more for each bin.
[0052] In aspects of graphics rendering, some graphics applications may render to a single target (i.e., a render target) one or more times. For example, in graphics rendering, a frame buffer on system memory may be updated multiple times. A frame buffer may be part of memory or random access memory (RAM) (e.g., containing a bitmap or storage device) to help store display data for the GPU. A frame buffer may also be a memory buffer containing a full frame of data. Additionally, a frame buffer may be a logical buffer. In some aspects, updating a frame buffer may be performed in bin or tile rendering, where, as discussed above, a surface is divided into multiple bins or tiles and then each bin or tile may be rendered individually. Further, in tile rendering, a frame buffer may be divided into multiple bins or tiles.
[0053] As noted herein, in some aspects, such as in a bin or tiled rendering architecture, frame buffers may have data stored or written to them repeatedly, e.g., when rendering from different types of memory. This may be referred to as resolving and unresolving the frame buffer or system memory. For example, when storing or writing to one frame buffer and then switching to another, data or information on the frame buffer may be resolved from GMEM at the GPU to memory in double data rate (DDR) RAM or dynamic RAM (DRAM), i.e., system memory.
[0054] In some aspects, system memory may also be a system on a chip (SoC) memory or another chip - based memory for storing data or information, e.g., on a device or smartphone. System memory may also be a physical data storage device shared by the CPU and / or GPU. In some aspects, system memory may be, for example, a DRAM chip on a device or smartphone. Thus, SoC memory may be a chip - based way of storing data.
[0055] In some aspects, GMEM can be on-chip memory at the GPU, which can be implemented by static RAM (SRAM). Additionally, GMEM can be stored on a device (e.g., a smart phone). As indicated herein, data or information can be transferred, for example, between system memory or DRAM and GMEM at the device. In some aspects, the system memory or DRAM can be located at the CPU or GPU. Additionally, data can be stored in DDR or DRAM. In certain aspects, such as in bin or tile rendering, a small portion of the memory can be stored at the GPU, e.g., in GMEM. In certain cases, storing data at GMEM may consume a greater processing workload and / or power consumption compared to storing data at the frame buffer or system memory.
[0056] Various techniques related to improving the performance of lossy bandwidth compression / decompression are disclosed herein. These improvements can be applied to lossy formats (e.g., Red (R) Green (G) Blue (B) Alpha (A) 8888 (RGBA8888) format). Lossy bandwidth compression / decompression can be utilized by many different components of a device, such as a display, GPU, video decoder, camera, and CPU. Lossy bandwidth compression / decompression can be useful for a System-on-Chip (SOC) as the SOC can be configured to perform memory-intensive tasks where memory bandwidth may be limited. Lossy bandwidth compression / decompression can help save memory bandwidth by compressing surfaces stored in system memory. For example, a surface can be compressed when it is written to and read from the main memory. In one example, the surface can be a VR surface associated with a virtual reality (VR) application.
[0057] Lossy bandwidth compression / decompression may be associated with banding artifacts due to truncation that occurs in the compression portion of the lossy bandwidth compression / decompression. Banding artifacts can refer to a form of banding in a digital image caused by each pixel's color in the digital image being rounded to the nearest digital color level. Dithering can be utilized to reduce the occurrence of banding artifacts. Dithering can refer to an intentionally applied form of noise used to randomize quantization error in order to prevent banding artifacts. Various dithering methods can be employed to mitigate banding artifacts, such as ordered dithering methods, error diffusion methods, and frequency-based methods.
[0058] S-CIELAB may refer to a spatial extension of the CIELAB color metric developed by the International Commission on Illumination (CIE). S-CIELAB can be used to measure the color reproduction error of digital images. Given two images (an original image and a compressed / reconstructed image), S-CIELAB can calculate an error term (ΔE) related to the human perception of the error difference. For an image, the error term (ΔE) can be on a per-pixel basis. Small ΔE values may be imperceptible to a human observer, while large ΔE values can be discerned by a human observer. S-CIELAB can be used as an error metric to determine whether a compression scheme is visually lossless. For example, if ΔE = 0, the compression scheme is lossless. On the patterned area of an image, the reproduction error measured using S-CIELAB can correspond better to the perceived color error than the error calculated without S-CIELAB. On the uniform spatial area of an image, the error calculated using S-CIELAB can be equal to the error calculated using CIELAB.
[0059] Figure 4 FIG. 400 is a diagram illustrating an example of an encoder 402 and a decoder 405. The encoder 402 and the decoder 405 can utilize truncation error signaling and truncation error adaptive dithering (described in more detail below). The encoder 402 and the decoder 405 can be included in the same device or different devices. In one example, the encoder 402 and / or the decoder 405 can be included in a CPU, a GPU, or a DPU. In one example, the encoder 402 and / or the decoder 405 can be included in a SOC. The encoder 402 and / or the decoder 405 can be associated with a codec. The encoder 402 and / or the decoder 405 can be associated with display processing, image processing, or data processing. The encoder 402 and the decoder 405 can be associated with a lossy compression scheme. Lossy compression (also known as irreversible compression) may refer to a compressed form of representing content using an inexact approximation of the content and partial data discard. Lossy compression can be associated with a reduced data size for storing, handling, and transmitting the content. Lossless compression (also known as reversible compression) may refer to a compressed form that allows the reconstruction of the content from the compressed content without loss of information. Lossy compression can provide a greater amount of data reduction than lossless compression.
[0060] The encoder 402 can obtain a source block 404. The source block 404 can be an image, an image frame that is part of video content, an audio frame, or other data, or can include an image, an image frame that is part of video content, an audio frame, or other data. In one example, the source block 404 can be associated with display processing, image processing, or data processing. The encoder 402 can perform truncation 406 (also referred to herein as the "truncation process") on the source block 404. The truncation process can refer to removing one or more least significant bits (LSBs) from the data. The truncation 406 can produce truncated data 408. In one example, the source block 404 can have a first number of bits, and the truncated data 408 can have a second number of bits, where the first number of bits is greater than the second number of bits. In one example, the truncation 406 can truncate one or more least significant bits (LSBs) associated with the source block 404.
[0061] The truncation 406 can be associated with a truncation error for a sample. The truncation error for a sample can refer to the difference between the original version of the sample and the quantized / reconstructed version of the sample. The encoder 402 can calculate the truncation error (TE) for a sample (s) using a quantization parameter (Q) according to Equation (I) below.
[0062] (I) TE(s) = s - ((s >> Q) << Q)
[0063] In equation (I), TE may refer to the truncation error for a sample. The quantization parameter (Q) may enable a series of values associated with a sample to be compressed into discrete values. As used herein, the symbol ">>" may refer to a right shift. In one aspect, the right shift may be an arithmetic right shift in which the least significant bit is lost and the most significant bit is replicated. In one example, in the case of a given bit sequence "1001", a right shift of 1 for 1 (i.e., 1001>>1) may result in the bit sequence "1101". In another example, in the case of a given bit sequence "0011", a right shift of 2 for 0011 (i.e., 0011>>2) may result in the bit sequence "0000". In yet another example, in the case of a given bit sequence "1011", a right shift of 2 for 1011 (i.e., 1011>>2) may result in the bit sequence "10" (i.e., the right shift may directly reduce the number of bits). As used herein, the symbol "<<" may refer to a left shift in which a new least significant bit with a value of "0" is added to the bit sequence and the existing bits are promoted. In one example, in the case of a given bit sequence "1011", a left shift of 2 for 1011 (i.e., 1011<<2) may result in the bit sequence "101100". The encoder 402 may compute (and signal) the truncation error for both the predictive mode of operation and the pulse code modulation (PCM) mode of operation. The predictive mode of operation may refer to a mode of operation that uses adjacent samples to predict the current sample. The PCM mode of operation may refer to a mode of operation that directly truncates the sample without prediction. The PCM mode of operation may be useful in high entropy scenarios where the prediction performance is not optimal (i.e., scenarios involving difficult content). For example, if the residual is one bit larger than the original sample, the PCM mode of operation may be utilized instead of the predictive mode of operation. After accumulating the truncation error over sub-block components (i.e., groups of samples) via equation (I), the encoder 402 may compute an average truncation error 410 for the sub-block component (explained in more detail below). Although the encoder 402 is described herein as computing the average truncation error 410 for the sub-block component, the encoder may also compute at least one representative value that collectively represents the truncation error for the sub-block component. The average truncation error 410 may be an example of at least one representative value; however, other types of truncation error values may be computed and utilized by the encoder 402.
[0064] Figure 5FIG. 500 is an illustration of an example of truncation error signaling technique 502. The truncation error signaling technique 502 can be utilized by an encoder 402 to calculate an average truncation error for data (e.g., average truncation error 410 for source block 404). The truncation error signaling technique 502 can include a first technique 504, a second technique 506, a third technique 508, and a fourth technique 510. In one aspect, the encoder 402 can select the first technique 504, the second technique 506, the third technique 508, or the fourth technique 510 based on the decoding mode employed by the encoder 402.
[0065] In the first technique 504, the encoder 402 can calculate a single value of the average truncation error for each sub-block component. As used herein, the term "sub-block component" can refer to a group of samples. The sub-block component can be associated with the source block 404. In other words, the group of samples can be associated with the sub-block component. The encoder 402 can calculate the average truncation error value (e.g., average truncation error 410) according to Equation (II) below.
[0066] (II) TE avg =(TE M×N +bias)>>shift
[0067] In Equation (II), TE avg can refer to the average truncation error value for the sub-block component. M×N can refer to the size of the sub-block component (i.e., the size of the group of samples). M and N can be positive integers. In one example, M can correspond to the height of the samples, and N can correspond to the width of the samples. TE M×N can refer to the sum of the truncation errors for the samples within the sub-block component. The encoder 402 can calculate the shift and the bias according to Equation (III) and Equation (IV) below, respectively.
[0068] (III) shift = log2(MN)
[0069] (IV) bias = 1<<(log2(MN)-1)
[0070] In an example regarding the first technique 504, the encoder 402 can calculate the average truncation error value for a 16×4 sub-block component (e.g., M = 16 and N = 4). The encoder 402 can calculate the average truncation error value for this example according to Equation (V) below.
[0071] (V) TE avg =(TE 16×4 +32)>>6
[0072] In the second technique 506, multiple syntax elements of truncation error can be used for a single sub-block component. The encoder 402 can define the size of the region within the sub-block component as M0×N0, where M0≤M and N0≤N. The encoder 402 can calculate the average truncation error value (e.g., average truncation error 410) according to the following equation (VI).
[0073]
[0074] In equation (VI), TE avg can refer to the average truncation error value for the region within the sub-block component. Additionally, can refer to the sum of the truncation errors of the samples within the region within the sub-block component. The encoder 402 can calculate the shift and bias in the second technique 506 according to the following equations (VII) and (VIII).
[0075] (VII) Shift = log2(M0N0)
[0076] (VIII) Bias = 1 << (log2(M0N0) - 1)
[0077] In an example regarding the second technique 506, the encoder 402 can divide a 16×4 (i.e., M = 16 and N = 4) sub-block component into four regions of size 4×4 (i.e., M0 = 4 and N0 = 4). The encoder 402 can calculate the average truncation error value for the region within the sub-block component, for example, according to the following equation (IX).
[0078] (IX) TE avg = (TE 4×4 + 8) >> 4
[0079] In the third technique 508, the encoder 402 can clip the truncation error to avoid using Q + 1 bits for signaling. For example, the encoder 402 can calculate the average truncation error value (e.g., average truncation error 410) according to the following equation (X).
[0080] (X) TE avg = min((1 << Q) - 1, (TE M×N + bias) >> shift)
[0081] In the fourth technique 510, the encoder 402 can modify the calculated average truncation error value in order to reduce the number of bits used for signaling. When Q is greater than a given threshold (i.e., Q > Q T ), the encoder 402 can perform the modification. The encoder 402 can add additional quantization parameter (Q extra)Applied to the average truncation error value. For example, the encoder 402 can calculate the average truncation error value (e.g., average truncation error 410) according to the following equations (XI), (XII), and (XIII).
[0082] (XI) TE avg =(TE M×N + bias) >> shift
[0083] (XII) shift = log2(MN) + Q extra
[0084] (XIII) bias = 1 << (log2(MN) + Q extra - 1)
[0085] In the fourth technique 510, the encoder 402 can signal the TE avg >> Q extra . In an example where Q extra is "1", TE avg can be modified to (TE avg >> 1) (i.e., a signal with one less bit) before signaling. During reconstruction, the decoder 405 can add a bit back. For example, for Q > Q T (i.e., TE avg <<= Q extra ), the decoder 405 can compensate by modifying the parsed value of TE avg . TE avg <<= Q extra is equivalent to TE avg = TE avg << Q extra (i.e., the parsed value of TE avg can be left-shifted by Q extra bits before being used for reconstruction).
[0086] Return reference Figure 5 , the first predictor 412 of the encoder 402 can receive the truncated data 408 generated by the truncator 406. The first predictor 412 can generate a residual 414 (i.e., a set of residual samples) for the truncated data 408 based on the truncated data 408. The entropy encoder 416 of the encoder 402 can apply entropy coding to the residual 414 to generate an entropy-coded residual 418 (i.e., an encoded set of residual samples for the truncated data 408). In one example, the entropy-coded residual 418 can be a block fixed-length coded (BFLC) prediction residual.
[0087] The syntax generator 420 of the encoder 402 may receive the entropy-coded residual 418 and the average truncation error 410. The syntax generator 420 may generate a bitstream 422 based on the average truncation error 410 and the entropy-coded residual 418. The bitstream 422 may include the average truncation error 410 and the entropy-coded residual 418. The bitstream 422 may also include a block length associated with the entropy-coded residual 418 (e.g., BFLC block length) and additional syntax in the tile header.
[0088] The encoder 402 may provide the bitstream 422 to the decoder 405. In one example, the encoder 402 (or a device or software associated with the encoder 402) may send the bitstream 422 to the decoder 405 via a wired connection. In another example, the encoder 402 (or a device or software associated with the encoder 402) may send the bitstream 422 to the decoder 405 via a wireless connection.
[0089] The syntax parser 424 of the decoder 405 may obtain the bitstream 422. The syntax parser 424 may parse the entropy-coded residual 418 and the average truncation error 410 from the bitstream 422. The entropy decoder 426 of the decoder 405 may apply entropy decoding to the entropy-coded residual 418 to generate a residual 414. The second predictor 428 of the decoder 405 may generate a predictor 430 based on the residual 414.
[0090] The dither generator 432 of the decoder 405 may obtain the average truncation error 410 (parsed via the syntax parser 424). The dither generator 432 may determine a dither value 434 (e.g., an ordered dither value) based on the average truncation error 410 and the quantization parameter (Q) discussed above. The dither value 434 may be obtained from a dither matrix (explained in more detail below). Aspects of the dither generator 432 will be discussed in more detail below.
[0091] The decoder 405 may perform a reconstruction 436 to generate a reconstructed tile 438, where the reconstructed tile 438 may be a reconstructed version of the source tile 404. The reconstruction 436 may be based on the dither value 434, the average truncation error 410, the predictor 430, and the residual 414. For example, the decoder 405 may perform the reconstruction 436 via equation (XIV).
[0092] (XIV)s rec = ((p + res) << Q) + TE avg + D scaled
[0093] In equation (XIV), s recmay refer to a reconstructed sample value (e.g., reconstructed tile 438), p may be a predictor (e.g., predictor 430), res may be a signaled residual (e.g., residual 414 generated by entropy decoder 426), Q may be a quantization parameter, TE avg may be the average truncation error 410 parsed from bitstream 422, and D scaled may be a dither matrix from which a dither value 434 is obtained. The reconstructed tile 438 may be associated with fewer block artifacts as compared to tiles reconstructed using techniques different from the techniques described herein.
[0094] Figure 6 is a diagram 600 illustrating example aspects of truncation error-based adaptive dithering. Decoder 405 may utilize dithering (e.g., ordered dithering) in order to reduce visual artifacts associated with truncation 406 described above. The strength of the dither matrix may be adaptive based on the average truncation error value of a sub-tile component (or region within a sub-tile component). Diagram 600 includes a first example 602 and a second example 604.
[0095] The first example 602 may correspond to a range of truncated LSB values from 0 to (1<<M)–1. In one example, M may be a quantization parameter such as quantization parameter Q described above in connection with Figure 4 and Figure 5 . In the first example 602, the average truncation error (e.g., average truncation error 410) may be approximately half of the range of truncated LSB values. In the first example 602, the dither strength (D min , D max ) may be a function of the average truncation error. The dither strength (D min , D max ) may correspond to added noise. In the first example 602, the dither strength may be symmetric about the reconstruction point (corresponding to the average truncation error) and may cover approximately half of the range of truncated LSB values. The dither strength may be symmetric about the reconstruction point in order to avoid biasing the reconstructed signal and to average to the same level of truncation error. In the first example 602, the dither strength may be maximized for a truncation error that is approximately half of the dynamic range (i.e., the range of truncated LSB values).
[0096] The second example 604 may correspond to a range of truncated LSB values from 0 to (1<<M)–1. In one example, M may be a quantization parameter such as quantization parameter Q described above in connection with Figure 4 and Figure 5 . In the second example 604, the average truncation error (e.g., average truncation error 410) may be close to 0. In the second example 604, the dither strength (D min , D max) can be a function of the average truncation error. The dither strength (D min , D max ) can correspond to the added noise. In the second example 604, the dither strength can be symmetric about the reconstruction point (with respect to the average truncation error) and can be a relatively small portion of the range of truncated LSB values. The dither strength can be symmetric about the reconstruction point to avoid biasing the reconstructed signal and to average the truncation error to the same level. In the second example 604, the dither strength can be minimized for truncation errors close to zero or close to (1 << M) – 1.
[0097] Return reference Figure 4 , in one aspect, the dither generator 432 of the decoder 405 can determine the dither matrix D avg ) based on the quantization parameter (e.g., Q) and the size of the average truncation error (TE max ) for each 4×4 region associated with the source tile 404, through the dither strength parameter (d scaled ) described in Table 1 below.
[0098] Q <![CDATA[TE avg range]]> <![CDATA[d max > x <![CDATA[TE avg = 0]]> 0 x <![CDATA[TE avg =(1<<Q)-1]]> 0 1 x 0 2 x 0 3 <![CDATA[1 ≤ TE avg ≤ 6]]> 1 4 <![CDATA[1≤TE avg <3]]> <![CDATA[TE avg > 4 <![CDATA[3 ≤ TE avg ≤ 12]]> 3 4 <![CDATA[12<TE avg ≤14]]> <![CDATA[15-TE avg > 5 <![CDATA[1≤TE avg <7]]> <![CDATA[TE avg > 5 <![CDATA[7 ≤ TE avg ≤ 24]]> 7 5 <![CDATA[24<TE avg ≤30]]> <![CDATA[31-TE avg >
[0099] Table 1: Mapping between the average truncation error and the dither strength parameter d for quantifying the choice of parameter Q max therebetween
[0100] In one aspect, the dither generator 432 of the decoder 405 can programmatically determine the dither strength parameter (d avg ) based on the average truncation error (TE max ) and the quantization parameter (Q) according to the following equations (XV), (XVI), and (XVII).
[0101]
[0102] (XVI) α = (1 << (Q - 2)) - 1
[0103] (XVII) β = ((1 << Q) - 1) - α
[0104] After the dither generator 432 of the decoder 405 determines the dither strength parameter d max programmatically (either via Table 1 or via equations (XV), (XVI), and (XVII)), the dither generator 432 can calculate the scaled dither matrix (D base ) based on the base dither matrix (D max ) and the dither strength parameter d scaled ). The following equations (XVIII) and (XIX) provide the calculation of the scaled dither matrix D scaled and the base dither matrix D base .
[0105] (XVIII) D scaled = (D base ·d max + 4) >> 3
[0106]
[0107] Although D in equation (XIX) base is a 4×4 matrix, other possibilities are expected. For example, the decoder 405 can utilize D of different sizes based on M and N (as described above) base . Further, in the above equation (XVIII), “+4” can facilitate rounding, and “>>3” can reduce the range of D scaled .
[0108] The dither generator 432 of the decoder 405 can determine the dither value 434 for a given sample during reconstruction by sampling D at the location of the samples (s y % 4, s x % 4) within the sub-tile component (s y , s x ). The symbol “%” can refer to the modulo operator. As described above, the decoder 405 can generate the reconstructed tile 438 based on the dither value 434, the average truncation error 410, the predictor 430, and the residual 414 scaled .
[0109] In one aspect, the encoder 402 can signal an additional average truncation error to improve performance. For example, the encoder 402 can signal an average truncation error value for each M×N region. Signaling the additional average error can improve performance when the quantization parameter Q is relatively high. In one aspect, the decoder 405 can rotate the base dither matrix D base to produce a matrix with the same performance. In one aspect, the decoder 405 can calculate the dither strength to be asymmetric about the reconstruction point (corresponding to the average truncation error). For example, if the average truncation error (TE avg ) is close to 0 or close to (1 << Q) – 1, the decoder 405 can allow the dither strength to be asymmetric, thereby providing a greater dither strength closer to 0 or closer to (1 << Q) – 1
[0110] Compared with other compression / decompression schemes, the aspects described above can be associated with a reduced error metric (such as a reduced mean squared error (MSE) metric or a reduced S-CIELAB metric)
[0111] Figure 7Call Flow Diagram 700 illustrates an example communication between an encoder 702 and a decoder 704 in accordance with one or more techniques of the present disclosure. In one example, encoder 702 and / or decoder 704 may be included in a DPU, GPU, or CPU. Encoder 702 and decoder 704 may be within the same device, or encoder 702 and decoder 704 may be in different devices. In one example, encoder 702 may be or may include encoder 402, and decoder 704 may be or may include decoder 405.
[0112] At 706, encoder 702 may obtain data that may be associated with data processing, image processing, or display processing. At 708, encoder 702 may perform a truncation process on the data, which results in truncated data. At 710, encoder 702 may generate a set of residual samples for the truncated data. At 712, encoder 702 may calculate a set of average truncation error values associated with the truncation process for the truncated data. At 714, encoder 702 may encode the set of residual samples for the truncated data. At 716, encoder 702 may generate a bitstream based on the set of residual samples and the set of average truncation error values. At 717, encoder 702 may store the bitstream. At 718, encoder 702 may send the bitstream to decoder 704.
[0113] At 720, decoder 704 may obtain a bitstream associated with a set of residual samples for the truncated data and a set of average truncation error values for the truncation process performed at 708. The bitstream may correspond to data associated with data processing, image processing, or display processing. At 722, decoder 704 may parse the set of residual samples and the set of average truncation error values from the bitstream. At 724, decoder 704 may decode the parsed set of residual samples for the truncated data. At 726, decoder 704 may perform a dithering process on the truncated data based on the set of average truncation error values. At 728, decoder 704 may reconstruct the truncated data based on the dithering process and the decoded set of residual samples. Reconstruction of the truncated data may result in untruncated data. At 730, decoder 704 may send, store, or process the untruncated data after reconstructing the truncated data.
[0114] Figure 8 Flowchart 800 is a flowchart of an example method of data processing in accordance with one or more techniques of the present disclosure. The method may be performed by a device such as: a device for data processing, GPU, CPU, display processing unit (DPU), or other display processor, wireless communication device, etc., as combined with Figures 1 to 7used in various aspects thereof. The method may be performed by encoder 402 or encoder 702. The method may be associated with various advantages, such as reducing visual artifacts generated as a byproduct of the compression / decompression process. In one example, the method may be performed by compressor / decompressor 198.
[0115] At 802, the apparatus performs a truncation process on data, where the data is associated with display processing, image processing, or data processing, and where the truncation process on the data produces truncated data. For example, Figure 7 At 708, it is shown that encoder 702 may perform a truncation process on data, where the data may be associated with display processing, image processing, or data processing. In another example, Figure 4 it is shown that encoder 402 may perform truncation 406 on source tile 404 (i.e., data). In additional examples, the truncated data may be truncated data 408. In one example, 802 may be performed by compressor / decompressor 198.
[0116] At 804, the apparatus calculates a set of truncation error values associated with the truncation process on the truncated data. For example, Figure 7 At 712, it is shown that encoder 702 may calculate a set of average truncation error values associated with the truncation process on the truncated data performed at 708. In another example, the set of truncation error values may be average truncation error 410 or may include the average truncation error. In additional examples, the apparatus may use one or more of the truncation error signaling techniques 502 to calculate the set of truncation error values. In yet another example, the set of truncation error values may correspond to the TE avg described above. In one example, 804 may be performed by compressor / decompressor 198.
[0117] At 806, the apparatus generates a set of residual samples for the truncated data. For example, Figure 7 At 710, it is shown that encoder 702 may generate a set of residual samples for the truncated data. In additional examples, Figure 7 At 714, it is shown that encoder 702 may encode the set of residual samples for the truncated data. In another example, Figure 4 it is shown that encoder 402 may encode residual 414 to generate entropy-coded residual 418. In one example, 806 may be performed by compressor / decompressor 198.
[0118] At 808, the apparatus generates a bitstream based on the set of residual samples for the truncated data and the set of truncation error values associated with the truncation process. For example, Figure 7At 716, it is shown that the encoder 702 can generate a bitstream based on a set of residual samples and a set of mean truncation error values. In another example, the bitstream can be the bitstream 422. In one example, 808 can be performed by the compressor / decompressor 198.
[0119] Figure 9 is a flowchart 900 of an example method of data processing according to one or more techniques of the present disclosure. The method can be performed by a device such as: a device for data processing, a GPU, a CPU, a DPU, or other display processors, wireless communication devices, etc., as used in connection with Figures 1 to 7 aspects thereof. The method can be performed by the encoder 402 or the encoder 702. The method can be associated with various advantages, such as reducing visual artifacts generated as a byproduct of the compression / decompression process. In one example, the method (including various aspects detailed below) can be performed by the compressor / decompressor 198.
[0120] At 904, the device performs a truncation process on the data, where the data is associated with display processing, image processing, or data processing, and where the truncation process on the data generates truncated data. For example, Figure 7 At 708, it is shown that the encoder 702 can perform a truncation process on the data, where the data can be associated with display processing, image processing, or data processing. In another example, Figure 4 it is shown that the encoder 402 can perform truncation 406 on the source tile 404 (i.e., the data). In additional examples, the truncated data can be the truncated data 408. In one example, 904 can be performed by the compressor / decompressor 198.
[0121] At 906, the device calculates a set of truncation error values associated with the truncation process on the truncated data. For example, Figure 7 At 712, it is shown that the encoder 702 can calculate a set of mean truncation error values associated with the truncation process on the truncated data performed at 708. In another example, the set of truncation error values can be the mean truncation error 410 or can include the mean truncation error. In additional examples, the device can use one or more of the truncation error signaling techniques 502 to calculate the set of truncation error values. In yet another example, the set of truncation error values can correspond to the TE avg described above. In one example, 906 can be performed by the compressor / decompressor 198.
[0122] At 908, the device generates a set of residual samples for the truncated data. For example, Figure 7 At 710, it is shown that the encoder 702 can generate a set of residual samples for the truncated data. In additional examples, Figure 7At 714, it is shown that the encoder 702 can encode a set of residual samples for the truncated data. In another example, Figure 4 it is shown that the encoder 402 can encode the residual 414 to generate an entropy-coded residual 418. In one example, 908 can be performed by the compressor / decompressor 198.
[0123] At 910, the device generates a bitstream based on a set of residual samples for the truncated data and a set of truncation error values associated with the truncation process. For example, Figure 7 At 716, it is shown that the encoder 702 can generate a bitstream based on the set of residual samples and the set of average truncation error values. In another example, the bitstream can be the bitstream 422. In one example, 910 can be performed by the compressor / decompressor 198.
[0124] In one aspect, at 912, the device can send or store the generated bitstream. For example, Figure 7 At 717, it is shown that the encoder 702 can store the bitstream generated at 716. In another example, Figure 7 At 718, it is shown that the encoder 702 can send the bitstream to the decoder 704. In additional examples, Figure 4 it is shown that the encoder 402 can provide the bitstream 422 to the decoder 405. In one example, 912 can be performed by the compressor / decompressor 198.
[0125] In one aspect, sending the generated bitstream can include sending the generated bitstream from the encoding device to the decoding device, or storing the generated bitstream can include storing the generated bitstream in a memory or a second memory, cache, or buffer associated with the encoding device. For example, the encoder 702 can be included in the encoding device, and the decoder 704 can be included in the decoding device. In another example, storing the bitstream at 717 can include storing the generated bitstream in a memory or a second memory, cache, or buffer associated with the encoder 702. In additional examples, the generated bitstream can be stored in one or more of the internal memory 121, the system memory 124, the internal memory 123, or the command buffer 250.
[0126] In one aspect, performing the truncation process on the data can include performing at least one bit-shift operation on the samples of the data based on quantization parameters. For example, the truncation process for the data can be performed according to equation (I) above.
[0127] In one aspect, calculating a set of truncation error values associated with a truncation process for truncated data may include calculating an average truncation error value based on a sum, shift, bias, and at least one bit shift operation of truncation errors for a group of samples associated with one or more of the data. For example, the set of truncation error values may be calculated according to equations (II), (III), and (IV) above. For example, calculating the set of truncation error values may include aspects described above with respect to the first technique 504.
[0128] In one aspect, calculating a set of truncation error values associated with a truncation process for truncated data may include calculating an average truncation error value based on a sum, shift, bias, and at least one bit shift operation of truncation errors for samples in a region within a group of samples associated with one or more of the data. For example, the set of truncation error values may be calculated according to equations (VI), (VII), and (VIII) above. For example, calculating the set of truncation error values may include aspects described above with respect to the second technique 506.
[0129] In one aspect, calculating a set of truncation error values associated with a truncation process for truncated data may include calculating an average truncation error value, and calculating the average truncation error value may include calculating a first average truncation error value based on at least one of a quantization parameter, a first value, and a first at least one bit shift operation. For example, the first average truncation error value may be (1<<Q)–1 as shown in equation (X) above. For example, calculating the set of truncation error values may include aspects described above with respect to the third technique 508.
[0130] In one aspect, calculating a set of truncation error values associated with a truncation process for truncated data may include calculating an average truncation error value, and calculating the average truncation error value may include calculating a second average truncation error value based on a sum, bias, shift, and a second at least one bit shift operation of truncation errors for a group of samples associated with one or more of the data. For example, the second average truncation error value may be (TE MxN + bias) >> shift as shown in equation (X) above. For example, calculating the set of truncation error values may include aspects described above with respect to the third technique 508.
[0131] In one aspect, calculating a set of truncation error values associated with a truncation process for truncated data may include calculating an average truncation error value, and calculating the average truncation error value may include selecting the smaller of a first average truncation error value or a second average truncation error value to be used as the average truncation error value. For example, the set of truncation error values may be calculated according to equation (X) above. In one example, the "min" operation in equation (X) may select the smaller of the first average truncation error value or the second average truncation error value to be used as the average truncation error value. For example, calculating the set of truncation error values may include aspects described above with respect to the third technique 508.
[0132] In one aspect, calculating a set of truncation error values associated with a truncation process for truncated data may include calculating an average truncation error value based on a sum, shift, bias, quantization parameter factor, and at least one bit shift operation of truncation errors for a set of samples associated with one or more of the data. For example, the set of truncation error values may be calculated according to equations (XI), (XII), and (XIII) above. For example, calculating the set of truncation error values may include aspects described above with respect to the fourth technique 510.
[0133] In one aspect, if the quantization parameter is greater than a threshold quantization parameter, the average truncation error value may be calculated based on a sum, shift, bias, quantization parameter factor, and at least one bit shift operation of truncation errors for the set of samples. For example, the quantization parameter may be Q above, and the threshold quantization parameter may be Q above T For example, calculating the set of truncation error values may include aspects described above with respect to the fourth technique 510.
[0134] In one aspect, a set of truncation error values for a truncation process in a bitstream may include an average truncation error value for each region within a set of samples associated with the data. For example, the average truncation error 410 may include an average truncation error value for each region within a set of samples associated with the data. In another example, the region may be defined by M and N as described above.
[0135] In one aspect, generating a set of residual samples for truncated data may include encoding the set of residual samples for the truncated data. For example, Figure 7 As shown at 710, the encoder 702 may generate a set of residual samples for the truncated data, and the set of residual samples may be encoded at 714. In another example, the set of residual samples may be the residual 414.
[0136] In one aspect, encoding a set of residual samples for truncated data may include entropy encoding the set of residual samples for truncated data. For example, encoding the set of residual samples at 714 may include entropy encoding the set of residual samples. In another example, entropy encoder 416 may encode the residual 414 to generate entropy decoded residual 418.
[0137] In one aspect, at 902, the apparatus may obtain data. For example, Figure 7 At 706, it is shown that encoder 702 may obtain data. In one example, 902 may be performed by compressor / decompressor 198.
[0138] In one aspect, obtaining data may include: receiving data from a GPU, CPU, or camera. For example, obtaining data at 706 may include receiving data from a GPU, CPU, or camera. In another example, a GPU, CPU, or camera may be included in device 104.
[0139] In one aspect, calculating a set of truncation error values may include calculating at least one representative value that collectively represents the set of truncation error values. For example, the at least one representative value may include average truncation error 410. In another example, the at least one representative value may be the set of average truncation error values calculated at 712.
[0140] Figure 10 is a flowchart 1000 of an example method of data processing according to one or more techniques of the present disclosure. The method may be performed by an apparatus such as: an apparatus for data processing, a GPU, a CPU, a DPU, or other display processor, a wireless communication device, etc., as used in connection with Figures 1 to 7 the aspects described. The method may be performed by decoder 405 or decoder 704. The method may be associated with various advantages, such as reducing visual artifacts generated as a byproduct of the compression / decompression process. In one example, the method may be performed by compressor / decompressor 198.
[0141] At 1002, the apparatus obtains a bitstream associated with a set of residual samples for truncated data and a set of truncation error values for a truncation process, where the bitstream corresponds to data associated with display processing, image processing, or data processing. For example, Figure 7At 720, it is shown that decoder 704 can obtain a bitstream associated with a set of residual samples for truncated data and a set of average truncation error values for a truncation process, where the bitstream corresponds to data associated with display processing, image processing, or data processing. In one example, the bitstream can be bitstream 422, the set of residual samples for truncated data can be residual 414 (or entropy-coded residual 418), the set of truncation error values can be average truncation error 410 or can include the average truncation error, the truncation process can be truncation 406, and the truncated data can be truncated data 408. In one example, 1002 can be performed by compressor / decompressor 198.
[0142] At 1004, the apparatus parses from the bitstream a set of residual samples for truncated data and a set of truncation error values to obtain a set of parsed residual samples for the truncated data. For example, Figure 7 At 722, it is shown that decoder 704 can parse from the bitstream a set of residual samples and a set of average truncation error values. In another example, Figure 4 It is shown that syntax parser 424 of decoder 405 can parse bitstream 422 to obtain average truncation error 410 and entropy-coded residual 418. For example, Figure 7 At 724, it is shown that decoder 704 can decode the set of parsed residual samples for the truncated data. In another example, Figure 4 It is shown that entropy decoder 426 can decode entropy-coded residual 418 to obtain residual 414. In one example, 1004 can be performed by compressor / decompressor 198.
[0143] At 1006, the apparatus reconstructs the truncated data based on the set of decoded residual samples and the set of truncation error values, where reconstruction of the truncated data results in untruncated data. For example, Figure 7 At 728, it is shown that decoder 704 can reconstruct the truncated data based on the set of determined residual samples obtained at 724 and the set of average truncation error values parsed from the bitstream at 722. Figure 4 It is also shown that reconstruction 436 can generate reconstructed tile 438 (i.e., untruncated data). In additional examples, the apparatus can reconstruct the truncated data according to equation (XIV) above. In one example, 1006 can be performed by compressor / decompressor 198.
[0144] Figure 11 Is a flowchart 1100 of an example method of data processing according to one or more techniques of the present disclosure. The method can be performed by an apparatus such as: an apparatus for data processing, GPU, CPU, DPU, or other display processor, wireless communication device, etc., as combined with Figures 1 to 7used in various aspects thereof. The method may be performed by decoder 405 or decoder 704. The method may be associated with various advantages, such as reducing visual artifacts generated as a byproduct of the compression / decompression process. In one example, the method (including various aspects detailed below) may be performed by compressor / decompressor 198.
[0145] At 1102, the apparatus obtains a bitstream associated with a set of residual samples for truncated data and a set of truncation error values for a truncation process, where the bitstream corresponds to data associated with display processing, image processing, or data processing. For example, Figure 7 At 720, it is shown that decoder 704 may obtain a bitstream associated with a set of residual samples for truncated data and a set of average truncation error values for a truncation process, where the bitstream corresponds to data associated with display processing, image processing, or data processing. In one example, the bitstream may be bitstream 422, the set of residual samples for truncated data may be residual 414 (or entropy-coded residual 418), the set of truncation error values may be average truncation error 410 or may include the average truncation error, the truncation process may be truncation 406, and the truncated data may be truncated data 408. In one example, 1102 may be performed by compressor / decompressor 198.
[0146] At 1104, the apparatus parses the set of residual samples for truncated data and the set of truncation error values from the bitstream to obtain a set of parsed residual samples for the truncated data. For example, Figure 7 At 722, it is shown that decoder 704 may parse the set of residual samples and the set of average truncation error values from the bitstream. In another example, Figure 4 It is shown that syntax parser 424 of decoder 405 may parse bitstream 422 to obtain average truncation error 410 and entropy-coded residual 418. For example, Figure 7 At 724, it is shown that decoder 704 may decode the set of parsed residual samples for truncated data. In another example, Figure 4 It is shown that entropy decoder 426 may decode entropy-coded residual 418 to obtain residual 414. In one example, 1104 may be performed by compressor / decompressor 198.
[0147] At 1108, the apparatus reconstructs the truncated data based on the set of parsed residual samples and the set of truncation error values, where the reconstruction of the truncated data results in untruncated data. For example, Figure 7 At 728, it is shown that decoder 704 may reconstruct the truncated data based on the set of determined residual samples obtained at 724 and the set of average truncation error values parsed from the bitstream at 722. Figure 4It is also shown that the reconstruction 436 can generate reconstructed tiles 438 (i.e., untruncated data). In another example, the device can reconstruct the truncated data according to the above equation (XIV). In one example, 1108 can be performed by the compressor / decompressor 198.
[0148] In one aspect, parsing a set of residual samples and a set of truncation error values for the truncated data can include entropy decoding the set of residual samples for the truncated data. For example, parsing the set of residual samples and the set of average truncation error values at 722 and decoding the parsed set of residual samples for the truncated data at 724 can include entropy decoding the parsed set of residual samples for the truncated data. In another example, Figure 4 It is shown that the entropy decoder 426 of the decoder 405 can entropy decode the entropy-coded residuals 418.
[0149] In one aspect, at 1110, the device can transmit, store, or process the untruncated data. For example, Figure 7 It is shown at 730 that the decoder 704 can transmit, store, or process the untruncated data after reconstructing the truncated data at 728. In one example, 1110 can be performed by the compressor / decompressor 198.
[0150] In one aspect, transmitting the untruncated data can include sending the untruncated data from the decoding device to a display, storing the untruncated data can include storing the untruncated data in a memory or a second memory, cache, or buffer associated with the decoding device, or processing the untruncated data can include processing the untruncated data at the decoding device. For example, transmitting, storing, or processing the untruncated data at 730 can include sending the untruncated data from the decoding device to a display, storing the untruncated data can include storing the untruncated data in a memory or a second memory, cache, or buffer associated with the decoding device, or processing the untruncated data can include processing the untruncated data at the decoding device. In another example, the decoder 704 can be included in the decoding device. In another example, the display can be the display 131 or can include the display. In another example, the untruncated data can be stored at the internal memory 121, the system memory 124, or the internal memory 123.
[0151] In one aspect, at 1106, the device can perform a dithering process for the truncated data based on the set of truncation error values, and the reconstruction of the truncated data can be further based on the dithering process. For example, Figure 7 It is shown at 726 that the decoder 704 can perform a dithering process for the truncated data based on the set of average truncation error values. In another example, Figure 4It is shown that the dither generator 432 can perform a dithering process on the truncated data 408 based on the average truncation error 410. In another example, the apparatus can perform the dithering process according to one or more of the above equations (XV), (XVI), (XVII), (XVIII), or (XIX) or Table 1 above. In additional examples, Figure 7 At 728, it is shown that the decoder 704 can reconstruct the truncated data based on the dithering process performed at 726 and the set of decoded residual samples obtained at 724. In another example, Figure 4 It is shown that the decoder 405 can perform a reconstruction 436 based on the dither value 434 generated by the dither generator 432 and the residual 414. Figure 4 It is also shown that the reconstruction 436 can generate a reconstructed tile 438 (i.e., untruncated data). In additional examples, the apparatus can reconstruct the truncated data according to the above equation (XIV). In one example, 1106 can be performed by the compressor / decompressor 198.
[0152] In one aspect, performing the dithering process can include determining a dither intensity parameter from a look-up table (LUT) based on a truncation error value associated with a set of truncation error values and a quantization parameter, where the LUT can include multiple dither intensity parameters for multiple truncation error values and multiple quantization parameters. For example, Table 1 above can be the LUT, the truncation error value can be TE avg , and the quantization parameter can be Q. Further, as illustrated above, Table 1 can include a mapping between the average truncation error and the dither intensity parameter d max for the selection of the quantization parameter.
[0153] In one aspect, performing the dithering process can include obtaining a dither intensity parameter, where if the truncation error value associated with the set of truncation error values is a first error value, the dither intensity parameter can be a first value, where if the truncation error value is a second error value, the dither intensity parameter can be a second value, where the first value is greater than the second value, and where the first error value is greater than the second error value. For example, Figure 6 the first example 602 and the second example 604 of min respectively show that if the average truncation error value is large, the dither intensity (D max ) can be relatively large, and if the average truncation error value is small, the dither intensity (D min ) can be relatively small. The dither intensity (D max ) can correspond to the dither intensity parameter d min , D max . max
[0154] In one aspect, performing a dithering process may include calculating at least one dither matrix. For example, the at least one dither matrix may be a scaled dither matrix (D scaled ) and a base dither matrix (D base ), or may include the scaled dither matrix and the base dither matrix.
[0155] In one aspect, the at least one dither matrix may include a scaled dither matrix and a base dither matrix. For example, the scaled dither matrix may be D scaled , and the base dither matrix may be D base .
[0156] In one aspect, calculating the scaled dither matrix may include calculating the scaled dither matrix based on one or more of the following: the base dither matrix, a dither strength parameter, at least one value associated with the base dither matrix and the dither strength parameter, and at least one bit shift operation. For example, the apparatus may calculate the scaled dither matrix according to Equation (XVIII) and Equation (XIX) above.
[0157] In one aspect, the dither strength parameter may be symmetric or asymmetric with respect to a truncation error value. For example, the first example 602 and the second example 604 illustrate that the dither strength (D min , D max ) may be symmetric with respect to the average truncation error (E). In another example, the dither strength (D min , D max ) may be asymmetric with respect to the average truncation error (E). The dither strength (D min , D max ) may correspond to the dither strength parameter d max .
[0158] In one aspect, the scaled dither matrix may be associated with at least one ordered dither value. For example, D scaled may be associated with at least one ordered dither value.
[0159] In one aspect, performing the dithering process may include rotating the base dither matrix. For example, the dithering process performed at 726 may include rotating the base dither matrix D base .
[0160] In one aspect, reconstructing the truncated data based on the dithering process and the set of parsed residual samples may include sampling the scaled dither matrix at the locations of the samples within the sample group. For example, reconstructing the truncated data at 728 may include sampling the scaled dither matrix at the locations of the samples within the sample group. In another example, the dither generator 432 of the decoder 405 may sample within the sub-block component (s y % 4, s x % 4) of the samples (sy ,s x ) to sample D at the location of scaled and determine the dither value 434 for a given sample during reconstruction.
[0161] In one aspect, performing the dithering process can include calculating a dither intensity parameter based on one or more of a truncation error value, quantization parameters, and at least one additional parameter. For example, the dither intensity parameter can be d as described above max . In another example, the dither intensity parameter can be calculated according to one or more of the above equations (XV), (XVI), and (XVII).
[0162] In the configuration, a method or apparatus for graphics processing is provided. The apparatus can be a GPU, a CPU, or some other processor capable of performing graphics processing. In various aspects, the apparatus can be the processing unit 120 within device 104, or can be some other hardware within device 104 or another device. The apparatus can include components for performing a truncation process on data, where the data is associated with display processing, image processing, or data processing, and where the truncation process on the data produces truncated data. The apparatus can also include components for calculating a set of truncation error values associated with the truncation process on the truncated data. The apparatus can also include components for generating a set of residual samples for the truncated data. The apparatus can also include components for generating a bitstream based on the set of residual samples for the truncated data and the set of truncation error values associated with the truncation process. The apparatus can also include components for transmitting or storing the generated bitstream. The components for transmitting the generated bitstream can include components for transmitting the generated bitstream from an encoding device to a decoding device, and the components for storing the generated bitstream can include components for storing the generated bitstream in a memory or a second memory, cache, or buffer associated with the encoding device. The components for performing a truncation process on data can include components for performing at least one bit shift operation on samples of the data based on a quantization parameter. The components for calculating a set of truncation error values associated with the truncation process can include components for calculating an average truncation error value based on the sum, shift, bias, and at least one bit shift operation of the truncation errors for a group of samples associated with one or more of the data. The components for calculating a set of truncation error values associated with the truncation process can include components for calculating an average truncation error value based on the sum, shift, bias, and at least one bit shift operation of the truncation errors for samples in a region within a group of samples associated with one or more of the data. The components for calculating a set of truncation error values associated with the truncation process can include components for calculating an average truncation value. The components for calculating an average truncation value can include components for calculating a first average truncation error value based on at least one of a quantization parameter, a first value, and a first at least one bit shift operation. The components for calculating an average truncation value can include components for calculating a second average truncation error value based on the sum, bias, shift, and a second at least one bit shift operation of the truncation errors for a group of samples associated with one or more of the data. The components for calculating an average truncation value can include components for selecting the smaller of the first average truncation error value or the second average truncation error value to be used as the average truncation error value. The components for generating a set of residual samples for the truncated data can include components for encoding the set of residual samples for the truncated data. The components for encoding the set of residual samples for the truncated data can include components for entropy encoding the set of residual samples for the truncated data. The apparatus can also include components for obtaining the data.The apparatus may further include components for obtaining a bitstream associated with a set of residual samples for truncated data and a set of truncation error values for a truncation process, where the bitstream corresponds to data associated with display processing, image processing, or data processing. The apparatus may further include components for parsing from the bitstream the set of residual samples for truncated data and a set of average truncation error values to obtain a parsed set of residual samples for the truncated data. The apparatus may further include components for reconstructing the truncated data based on the parsed set of residual samples and the set of truncation error values, where reconstruction of the truncated data results in untruncated data. The components for parsing the set of residual samples may include components for entropy decoding the set of residual samples for the truncated data. The apparatus may further include components for transmitting, storing, or processing the untruncated data. The components for transmitting the untruncated data may include components for transmitting the untruncated data from the decoding device to a display. The components for storing the untruncated data may include components for storing the untruncated data in a memory associated with the decoding device or a second memory, cache, or buffer. The components for processing the untruncated data may include components for processing the untruncated data at the decoding device. The apparatus may further include components for performing a dithering process for the truncated data based on the set of truncation error values prior to reconstructing the truncated data, where reconstruction of the truncated data is further based on the dithering process. The components for performing the dithering process may include components for calculating a dithering intensity parameter based on one or more of the following: a truncation error value associated with the set of truncation error values, a quantization parameter, and at least one additional parameter. The components for performing the dithering process may include components for determining a dithering intensity parameter from a look-up table (LUT) based on the average truncation error value associated with the set of truncation error values and the quantization parameter, where the LUT includes a plurality of dithering intensity parameters for a plurality of truncation error values and a plurality of quantization parameters. The components for performing the dithering process may include components for calculating at least one dither matrix. The components for calculating a scaled dither matrix may include components for calculating the scaled dither matrix based on one or more of the following: a base dither matrix, the dithering intensity parameter, at least one value associated with the base dither matrix and the dithering intensity parameter, and at least one bit shift operation. The components for performing the dithering process may include components for rotating the base dither matrix. The components for reconstructing the truncated data based on the dithering process and the decoded set of residual samples may include components for sampling the scaled dither matrix at the location of samples within a group of samples.
[0163] In various configurations, a method or apparatus for display processing is provided. The apparatus can be a DPU, a display processor, or some other processor capable of performing display processing. In various aspects, the apparatus can be the display processor 127 within device 104, or can be some other hardware within device 104 or another device. The apparatus can include components for performing a truncation process on data, where the data is associated with display processing, image processing, or data processing, and where the truncation process on the data produces truncated data. The apparatus can further include components for calculating a set of truncation error values associated with the truncation process on the truncated data. The apparatus can further include components for generating a set of residual samples for the truncated data. The apparatus can further include components for generating a bitstream based on the set of residual samples for the truncated data and the set of truncation error values associated with the truncation process. The apparatus can further include components for transmitting or storing the generated bitstream. The components for transmitting the generated bitstream can include components for transmitting the generated bitstream from an encoding device to a decoding device, and the components for storing the generated bitstream can include components for storing the generated bitstream in a memory or a second memory, cache, or buffer associated with the encoding device. The components for performing the truncation process on data can include components for performing at least one bit shift operation on samples of the data based on quantization parameters. The components for calculating the set of truncation error values associated with the truncation process can include components for calculating an average truncation error value based on a sum, shift, bias, and at least one bit shift operation of truncation errors for a group of samples associated with one or more of the data. The components for calculating the set of truncation error values associated with the truncation process can include components for calculating an average truncation error value based on a sum, shift, bias, and at least one bit shift operation of truncation errors for samples in a region within a group of samples associated with one or more of the data. The components for calculating the set of truncation error values associated with the truncation process can include components for calculating an average truncation value. The components for calculating the average truncation value can include components for calculating a first average truncation error value based on at least one of quantization parameters, a first value, and a first at least one bit shift operation. The components for calculating the average truncation value can include components for calculating a second average truncation error value based on a sum, bias, shift, and a second at least one bit shift operation of truncation errors for a group of samples associated with one or more of the data. The components for calculating the average truncation value can include components for selecting the smaller of the first average truncation error value or the second average truncation error value to be used as the average truncation error value. The components for generating the set of residual samples for the truncated data can include components for encoding the set of residual samples for the truncated data. The components for encoding the set of residual samples for the truncated data can include components for entropy encoding the set of residual samples for the truncated data. The apparatus can further include components for obtaining the data.The apparatus may further include components for obtaining a bitstream associated with a set of residual samples for truncated data and a set of truncation error values for a truncation process, where the bitstream corresponds to data associated with display processing, image processing, or data processing. The apparatus may further include components for parsing from the bitstream the set of residual samples for truncated data and a set of average truncation error values to obtain a parsed set of residual samples for the truncated data. The apparatus may further include components for reconstructing the truncated data based on the parsed set of residual samples and the set of truncation error values, where reconstruction of the truncated data yields untruncated data. The components for parsing the set of residual samples may include components for entropy decoding the set of residual samples for the truncated data. The apparatus may further include components for transmitting, storing, or processing the untruncated data. The components for transmitting the untruncated data may include components for sending the untruncated data from the decoding device to a display. The components for storing the untruncated data may include components for storing the untruncated data in a memory associated with the decoding device or a second memory, cache, or buffer. The components for processing the untruncated data may include components for processing the untruncated data at the decoding device. The apparatus may further include components for performing a dithering process for the truncated data based on the set of truncation error values prior to reconstructing the truncated data, where reconstruction of the truncated data is further based on the dithering process. The components for performing the dithering process may include components for calculating a dithering intensity parameter based on one or more of the following: a truncation error value associated with the set of truncation error values, a quantization parameter, and at least one additional parameter. The components for performing the dithering process may include components for determining a dithering intensity parameter from a look-up table (LUT) based on the average truncation error value associated with the set of truncation error values and the quantization parameter, where the LUT includes a plurality of dithering intensity parameters for a plurality of truncation error values and a plurality of quantization parameters. The components for performing the dithering process may include components for calculating at least one dither matrix. The components for calculating a scaled dither matrix may include components for calculating the scaled dither matrix based on one or more of the following: a base dither matrix, the dithering intensity parameter, at least one value associated with the base dither matrix and the dithering intensity parameter, and at least one shift operation. The components for performing the dithering process may include components for rotating the base dither matrix. The components for reconstructing the truncated data based on the dithering process and the decoded set of residual samples may include components for sampling the scaled dither matrix at locations of samples within a group of samples.
[0164] It should be understood that the specific order or hierarchy of the blocks / steps in the processes, flowcharts, and / or call flowcharts disclosed herein are merely illustrative of example methods. It should be understood that based on design preferences, the specific order or hierarchy of these blocks / steps in the processes, flowcharts, and / or call flowcharts can be rearranged. Additionally, some blocks / steps may be combined or omitted. Other blocks / steps may also be added. The appended method claims present the elements of the various blocks / steps in a sample order, but are not meant to be limited to the specific order or hierarchy presented.
[0165] The foregoing description is provided to enable any person skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects. Thus, the claims are not intended to be limited to the aspects shown herein, but are to be accorded the full scope consistent with the language of the claims, wherein the reference to an element in the singular is not intended to mean "one and only one" unless specifically so stated, but rather "one or more". The word "exemplary" is used herein to mean "serving as an example, instance, or illustration". Any aspect described herein as "exemplary" is not necessarily to be construed as preferred or having an advantage over other aspects.
[0166] Unless otherwise specifically stated, the term "some" refers to one or more, and the term "or" can be construed as "and / or" where the context does not otherwise dictate. Combinations such as "at least one of A, B, or C", "one or more of A, B, or C", "at least one of A, B, and C", "one or more of A, B, and C", and "any combination of A, B, C, or any of them", including any combination of A, B, and / or C, can include multiple A's, multiple B's, or multiple C's. Specifically, combinations such as "at least one of A, B, or C", "one or more of A, B, or C", "at least one of A, B, and C", "one or more of A, B, and C", and "any combination of A, B, C, or any of them" can be A only, B only, C only, A and B, A and C, B and C, or A and B and C, where any such combination can include one or more members of A, B, or C. All structural and functional equivalents of the elements of the various aspects described throughout this disclosure that are known or later will be known to those of ordinary skill in the art are expressly incorporated herein by reference and are intended to be covered by the claims. Additionally, nothing disclosed herein is intended to be dedicated to the public, whether or not such disclosure is expressly recited in the claims. The words "module", "mechanism", "element", "device", etc. are not intended to substitute for the word "component". Thus, no claim element should be construed as a means-plus-function unless the element is expressly recited using the phrase "means for...".
[0167] In one or more examples, the functions described herein may be implemented in hardware, software, firmware, or any combination thereof. For example, although the term "processing unit" is used throughout this disclosure, such processing units may be implemented in hardware, software, firmware, or any combination thereof. If any functions, processing units, techniques, or other modules described herein are implemented in software, the functions, processing units, techniques, or other modules described herein may be stored on or transmitted over a computer-readable medium as one or more instructions or code.
[0168] A computer-readable medium may include computer data storage media and communication media including any medium that facilitates transfer of a computer program from one place to another. In this manner, a computer-readable medium generally may correspond to: (1) a non-transitory tangible computer-readable storage medium; or (2) a communication medium such as a signal or carrier. A data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. By way of example, and not limitation, such computer-readable media may include RAM, ROM, EEPROM, compact disc read only memory (CD-ROM) or other optical disc storage, magnetic disk storage or other magnetic storage devices. As used herein, disks and optical discs include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where disks generally reproduce data magnetically, while optical discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media. A computer program product may include a computer-readable medium.
[0169] The techniques of this disclosure may be implemented in a variety of devices or apparatuses, including wireless handsets, integrated circuits (ICs) or IC sets (e.g., chip sets). Various components, modules, or units are described in this disclosure to emphasize functional aspects of devices configured to perform the disclosed techniques, but need not necessarily be implemented by different hardware units. Instead, as described above, the various units may be combined in any hardware unit or provided by a collection of interoperable hardware units (including one or more processors as described above) in conjunction with suitable software and / or firmware. Thus, the term "processor" as used herein may refer to any of the above structures or any other structure suitable for implementing the techniques described herein. Similarly, these techniques may be fully implemented in one or more circuits or logic elements.
[0170] The following aspects are merely illustrative and may be combined with other aspects or teachings described herein without limitation.
[0171] Aspect 1 is a method for data processing, the method comprising: performing a truncation process on data, wherein the data is associated with display processing, image processing, or the data processing, and wherein the truncation process on the data generates truncated data; calculating a set of truncation error values associated with the truncation process on the truncated data; generating a set of residual samples for the truncated data; and generating a bitstream based on the set of residual samples for the truncated data and the set of truncation error values associated with the truncation process.
[0172] Aspect 2 can be combined with aspect 1 and further comprises transmitting or storing the generated bitstream.
[0173] Aspect 3 can be combined with aspect 2 and comprises: transmitting the generated bitstream includes transmitting the generated bitstream from an encoding device to a decoding device, or wherein storing the generated bitstream includes storing the generated bitstream in a memory or a second memory, cache, or buffer associated with the encoding device.
[0174] Aspect 4 can be combined with any one of aspects 1 to 3 and comprises: performing the truncation process on the data includes performing at least one bit shift operation on samples of the data based on a quantization parameter.
[0175] Aspect 5 can be combined with any one of aspects 1 to 4 and comprises: calculating the set of truncation error values associated with the truncation process on the truncated data includes calculating an average truncation error value based on the sum, shift, bias, and at least one bit shift operation of the truncation errors for a group of samples associated with one or more of the data.
[0176] Aspect 6 can be combined with any one of aspects 1 to 4 and comprises: calculating the set of truncation error values associated with the truncation process on the truncated data includes calculating an average truncation error value based on the sum, shift, bias, and at least one bit shift operation of the truncation errors for samples in a region within a group of samples associated with one or more of the data.
[0177] Aspect 7 can be combined with any one of aspects 1 to 4 and comprises: calculating the set of truncation error values associated with the truncation process on the truncated data includes calculating an average truncation error value, wherein calculating the average truncation error value includes: calculating a first average truncation error value based on at least one of a quantization parameter, a first value, and a first at least one bit shift operation; calculating a second average truncation error value based on the sum, bias, shift, and a second at least one bit shift operation of the truncation errors for a group of samples associated with one or more of the data; and selecting the smaller of the first average truncation error value and the second average truncation error value to be used as the average truncation error value.
[0178] Aspect 8 can be combined with any one of Aspects 1 to 4, and includes: calculating the set of truncation error values associated with the truncation process for the truncated data includes calculating an average truncation error value based on a sum, shift, bias, quantization parameter factor, and at least one bit shift operation of truncation errors for a sample group associated with one or more of the data.
[0179] Aspect 9 can be combined with Aspect 8, and includes that in the case where the quantization parameter is greater than a threshold quantization parameter, the average truncation error value is calculated based on the sum, the shift, the bias, the quantization parameter factor, and the at least one bit shift operation of the truncation errors for the sample group.
[0180] Aspect 10 can be combined with any one of Aspects 1 to 9, and includes: the set of truncation error values for the truncation process in the bitstream includes an average truncation error value for each region within a sample group associated with the data.
[0181] Aspect 11 can be combined with any one of Aspects 1 to 10, and includes: generating the set of residual samples for the truncated data includes encoding the set of residual samples for the truncated data.
[0182] Aspect 12 can be combined with Aspect 11, and includes: encoding the set of residual samples for the truncated data includes entropy encoding the set of residual samples for the truncated data.
[0183] Aspect 13 can be combined with any one of Aspects 1 to 12, and includes: calculating the set of truncation error values includes calculating at least one representative value that collectively represents the set of truncation error values.
[0184] Aspect 14 can be combined with Aspect 13, and includes: obtaining the data includes: receiving the data from a graphics processing unit (GPU), a central processing unit (CPU), or a camera.
[0185] Aspect 15 is an apparatus for data processing, the apparatus includes at least one processor, the at least one processor is coupled to a memory, and at least partially based on information stored in the memory, the at least one processor is configured to implement the method according to any one of Aspects 1 to 14.
[0186] Aspect 16 can be combined with Aspect 15, and includes that the apparatus is a wireless communication device, the apparatus further includes at least one of an antenna or a transceiver coupled to the at least one processor, and Aspect 16 further includes obtaining the data via at least one of the antenna or the transceiver.
[0187] Aspect 17 is an apparatus for data processing, the apparatus comprising components for implementing the method according to any one of Aspects 1 to 14.
[0188] Aspect 18 is a computer-readable medium (e.g., a non-transitory computer-readable medium) storing computer-executable code, the code causing at least one processor to implement the method according to any one of Aspects 1 to 14 when executed by the at least one processor.
[0189] Aspect 19 is a method for data processing, the method comprising: obtaining a bitstream associated with a set of residual samples for truncated data and a set of truncation error values for a truncation process, wherein the bitstream corresponds to data associated with display processing, image processing, or the data processing; parsing the set of residual samples for the truncated data and the set of average truncation error values from the bitstream to obtain a parsed set of residual samples for the truncated data; and reconstructing the truncated data based on the decoded set of residual samples and the set of truncation error values, wherein the reconstruction of the truncated data results in untruncated data.
[0190] Aspect 20 can be combined with Aspect 19 and includes: parsing the set of residual samples for the truncated data and the set of truncation error values includes entropy decoding the set of residual samples for the truncated data.
[0191] Aspect 21 can be combined with any one of Aspects 19 to 20 and further includes transmitting, storing, or processing the untruncated data.
[0192] Aspect 22 can be combined with Aspect 21 and includes: transmitting the untruncated data includes sending the untruncated data from a decoding device to a display, wherein storing the untruncated data includes storing the untruncated data in a memory or a second memory, cache, or buffer associated with the decoding device, or wherein processing the untruncated data includes processing the untruncated data at the decoding device.
[0193] Aspect 23 can be combined with any one of Aspects 19 to 22 and further includes performing a dithering process for the truncated data based on the set of truncation error values before reconstructing the truncated data, wherein reconstructing the truncated data is further based on the dithering process.
[0194] Aspect 24 can be combined with Aspect 23 and includes: performing the dithering process includes determining a dithering intensity parameter from a look-up table (LUT) based on a truncation error value associated with the set of truncation error values and a quantization parameter, wherein the LUT includes a plurality of dithering intensity parameters for a plurality of truncation error values and a plurality of quantization parameters.
[0195] Aspect 25 can be combined with any one of aspects 23 to 24, and includes: performing the dithering process includes obtaining a dithering intensity parameter, wherein when the truncation error value associated with the set of truncation error values is a first error value, the dithering intensity parameter is a first value, wherein when the truncation error value is a second error value, the dithering intensity parameter is a second value, wherein the first value is greater than the second value, and wherein the first error value is greater than the second error value.
[0196] Aspect 26 can be combined with any one of aspects 23 to 25, and includes: performing the dithering process includes calculating at least one dither matrix.
[0197] Aspect 27 can be combined with aspect 26, and includes: the at least one dither matrix includes a scaled dither matrix and a base dither matrix.
[0198] Aspect 28 can be combined with aspect 27, and includes: calculating the scaled dither matrix includes calculating the scaled dither matrix based on one or more of the following: the base dither matrix, the dithering intensity parameter, at least one value associated with the base dither matrix and the dithering intensity parameter, and at least one shift operation.
[0199] Aspect 29 can be combined with aspect 28, and includes that the dithering intensity parameter is symmetric or asymmetric with respect to the truncation error value.
[0200] Aspect 30 can be combined with any one of aspects 27 to 29, and includes that the scaled dither matrix is associated with at least one ordered dither value.
[0201] Aspect 31 can be combined with any one of aspects 27 to 30, and includes: performing the dithering process includes rotating the base dither matrix.
[0202] Aspect 32 can be combined with any one of aspects 27 to 31, and includes: reconstructing the truncated data based on the dithering process and the set of decoded residual samples includes sampling the scaled dither matrix at the locations of the samples within the sample group.
[0203] Aspect 33 can be combined with aspect 23 and any one of aspects 25 to 32, and includes: performing the dithering process includes calculating the dithering intensity parameter based on one or more of the following: the truncation error value associated with the set of truncation error values, the quantization parameter, or at least one additional parameter.
[0204] Aspect 34 is an apparatus for data processing, the apparatus including at least one processor, the at least one processor being coupled to a memory, and the at least one processor being configured to implement the method according to any one of aspects 19 to 33, at least in part based on information stored in the memory.
[0205] Aspect 35 can be combined with aspect 34 and includes that the apparatus is a wireless communication device, the apparatus further including at least one of an antenna or a transceiver coupled to the at least one processor, wherein, in order to obtain the bit stream, the at least one processor is configured to receive the bit stream via at least one of the antenna or the transceiver.
[0206] Aspect 36 is an apparatus for data processing, the apparatus including components for implementing the method according to any one of aspects 19 to 33.
[0207] Aspect 37 is a computer-readable medium (e.g., a non-transitory computer-readable medium) storing computer-executable code that, when executed by at least one processor, causes the at least one processor to implement the method according to any one of aspects 19 to 33.
[0208] Aspects have been described herein. These aspects and other aspects are within the scope of the following claims.
Claims
1. An apparatus for data processing, the apparatus comprising: a memory; and at least one processor coupled to the memory and configured to, at least in part based on information stored in the memory: perform a truncation process on data, where the data is associated with display processing, image processing, or the data processing, and where the truncation process on the data produces truncated data; calculate a set of truncation error values associated with the truncation process on the truncated data; generate a set of residual samples for the truncated data; and generate a bitstream based on the set of residual samples for the truncated data and the set of truncation error values associated with the truncation process.
2. The apparatus according to claim 1, wherein the at least one processor is further configured to: send or store the generated bitstream.
3. The apparatus according to claim 2, wherein, in order to send the generated bitstream, the at least one processor is configured to send the generated bitstream from an encoding device to a decoding device, or wherein, in order to store the generated bitstream, the at least one processor is configured to store the generated bitstream in the memory or a second memory, cache, or buffer associated with the encoding device.
4. The apparatus according to claim 1, wherein, in order to perform the truncation process on the data, the at least one processor is configured to perform at least one bit shift operation on samples of the data based on a quantization parameter.
5. The apparatus according to claim 1, wherein, in order to calculate the set of truncation error values associated with the truncation process on the truncated data, the at least one processor is configured to calculate an average truncation error value based on a sum, shift, bias, or at least one bit shift operation of truncation errors for a group of samples associated with one or more of the data.
6. The apparatus according to claim 1, wherein, in order to calculate the set of truncation error values associated with the truncation process on the truncated data, the at least one processor is configured to calculate an average truncation error value based on a sum, shift, bias, or at least one bit shift operation of truncation errors for samples in a region within a group of samples associated with one or more of the data.
7. The apparatus according to claim 1, wherein, in order to calculate the set of truncation error values associated with the truncation process on the truncated data, the at least one processor is configured to calculate an average truncation error value, and wherein, in order to calculate the average truncation error value, the at least one processor is configured to: calculate a first average truncation error value based on at least one of a quantization parameter, a first value, and a first at least one bit shift operation; calculate a second average truncation error value based on a sum of truncation errors for a group of samples associated with one or more of the data, a bias, a shift, or a second at least one bit shift operation; and select the smaller of the first average truncation error value or the second average truncation error value to be used as the average truncation error value.
8. The apparatus according to claim 1, wherein, in order to calculate the set of truncation error values associated with the truncation process for the truncated data, the at least one processor is configured to calculate an average truncation error value based on a sum, shift, bias, quantization parameter factor, or at least one shift operation of truncation errors for a set of samples associated with one or more of the data.
9. The apparatus according to claim 8, wherein, in order to calculate the average truncation error value, the at least one processor is configured to calculate the average truncation error value based on the sum, the shift, the bias, the quantization parameter factor, and the at least one shift operation of the truncation errors for the set of samples when the quantization parameter is greater than a threshold quantization parameter.
10. The apparatus according to claim 1, wherein the set of truncation error values for the truncation process in the bitstream includes an average truncation error value for each region within the set of samples associated with the data.
11. The apparatus according to claim 1, wherein, in order to generate the set of residual samples for the truncated data, the at least one processor is configured to encode the set of residual samples for the truncated data.
12. The apparatus according to claim 11, wherein, in order to encode the set of residual samples for the truncated data, the at least one processor is configured to perform entropy encoding on the set of residual samples for the truncated data.
13. The apparatus according to claim 1, wherein, in order to calculate the set of truncation error values, the at least one processor is configured to: calculate at least one representative value that collectively represents the set of truncation error values.
14. The apparatus according to claim 1, wherein the apparatus is a wireless communication device, and the apparatus further includes at least one of an antenna or a transceiver coupled to the at least one processor, and wherein the at least one processor is further configured to: obtain the data via at least one of the antenna or the transceiver.
15. An apparatus for data processing, the apparatus comprising: a memory; and at least one processor coupled to the memory, and at least partially based on information stored in the memory, the at least one processor is configured to: obtain a bitstream associated with a set of residual samples for truncated data and a set of truncation error values for a truncation process, wherein the bitstream corresponds to data associated with display processing, image processing, or the data processing; parse the set of residual samples for the truncated data and the set of truncation error values from the bitstream to obtain a parsed set of residual samples for the truncated data; and reconstruct the truncated data based on the parsed set of residual samples and the set of truncation error values, wherein the reconstruction of the truncated data results in untruncated data.
16. The apparatus according to claim 15, wherein, in order to parse the set of residual samples and the set of truncation error values for the truncated data, the at least one processor is configured to perform entropy decoding on the set of residual samples for the truncated data.
17. The apparatus according to claim 15, wherein the at least one processor is further configured to: Transmit, store, or process the untruncated data.
18. The apparatus according to claim 17, wherein, in order to transmit the untruncated data, the at least one processor is configured to transmit the untruncated data from the decoding device to a display, or wherein, in order to store the untruncated data, the at least one processor is configured to store the untruncated data in the memory or a second memory, cache, or buffer associated with the decoding device, or wherein, in order to process the untruncated data, the at least one processor is configured to process the untruncated data at the decoding device.
19. The apparatus according to claim 15, wherein the at least one processor is further configured to: Before the at least one processor is configured to reconstruct the truncated data, perform a dithering process on the truncated data based on the set of truncation error values, wherein, in order to reconstruct the truncated data, the at least one processor is configured to further reconstruct the truncated data based on the dithering process.
20. The apparatus according to claim 19, wherein, in order to perform the dithering process, the at least one processor is configured to determine a dithering intensity parameter from a look-up table (LUT) based on a truncation error value associated with the set of truncation error values and a quantization parameter, wherein the LUT includes a plurality of dithering intensity parameters for a plurality of truncation error values and a plurality of quantization parameters.
21. The apparatus according to claim 19, wherein, in order to perform the dithering process, the at least one processor is configured to obtain a dithering intensity parameter, wherein, when the truncation error value associated with the set of truncation error values is a first error value, the dithering intensity parameter is a first value, and wherein, when the truncation error value is a second error value, the dithering intensity parameter is a second value, wherein the first value is greater than the second value, and wherein the first error value is greater than the second error value.
22. The apparatus according to claim 19, wherein, in order to perform the dithering process, the at least one processor is configured to calculate at least one dithering matrix.
23. The apparatus according to claim 22, wherein the at least one dithering matrix includes a scaled dithering matrix and a base dithering matrix.
24. The apparatus according to claim 23, wherein, in order to calculate the scaled dithering matrix, the at least one processor is configured to calculate the scaled dithering matrix based on one or more of the following: the base dithering matrix, the dithering intensity parameter, at least one value associated with the base dithering matrix and the dithering intensity parameter, or at least one shift operation.
25. The apparatus according to claim 24, wherein the dither intensity parameter is symmetric or asymmetric with respect to the truncation error value.
26. The apparatus according to claim 23, wherein, in order to perform the dithering process, the at least one processor is configured to rotate the base dither matrix.
27. The apparatus according to claim 19, wherein, in order to perform the dithering process, the at least one processor is configured to calculate a dither intensity parameter based on one or more of: a truncation error value associated with the set of truncation error values, a quantization parameter, or at least one additional parameter.
28. The apparatus according to claim 15, wherein the apparatus is a wireless communication device, and the apparatus further includes at least one of an antenna or a transceiver coupled to the at least one processor, and wherein, in order to obtain the bitstream, the at least one processor is configured to receive the bitstream via at least one of the antenna or the transceiver.
29. A method for data processing, the method comprising: performing a truncation process on data, wherein the data is associated with display processing, image processing, or the data processing, and wherein the truncation process on the data produces truncated data; calculating a set of truncation error values associated with the truncation process for the truncated data; generating a set of residual samples for the truncated data; and generating a bitstream based on the set of residual samples for the truncated data and the set of truncation error values associated with the truncation process.
30. A method for data processing, the method comprising: obtaining a bitstream associated with a set of residual samples for truncated data and a set of truncation error values for a truncation process, wherein the bitstream corresponds to data associated with display processing, image processing, or the data processing; parsing the set of residual samples for the truncated data and the set of truncation error values from the bitstream to obtain a parsed set of residual samples for the truncated data; and reconstructing the truncated data based on the parsed set of residual samples and the set of truncation error values, wherein the reconstruction of the truncated data produces untruncated data.