Pixel preconditioning process for non-native compression
By performing pixel preconditioning on the bit set and processing the first subset based on the parity of the second subset, the problem of large prediction error in non-native compression is solved, achieving a higher compression ratio and resource saving.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- QUALCOMM INC
- Filing Date
- 2024-12-05
- Publication Date
- 2026-07-21
AI Technical Summary
Existing technologies suffer from high prediction errors in non-native compression, leading to reduced compression ratios and increased bandwidth requirements.
By performing pixel preconditioning on the bit set and processing the first subset based on the parity of the second subset, discontinuities are removed to improve the compression ratio.
It improves the compression ratio and saves computing and network resources.
Smart Images

Figure CN122439313A_ABST
Abstract
Description
Cross-references to related applications
[0001] This application claims the benefit of U.S. Non-Provisional Patent Application Serial No. 18 / 398,032, filed on December 27, 2023, entitled “PIXEL PRECONDITIONING FOR NON-NATIVE COMPRESSION”, the entire contents of which are expressly incorporated herein by reference. Technical Field
[0002] This disclosure relates in general to processing systems, and more specifically to one or more techniques for data processing. Background Technology
[0003] Computing devices typically perform graphics and / or display processing (e.g., utilizing a graphics processing unit (GPU), a central processing unit (CPU), a display processor, etc.) to render and display visual content. Such computing devices can include, for example, computer workstations, mobile phones (such as smartphones), embedded systems, personal computers, tablet computers, and video game consoles. A GPU is configured to execute a graphics processing pipeline that includes one or more processing stages that operate together to execute graphics processing commands and output frames. A CPU controls the operation of a GPU by issuing one or more graphics processing commands to it. Modern CPUs are typically capable of executing multiple applications concurrently, each of which may require the use of a GPU during execution. A display processor can be configured to convert digital information received from the CPU into analog values and can issue commands to a display panel to display visual content. Devices that provide content for visual presentation on a display can utilize a CPU, GPU, and / or display processor.
[0004] Current techniques for non-native compression are often associated with relatively high prediction errors. Improvements in techniques related to non-native compression are needed. Summary of the Invention
[0005] The following is a simplified summary of one or more aspects to provide a basic understanding of these aspects. This summary is not a broad overview of all anticipated aspects, nor is it intended to identify key or essential elements of all aspects, nor to describe the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simplified form as a prelude to the more detailed description that follows.
[0006] In one aspect of this disclosure, a method, computer-readable medium, and apparatus for data processing are provided. The apparatus includes: a memory; and a processor coupled to the memory, and configured, based on information stored in the memory, to: obtain a set of bits including a first subset and a second subset, wherein the bit width of the bit set is greater than the configured bit width of compression hardware; process the first subset based on parity corresponding to the second subset; and output a decoder for the compression hardware a set of bits including the processed first and second subsets.
[0007] In one aspect of this disclosure, a method, computer-readable medium, and apparatus for data processing are provided. The apparatus includes: a memory; and a processor coupled to the memory, and configured, based on information stored in the memory, to: obtain a set of bits including a first subset and a second subset from an encoder of compression hardware, wherein the bit width of the bit set is greater than the configured bit width of the compression hardware; process the first subset based on parity corresponding to the second subset; and output a set of bits including the processed first subset and second subset.
[0008] To achieve the foregoing and related objectives, one or more aspects include the features fully described below and specifically pointed out in the claims. The following description and drawings set forth some exemplary features of one or more aspects in detail. However, these features indicate only a few of the various ways in which the principles of the various aspects may be employed, and this description is intended to include all such aspects and their equivalents. Attached Figure Description
[0009] Figure 1 This is a block diagram illustrating an example of a system for generating content based on one or more techniques of this disclosure.
[0010] Figure 2 Example GPUs based on one or more techniques according to this disclosure are illustrated.
[0011] Figure 3 Example images or surfaces are illustrated according to one or more techniques of this disclosure.
[0012] Figure 4 This is a diagram illustrating an example of native compression according to one or more techniques of this disclosure.
[0013] Figure 5 This is a diagram illustrating an example of non-native compression according to one or more techniques of this disclosure.
[0014] Figure 6 This is a diagram illustrating an example of native decompression according to one or more techniques of this disclosure.
[0015] Figure 7 This is a diagram illustrating an example of non-native decompression according to one or more techniques of this disclosure.
[0016] Figure 8 This is an illustration of a first plot of the most significant bit (MSB) and the native bit component values and a second plot of the least significant bit (LSB) and the native bit component values according to one or more techniques of this disclosure.
[0017] Figure 9 These are illustrations of example aspects related to pixel preconditioning according to one or more techniques of this disclosure.
[0018] Figure 10 These are illustrations of example aspects of encoding related to one or more techniques according to this disclosure.
[0019] Figure 11 These are illustrations illustrating example aspects of decoding related to one or more techniques according to this disclosure.
[0020] Figure 12 This is a call flowchart illustrating example communication between an encoder and a decoder according to one or more techniques of this disclosure.
[0021] Figure 13 This is a flowchart of an example method for graphical processing according to one or more techniques of this disclosure.
[0022] Figure 14 This is a flowchart of an example method for graphical processing according to one or more techniques of this disclosure.
[0023] Figure 15 This is a flowchart of an example method for graphical processing according to one or more techniques of this disclosure.
[0024] Figure 16 This is a flowchart of an example method for graphical processing according to one or more techniques of this disclosure. Detailed Implementation
[0025] Various aspects of the systems, apparatuses, computer program products, and methods will be described more fully below with reference to the accompanying drawings. However, this disclosure may be embodied in many different forms and should not be construed as limited to any particular structure or function presented throughout this disclosure. Rather, these aspects are provided to make this disclosure comprehensive and complete, and to fully convey the scope of this disclosure to those skilled in the art. Based on the teachings herein, those skilled in the art will understand that the scope of this disclosure is intended to cover any aspect of the systems, apparatuses, computer program products, and methods disclosed herein, whether implemented independently of or in combination with other aspects of this disclosure. For example, any number of aspects set forth herein may be used to implement an apparatus or practice. Furthermore, the scope of this disclosure is intended to cover such apparatuses or methods implemented using structures, functionalities, or structures and functionalities other than or different from the various aspects of the disclosure set forth herein. Any aspect disclosed herein may be embodied by one or more elements of the claims.
[0026] Although various aspects are described herein, many variations and substitutions of these aspects fall within the scope of this disclosure. While some potential benefits and advantages of the aspects of this disclosure are mentioned, the scope of this disclosure is not intended to be limited to a particular benefit, use, or objective. Rather, the aspects of this disclosure are intended to be broadly applicable to different wireless technologies, system configurations, processing systems, networks, and transmission protocols, some of which are illustrated by way of example in the accompanying drawings and the description below. The detailed description and drawings are merely illustrative and not limiting of this disclosure, and the scope of this disclosure is defined by the appended claims and their equivalents.
[0027] Several aspects are presented with reference to various apparatuses and methods. These apparatuses and methods are described in detail and illustrated in the accompanying drawings by various blocks, components, circuits, processes, algorithms, etc. (collectively referred to as "elements"). These elements can be implemented using electronic hardware, computer software, or any combination thereof. Whether such elements are implemented as hardware or software depends on the specific application and the design constraints imposed on the system as a whole.
[0028] For example, an element, any part of an element, or any combination of elements can be implemented as a “processing system” including one or more processors (which may also be referred to as processing units). Examples of processors include microprocessors, microcontrollers, graphics processing units (GPUs), general-purpose GPUs (GPGPUs), central processing units (CPUs), application processors, digital signal processors (DSPs), reduced instruction set computing (RISC) processors, system-on-a-chip (SoCs), baseband processors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), programmable logic devices (PLDs), state machines, gated logic units, discrete hardware circuits, and other suitable hardware configured to perform the various functionalities described throughout this disclosure. One or more processors in the processing system can execute software. Whether referred to as software, firmware, middleware, microcode, hardware description language, or other names, software is broadly understood to mean instructions, instruction sets, code, code segments, program code, programs, subroutines, software components, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, functions, etc.
[0029] The term "application" can refer to software. As described herein, one or more technologies can refer to an application (e.g., software) configured to perform one or more functions. In such examples, the application may be stored in memory (e.g., on-chip memory of a processor, system memory, or any other memory). Hardware described herein, such as a processor, may be configured to execute the application. For example, an application may be described as including code that, when executed by the hardware, causes the hardware to perform one or more technologies described herein. As an example, the hardware may access and execute code accessed from memory to perform one or more technologies described herein. In some examples, components are identified in this disclosure. In such examples, a component may be hardware, software, or a combination thereof. Each component may be a separate component or a subcomponent of a single component. As used herein, the term "bit component" may refer to a set of bits.
[0030] In one or more examples described herein, the described functionality can be implemented in hardware, software, or any combination thereof. If implemented in software, the functionality can be stored or encoded as one or more instructions or code on a computer-readable medium. Computer-readable media includes computer storage media. Storage media can be any available medium accessible by a computer. By way of example, and not limitation, such computer-readable media can include random access memory (RAM), read-only memory (ROM), electrically erasable programmable ROM (EEPROM), optical disc storage devices, magnetic disk storage devices, other magnetic storage devices, combinations of computer-readable media of the types described above, or any other medium that can be used to store computer-executable code in the form of instructions or data structures accessible by a computer.
[0031] As used herein, instances of the term "content" may refer to "graphic content," "image," etc., regardless of whether the term is used as an adjective, noun, or other part of speech. In some examples, as used herein, the term "graphic content" may refer to content produced by one or more processes in a graphics processing pipeline. In other examples, as used herein, the term "graphic content" may refer to content produced by a processing unit configured to perform graphics processing. In yet another example, as used herein, the term "graphic content" may refer to content produced by a graphics processing unit.
[0032] Data compression can refer to the process of encoding, reconstructing, and / or otherwise modifying data to reduce its size. Data decompression can refer to the process of reassembling data compressed through the data compression process. Data compression / decompression can be used for various purposes, such as reducing the size of data transmitted between a first device and a second device, or reducing the size of data transmitted between a first component and a second component of a device, which can reduce bandwidth requirements and thus be associated with improved performance. Data compression / decompression can be based on a compression / decompression format, where the compression / decompression format can be associated with a bit width. In one example, the compression / decompression format can be associated with processing ten-bit data groups at a time; that is, the compression / decompression format can have a ten-bit bit width. In another example, the compression / decompression format can be associated with processing twelve-bit data groups at a time; that is, the compression / decompression format can have a twelve-bit bit width.
[0033] The device (or a component of the device) may include compression hardware and / or decompression hardware, wherein the compression hardware is configured to compress data, and wherein the decompression hardware is configured to decompress data. The compression / decompression hardware may be configured with a second bit width. The second bit width of the compression / decompression hardware may be the same as the first bit width of the compression / decompression format, or the second bit width of the compression / decompression hardware may be different from the first bit width of the compression / decompression format. When the compression / decompression hardware uses a compression / decompression format with a bit width identical to the bit width of the compression / decompression hardware to compress / decompress data, the compression / decompression of the data may be referred to as "native compression / decompression." When the compression / decompression hardware uses a compression / decompression format with a bit width different from the bit width of the compression / decompression hardware to compress / decompress data, the compression / decompression of the data may be referred to as "non-native compression / decompression."
[0034] In examples relative to non-native compression, the bit width of the compression / decompression format can be larger than the bit width of the compression / decompression hardware. In this example, the bit set (i.e., the data) can be split into a first subset and a second subset, and the first and second subsets can be processed independently by the compression / decompression hardware. For example, a first compression unit of the compression / decompression hardware can compress the first subset, and a second compression unit of the compression / decompression hardware can compress the second subset. However, splitting and compressing the bit set in the above manner can result in a relatively large prediction error for the bit set, which may lead to a reduced compression ratio. A reduced compression ratio may result in a relatively large amount of bandwidth being used for the compression / decompression process.
[0035] This document describes various techniques related to pixel preconditioning processing for non-native compression. In one example, an apparatus (e.g., an encoder) obtains a bit set comprising a first subset and a second subset, wherein the bit width of the bit set is greater than the configured bit width of the compression hardware. The apparatus (e.g., the encoder) processes the first subset based on the parity corresponding to the second subset. Parity may refer to whether the number represented by the bit set or subset is even or odd. The decoder output of the apparatus (e.g., the encoder) for the compression hardware includes the bit set comprising the processed first subset and second subset. Compared to processing the first subset based on the parity corresponding to the second subset, the apparatus may remove discontinuities associated with the first subset, which can improve the compression ratio when the bit set is compressed. The improved compression ratio may be associated with savings in computational and / or network resources. In another example, an apparatus (e.g., a decoder) obtains a bit set comprising a first subset and a second subset from the encoder of the compression hardware, wherein the bit width of the bit set is greater than the configured bit width of the compression hardware. The apparatus (e.g., a decoder) processes the first subset based on the parity corresponding to the second bit subset. The apparatus (e.g., the decoder) outputs a set of bits comprising the processed first and second bit subsets. Compared to processing the first subset based on the parity corresponding to the second bit subset, the apparatus can remove discontinuities associated with the first subset, which saves computational resources.
[0036] The examples described herein may relate to the use and functionality of a graphics processing unit (GPU). As used herein, a GPU can be any type of graphics processor, and a graphics processor can be any type of processor designed or configured to process graphical content. For example, a graphics processor or GPU can be a dedicated circuit designed to process graphical content. As an additional example, a graphics processor or GPU can be a general-purpose processor configured to process graphical content.
[0037] Figure 1This is a block diagram illustrating an example content generation system 100 configured to implement one or more technologies of this disclosure. The content generation system 100 includes a device 104. Device 104 may include one or more components or circuitry for performing the various functions described herein. In some examples, one or more components of device 104 may be components of a System-on-a-Chip (SOC). Device 104 may include one or more components configured to perform one or more technologies of this disclosure. In the illustrated example, device 104 may include a processing unit 120, a content encoder / decoder 122, and a system memory 124. In some aspects, device 104 may include multiple components (e.g., a communication interface 126, a transceiver 132, a receiver 128, a transmitter 130, a display processor 127, and one or more displays 131). Display 131 may refer to one or more displays 131. For example, display 131 may include a single display or multiple displays, which may include a first display and a second display. The first display may be a left-eye display, and the second display may be a right-eye display. In some examples, the first and second displays may receive different frames for presentation on the first and second displays. In other examples, the first and second displays may receive the same frames used for rendering on both displays. In yet another example, the results of graphics processing may not be displayed on the device; for example, the first and second displays may not receive any frames used for rendering on either display. Instead, the frames or graphics processing results may be transferred to another device. In some respects, this is referred to as split rendering.
[0038] Processing unit 120 may include internal memory 121. Processing unit 120 may be configured to perform graphics processing using graphics processing pipeline 107. Content encoder / decoder 122 may include internal memory 123. In some examples, device 104 may include a processor configured to perform one or more display processing techniques on one or more frames generated by processing unit 120, and then display those frames through one or more displays 131. Although the processor in example content generation system 100 is configured as display processor 127, it should be understood that display processor 127 is one example of a processor and other types of processors, controllers, etc., may be used instead of display processor 127. Display processor 127 may be configured to perform display processing. For example, display processor 127 may be configured to perform one or more display processing techniques on one or more frames generated by processing unit 120. One or more displays 131 may be configured to display or otherwise present the frames processed by display processor 127. In some examples, one or more displays 131 may include one or more of the following: liquid crystal display (LCD), plasma display, organic light-emitting diode (OLED) display, projection display device, augmented reality display device, virtual reality display device, head-mounted display, or any other type of display device.
[0039] Memory (such as system memory 124) external to processing unit 120 and content encoder / decoder 122 may be accessible to processing unit 120 and content encoder / decoder 122. For example, processing unit 120 and content encoder / decoder 122 may be configured to read from and / or write to external memory (such as system memory 124). Processing unit 120 may be communicatively coupled to system memory 124 via a bus. In some examples, processing unit 120 and content encoder / decoder 122 may be communicatively coupled to internal memory 121 via the bus or via a different connection.
[0040] Content encoder / decoder 122 can be configured to receive graphic content from any source, such as system memory 124 and / or communication interface 126. System memory 124 can be configured to store received encoded or decoded graphic content. Content encoder / decoder 122 can be configured to receive encoded or decoded graphic content from system memory 124 and / or communication interface 126, for example, in the form of encoded pixel data. Content encoder / decoder 122 can be configured to encode or decode any graphic content.
[0041] Internal memory 121 or system memory 124 may include one or more volatile or non-volatile memories or storage devices. In some examples, internal memory 121 or system memory 124 may include RAM, static random access memory (SRAM), dynamic random access memory (DRAM), erasable programmable ROM (EPROM), EEPROM, flash memory, magnetic data media or optical storage media, or any other type of memory. According to some examples, internal memory 121 or system memory 124 may be a non-transitory storage medium. The term "non-transitory" may indicate that the storage medium is not embodied in a carrier wave or propagating signal. However, the term "non-transitory" should not be construed as meaning that internal memory 121 or system memory 124 is not removable or that its contents are static. For example, system memory 124 may be removed from device 104 and moved to another device. Alternatively, system memory 124 may not be removable from device 104.
[0042] Processing unit 120 may be a CPU, GPU, GPGPU, or any other processing unit configured to perform graphics processing. In some examples, processing unit 120 may be integrated into the motherboard of device 104. In other examples, processing unit 120 may reside on a graphics card mounted in a port on the motherboard of device 104, or may otherwise be incorporated into a peripheral device configured to interoperate with device 104. Processing unit 120 may include one or more processors, such as one or more microprocessors, GPUs, ASICs, FPGAs, arithmetic logic units (ALUs), DSPs, discrete logic components, software, hardware, firmware, other equivalent integrated or discrete logic circuits, or any combination thereof. If the technology is partially implemented in software, processing unit 120 may store instructions for software in a suitable non-transitory computer-readable storage medium (e.g., internal memory 121) and may use one or more processors to execute instructions in hardware to perform the technology of this disclosure. Any of the foregoing (including hardware, software, combinations of hardware and software, etc.) may be considered as one or more processors.
[0043] The content encoder / decoder 122 can be any processing unit configured to perform content decoding. In some examples, the content encoder / decoder 122 may be integrated into the motherboard of device 104. The content encoder / decoder 122 may include one or more processors, such as one or more microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), arithmetic logic units (ALUs), digital signal processors (DSPs), video processors, discrete logic components, software, hardware, firmware, other equivalent integrated or discrete logic circuits, or any combination thereof. If the technology is partially implemented in software, the content encoder / decoder 122 may store instructions for software in a suitable non-transitory computer-readable storage medium (e.g., internal memory 123) and may use one or more processors to execute instructions in hardware to perform the technology of this disclosure. Any of the foregoing (including hardware, software, combinations of hardware and software, etc.) can be considered as one or more processors.
[0044] In some aspects, the content generation system 100 may include a communication interface 126. The communication interface 126 may include a receiver 128 and a transmitter 130. The receiver 128 may be configured to perform any of the receiving functions described herein with respect to device 104. Additionally, the receiver 128 may be configured to receive information from another device, such as eye or head positioning information, rendering commands, and / or location information. The transmitter 130 may be configured to perform any of the transmitting functions described herein with respect to device 104. For example, the transmitter 130 may be configured to transmit information to another device, which may include a request for content. The receiver 128 and the transmitter 130 may be combined to form a transceiver 132. In such an example, the transceiver 132 may be configured to perform any of the receiving and / or transmitting functions described herein with respect to device 104.
[0045] Refer again Figure 1In some aspects, processing unit 120 may include compressor 198, configured to: obtain a bit set including a first subset and a second subset, wherein the bit width of the bit set is greater than the configured bit width of the compression hardware; process the first subset based on parity corresponding to the second subset; and output a decoder for the compression hardware including the processed first and second subsets of bits. In some aspects, processing unit 120 may include decompressor 199, configured to obtain a bit set including a first subset and a second subset from the encoder of the compression hardware, wherein the bit width of the bit set is greater than the configured bit width of the compression hardware; process the first subset based on parity corresponding to the second subset; and output a bit set including the processed first and second subsets of bits. Although the following description may focus on data processing and graphics processing, the concepts described herein are applicable to other similar processing techniques. For example, the concepts presented herein are applicable to compressing and decompressing data types other than graphics-related data, such as pixel data. A pixel may refer to the smallest addressable element in an image. A frame may refer to an image in an image sequence. Furthermore, although the descriptions in this document may focus on bit widths of 10, 12, 14, or 16 bits, the concepts described herein may also be applied to other bit widths (e.g., 15, 18, etc.).
[0046] Devices such as device 104 can refer to any device, apparatus, or system configured to perform one or more of the technologies described herein. For example, a device can be a server, base station, user equipment, client device, station, access point, computer (such as a personal computer, desktop computer, laptop computer, tablet computer, computer workstation, or mainframe computer), end product, apparatus, telephone, smartphone, server, video game platform or console, handheld device (such as a portable video game device or personal digital assistant (PDA)), wearable computing device (such as a smartwatch, augmented reality device, or virtual reality device), non-wearable device, display or display device, television, set-top box, intermediate network device, digital media player, video streaming device, content streaming device, in-vehicle computer, any mobile device, any device configured to generate graphical content, or any device configured to perform one or more of the technologies described herein. The processes described herein may be described as being performed by a specific component (e.g., GPU), but in other embodiments, other components (e.g., CPU) consistent with the disclosed embodiments may be used to perform them.
[0047] A GPU can process various types of data or data packets within its pipeline. For example, in some aspects, a GPU can process two types of data or data packets, such as context register packets and draw call data. Context register packets can be a set of global state information, such as information about global registers, shaders, or constant data, which can adjust how the graphics context will be processed. For example, a context register packet may include information about the color format. In some aspects of a context register packet, there may be one or more bits indicating which workload belongs to the context register. Additionally, multiple functions or programs can run simultaneously and / or in parallel. For example, a function or program may describe an operation, such as a color mode or color format. Therefore, context registers can define various states of the GPU.
[0048] Context states can be used to determine how individual processing units (e.g., vertex extractors (VFDs), vertex shaders (VSs), shader processors, or geometry processors) operate and / or in which mode they operate. To do this, the GPU uses context registers and programming data. In some aspects, the GPU can generate workloads in the pipeline based on the context register definitions of modes or states, such as vertex or pixel workloads. Certain processing units (e.g., VFDs) can use these states to determine certain functions, such as how to aggregate vertices. Because these modes or states can change, the GPU may need to modify the corresponding context. Additionally, the workload corresponding to a mode or state may follow the changed mode or state.
[0049] Figure 2 Example GPU 200 is illustrated according to one or more technologies according to this disclosure. For example... Figure 2 As shown, GPU 200 includes a command processor (CP) 210, a draw call group 212, a VFD 220, a VS 222, a vertex cache (VPC) 224, a triangle setup engine (TSE) 226, a rasterizer (RAS) 228, a Z-process engine (ZPE) 230, a pixel interpolator (PI) 232, a fragment shader (FS) 234, a rendering backend (RB) 236, an L2 cache (UCHE) 238, and system memory 240. Although Figure 2 The GPU 200 includes processing units 220 to 238, but the GPU 200 may include multiple additional processing units. Additionally, processing units 220 to 238 are merely examples, and the GPU may use any combination or order of processing units in accordance with this disclosure. The GPU 200 also includes a command buffer 250, a context register group 260, and a context state 261.
[0050] like Figure 2As shown, the GPU can use a CP (e.g., CP 210) or a hardware accelerator to resolve the command buffer into context register groups (e.g., context register group 260) and / or draw call data groups (e.g., draw call group 212). Subsequently, CP 210 can transfer the context register group 260 or the draw call group 212 to a processing unit or block within the GPU via a separate path. Furthermore, the command buffer 250 can alternate between different states of the context registers and draw calls. For example, the command buffer can simultaneously store the following information: the context register of context N, the draw call of context N, the context register of context N+1, and the draw call of context N+1.
[0051] GPUs can render images in a variety of different ways. In some cases, GPUs can render images using direct rendering and / or tiled rendering. In a tiled rendering GPU, an image can be divided or separated into different parts or tiles. After the image is divided, each part or tile can be rendered individually. A tiled rendering GPU can divide a computer graphics image into a grid format, so that each part of the grid (i.e., a tile) is rendered individually. In some aspects of tiled rendering, the image can be divided into different bins or tiles during binning passes. In some aspects, a visibility stream can be constructed during binning passes, where visible primitives or draw calls can be identified. A rendering pass can be performed after a binning pass. In contrast to tiled rendering, direct rendering does not divide a frame into smaller bins or tiles. Instead, in direct rendering, the entire frame is rendered at once (i.e., without binning passes). Additionally, some types of GPUs allow both tiled rendering and direct rendering (e.g., flex rendering).
[0052] In some respects, a GPU can apply the drawing or rendering process to different bins or tiles. For example, a GPU can render a bin and perform all drawing for the primitives or pixels within that bin. During the bin-based rendering process, the rendering target can be located in GPU Internal Memory (GMEM). In some instances, after rendering a bin, the contents of the rendering target can be moved to system memory, and GMEM can be freed to render the next bin. Additionally, a GPU can render another bin and perform drawing for the primitives or pixels within that bin. Thus, in some respects, there may be a small number of bins covering all the drawing on a surface, for example, four bins. Furthermore, a GPU can loop through all the drawing in a bin but perform drawing only for visible drawing calls, i.e., drawing calls that include visible geometry. In some respects, a visibility stream can be generated, for example, in binning passes, to determine the visibility information of each primitive in an image or scene. For example, such a visibility stream can identify whether a primitive is visible. In some respects, this information can be used to remove invisible primitives, such that, for example, invisible primitives are not rendered in a rendering pass. Additionally, at least some primitives that are marked as visible can be rendered in the rendering pass.
[0053] In some aspects of tile rendering, there can be multiple processing stages or passes. For example, rendering can be performed in two passes, such as a binning, visibility, or box visibility pass and a rendering or box rendering pass. During a visibility pass, the GPU can input a rendering workload, record the positions of primitives or triangles, and then determine which primitives or triangles fall into which bins or regions. In some aspects of a visibility pass, the GPU can also identify or mark the visibility of each primitive or triangle in the visibility stream. During a rendering pass, the GPU can input a visibility stream and process one bin or region at a time. In some aspects, the visibility stream can be analyzed to determine which primitives or primitive vertices are visible or invisible. Thus, visible primitives or primitive vertices can be processed. By doing so, the GPU can reduce the unnecessary workload of processing or rendering invisible primitives or triangles.
[0054] In some aspects, certain types of primitive geometry, such as localized geometry, can be processed during visibility passes. Additionally, primitives can be categorized into different bins or regions based on their localization or position. In some instances, categorizing primitives or triangles into different bins can be performed by determining visibility information for those primitives or triangles. For example, the GPU can determine the visibility information for each primitive in each bin or region or write it to, for example, system memory. This visibility information can be used to determine or generate a visibility stream. In a rendering pass, the primitives in each bin can be rendered individually. In these cases, the visibility stream can be retrieved from memory and used to remove primitives that are not visible to that bin.
[0055] Some aspects of a GPU or GPU architecture can provide multiple different options for rendering (e.g., software rendering and hardware rendering). In software rendering, the driver or CPU can process each view... Figure 1 The entire frame geometry is copied each time. Additionally, some different states can change depending on the viewpoint. Therefore, in software rendering, the software can copy the entire workload by changing some states that can be used for rendering for each viewpoint in the image. In some respects, this can lead to increased overhead because the GPU may submit the same workload multiple times for each viewpoint in the image. In hardware rendering, the hardware or GPU may be responsible for copying or processing the geometry for each viewpoint in the image. Therefore, the hardware can manage the copying or processing of primitives or triangles for each viewpoint in the image.
[0056] Figure 3 An image or surface 300 according to one or more techniques of this disclosure is illustrated, including multiple elements divided into multiple boxes. For example... Figure 3 As shown, the image or surface 300 includes a region 302, which includes primitives 321, 322, 323, and 324. Primitives 321, 322, 323, and 324 are divided or placed into different bins, such as bins 310, 311, 312, 313, 314, and 315. Figure 3 This example illustrates tile rendering using multiple viewpoints for primitives 321-324. For instance, primitives 321-324 are in a first viewpoint 350 and a second viewpoint 351. Therefore, GPU processing or rendering of an image or surface 300 including region 302 can utilize multi-view or multi-view rendering.
[0057] As indicated in this article, GPUs or graphics processors can use tile rendering architectures to reduce power consumption or save memory bandwidth. As further stated above, this rendering method divides the scene into multiple bins, along with visibility paths that identify the visible triangles within each bin. Therefore, in tile rendering, the entire screen can be divided into multiple bins or tiles. The scene can then be rendered multiple times, for example, once or multiple times for each bin.
[0058] In various aspects of graphics rendering, some graphics applications may render a single target (i.e., the rendering target) once or multiple times. For example, in graphics rendering, the frame buffer on system memory can be updated multiple times. The frame buffer can be part of memory or random access memory (RAM) (e.g., containing bitmaps or storage devices) to help store display data for the GPU. The frame buffer can also be a memory buffer containing a complete frame of data. Additionally, the frame buffer can be a logical buffer. In some aspects, updating the frame buffer can be performed in bin or tile rendering, where, as discussed above, the surface is divided into multiple bins or tiles, and each bin or tile can then be rendered individually. Furthermore, in tile rendering, the frame buffer can be divided into multiple bins or tiles.
[0059] As this article points out, in some respects, such as in boxed or tiled rendering architectures, frame buffers allow data to be repeatedly stored or written to them, for example, when rendering from different types of memory. This can be referred to as unresolving the frame buffers or system memory. For example, when storing or writing to one frame buffer and then switching to another, the data or information on the frame buffer can be resolved from the GMEM at the GPU to system memory, i.e., memory in dual data rate (DDR) RAM or dynamic RAM (DRAM).
[0060] In some respects, system memory can also be system-on-chip (SoC) memory or another chip-based memory, such as on a device or smartphone, used for storing data or information. System memory can also be a physical data storage device shared by the CPU and / or GPU. In some respects, system memory can be, for example, a DRAM chip on a device or smartphone. Therefore, SoC memory can be a chip-based method for storing data.
[0061] In some respects, GMEM can be on-chip memory at the GPU, which can be implemented using static RAM (SRAM). Alternatively, GMEM can be stored on the device (e.g., a smartphone). As indicated herein, data or information can be transferred between system memory or DRAM and GMEM, for example, at the device. In some respects, system memory or DRAM can reside at the CPU or GPU. Furthermore, data can be stored in DDR or DRAM. In some respects, such as in bin or tiled rendering, a small portion of the memory can be stored at the GPU, for example, in GMEM. In some cases, storing data at GMEM may utilize a larger processing workload and / or consume more power compared to storing data at the frame buffer or system memory.
[0062] Figure 4Figure 400 illustrates an example of native compression 402 according to one or more techniques of this disclosure. Native compression 402 may also be referred to as "native encoding". In one example, native compression 402 may be performed by device 104. Universal Bandwidth Compression (UBWC) may refer to a compression / decompression standard that can be used to compress / decompress graphics buffers (i.e., graphics data) such as graphics processor buffers, video buffers, and / or camera buffers. UBWC can operate on a per-tile basis, where each tile may include a defined number of pixels (e.g., four pixels, eight pixels, etc.). UBWC can be used to compress / decompress graphics buffers transmitted within a device (e.g., from one hardware block of the device to another hardware block of the device). UBWC can also be used to compress / decompress graphics buffers transmitted between devices (e.g., from a first device to a second device). The hardware of a device may be configured to perform compression / decompression via UBWC. With UBWC, as the bit width of the bit components (i.e., bit sets) increases, the hardware complexity (e.g., timing, area) may increase. A bit can refer to a logical state that has one of two possible values (zero or one). Bit width can refer to the number of bits in the bit set.
[0063] UBWC can include different variants. Some example variants of UBWC that operate on 10-bit components include UBWC_RGBA8888, UBWC_RGBA1010102, UBWC_NV12-Y, UBWC_NV-12-UV, UBWC_TP10-Y, UBWC_TP10-UV, UBWC_NV124R-Y, and UBWC_NV124R-UV. Some example variants of UBWC that operate on more than 10-bit components include UBWC_P016 (16-bit), UBWC_TBAYER12 (12-bit), and UBWC_TBAYER14 (14-bit).
[0064] Compression / decompression hardware can be configured to operate on data based on the first bit (i.e., the first bit width), meaning it can be configured to process data based on the number of bits. For example, if you want to compress one hundred bits and the first bit is ten, the first compression unit of the compression / decompression hardware can compress bits zero through nine of the data, the second compression unit can compress bits ten through nineteen, and so on.
[0065] A compression / decompression format (e.g., UBWC) can be associated with processing a second bit of data (i.e., a second bit width). For example, if you want to compress one hundred bits and the second bit of the compression / decompression format is ten, then the compression / decompression format can instruct that the one hundred bits be processed (e.g., compressed) in groups of ten.
[0066] The first digit can be the same as the second digit, or the first digit can be different from the second digit. When the first digit equals the second digit, compression can be called "native compression," and decompression can be called "native decompression." When the first digit is not equal to the second digit (i.e., the first digit is different from the second digit), compression can be called "non-native compression," and decompression can be called "non-native decompression."
[0067] In an example where the second bit of the compression / decompression format is greater than the first bit of the compression / decompression hardware, the device can split the bit component of the data to be processed into a first bit component and a second bit component, where the sum of the bits in the first bit component and the bits in the second bit component equals the number of bits in the bit component. The device's first compression unit processes the first bit component, and the device's second compression unit processes the second bit component. The aspects presented herein relate to improving the compression ratio of non-native compression. Compression ratio can refer to the relative reduction in the size of the data representation produced by a data compression algorithm. Compression ratio can be expressed as the quotient of the uncompressed size of the data and the compressed size of the data.
[0068] In the example relative to native compression 402, a device (e.g., device 104) can acquire image data (e.g., frames), and the device can divide the image into multiple input tiles, wherein the multiple input tiles may include input tile 404. In one example, input tile 404 may be a rectangular subdivision of image data (e.g., a rectangular subdivision of a frame).
[0069] The device's component extractor 406 can extract a zero-bit component 408, a first-bit component 410, and an n-th bit component 412 from the input patch 404, where n is a positive integer greater than or equal to 2. In one example, each of the zero-bit component 408, the first-bit component 410, and the n-th bit component 412 can have a bit width of ten. In one example, the zero-bit component 408 can correspond to the red (R) value of a pixel, the first-bit component 410 can correspond to the green (G) value of a pixel, and the n-th bit component 412 can correspond to the blue (B) value of a pixel.
[0070] The zeroth compression unit 414 of the device can process the zeroth bit component 408, the first compression unit 416 of the device can process the first bit component 410, and the (n-1)th compression unit 418 can process the nth bit component 412. For example, the zeroth compression unit 414 of the device can compress the zeroth bit component 408, the first compression unit 416 of the device can compress the first bit component 410, and the (n-1)th compression unit 418 can compress the nth bit component 412. The sum of the bits of the (compressed) zeroth bit component 408, the (compressed) first bit component 410, and the (compressed) nth bit component 412 may be less than the number of bits of the input block 404. The (compressed) zeroth bit component 408 may be referred to as and / or may include the zeroth index 420, the (compressed) first bit component 410 may be referred to as and / or may include the first index 422, and the (compressed) nth bit component 412 may be referred to as and / or may include the nth index 424.
[0071] The device's packer 426 can pack the zero index 420, the first index 422, and the nth index 424 into a bit stream 428. The bit stream 428 can be sent to another hardware component within the device, or the bit stream can be sent to another device.
[0072] Figure 5 Figure 500 illustrates an example of non-native compression 502 according to one or more techniques of this disclosure. Non-native compression 502 may also be referred to as "native decoding". In one example, non-native compression 502 may be performed by device 104. As mentioned above, for compression / decompression of bit components with higher bit widths, the hardware complexity in terms of timing and area (i.e., the complexity of the compression / decompression hardware) may increase; that is, the hardware complexity may increase as the number of bits in the bit component increases. One way to compress bit components with a bit width higher than the configured bit width of the compression / decompression hardware is to split the bit component into a first bit component and a second bit component, and the compression / decompression hardware can independently process the first bit component and the second bit component using different instances of the same compression hardware (e.g., using a first compression unit and a second compression unit).
[0073] In the example relative to non-native compression 502, the device's compression / decompression hardware can be configured to process M-bit components (i.e., M-bit pixel components) for compression / decompression purposes, where M is a positive integer. In this example, a compression / decompression format for processing N-bit components can be defined, where N is a positive integer and where N > M. The N-bit component can be compressed by the same compression / decompression hardware (i.e., different instances of the same compression / decompression hardware) by splitting the N-bit component into j bits including the most significant bit (MSB) of the N-bit component and k bits including the least significant bit (LSB) of the N-bit component, where j and k are positive integers and where j ≤ M and k ≤ M. In one example, N can be sixteen, M can be ten, k can be 8, and j can be 8. The MSB can refer to the bit in the bit set that represents the position of the most significant bit of a binary integer. The LSB can refer to the bit in the bit set that represents the position of the least significant bit of a binary integer. In one example, for the set "1001", the MSB can be the leftmost "1" bit and the MSB can be the rightmost "1" bit.
[0074] For example, in an example relative to non-native compression 502, a device (e.g., device 104) may acquire image data (e.g., frames), and the device may divide the image into multiple input tiles, wherein the multiple input tiles may include input tile 504. In one example, input tile 504 may be a rectangular subdivision of image data (e.g., a rectangular subdivision of a frame).
[0075] The component extractor 506 of the device can extract the zeroth bit component 508A and the zeroth bit component 508B from the input block 504. The zeroth bit component 508A can correspond to the bits [N-1:k] of the zeroth bit component. Figure 5 (Not illustrated in the example), that is, the zeroth bit component 508A may include j MSBs of the zeroth bit component. In one example, N may be 16, k may be 8, and j may be 8, and therefore, the zeroth bit component 508A may correspond to bits [15:8] of the zeroth bit component, and the zeroth bit component 508A may include 8 MSBs of the zeroth bit component. The zeroth bit component 508B may correspond to bits [k-1:0] of the zeroth bit component, that is, the zeroth bit component 508B may include k LSBs of the zeroth bit component. According to the example above, the zeroth bit component 508B may correspond to bits [7:0] of the zeroth bit component, and the zeroth bit component 508B may include 8 LSBs of the zeroth bit component. In one example, the combination of the zeroth bit component 508A and the zeroth bit component 508B may correspond to the red (R) value of the pixel.
[0076] The component extractor 506 of the device can extract the (n-1)th bit component 510A and the (n-1)th bit component 510B from the input block 504. The (n-1)th bit component 510A can correspond to the bits [N-1:k] of the (n-1)th bit component. Figure 5 (Not illustrated in the example), that is, the (n-1)th bit component 510A may include j MSBs of the (n-1)th bit component. In one example, N may be 16, k may be 8, and j may be 8, and therefore, the (n-1)th bit component 510A may correspond to bits [15:8] of the (n-1)th bit component, and the (n-1)th bit component 510A may include 8 MSBs of the (n-1)th bit component. The (n-1)th bit component 510B may correspond to bits [k-1:0] of the (n-1)th bit component, that is, the (n-1)th bit component 510B may include k LSBs of the (n-1)th bit component. According to the example above, the (n-1)th bit component 510B may correspond to bits [7:0] of the (n-1)th bit component, and the (n-1)th bit component 510B may include 8 LSBs of the (n-1)th bit component. In one example, the combination of the (n-1)th bit component 510A and the (n-1)th bit component 510B may correspond to the green (G) value of the pixel.
[0077] The zeroth compression unit 512 of the device can process the zeroth bit component 508A, the first compression unit 514 of the device can process the zeroth bit component 508B, the (2n-2)th compression unit 516 of the device can process the (n-1)th bit component 510A, and the (2n-1)th compression unit 518 of the device can process the (n-1)th bit component 510B. For example, the zeroth compression unit 512 of the device can compress the zeroth bit component 508A, the first compression unit 514 of the device can compress the zeroth bit component 508B, the (2n-2)th compression unit 516 of the device can compress the (n-1)th bit component 510A, and the (2n-1)th compression unit 518 of the device can compress the (n-1)th bit component 510B. The sum of the bits of the compressed zero-th bit component 508A, the compressed zero-th bit component 508B, the compressed (n-1)-th bit component 510A, and the compressed (n-1)-th bit component 510B may be less than the number of bits in the input block 404. The compressed zero-th bit component 508A may be referred to as and / or may include the zero index 520A corresponding to the MSB of the zero-th bit component, the compressed zero-th bit component 508B may be referred to as and / or may include the zero index 520B corresponding to the LSB of the zero-th bit component, the compressed (n-1)-th bit component 510A may be referred to as and / or may include the (n-1)-th index 522A corresponding to the MSB of the (n-1)-th bit component, and the compressed (n-1)-th bit component 510B may be referred to as and / or may include the (n-1)-th index 522B corresponding to the LSB of the (n-1)-th bit component.
[0078] The packer 524 of the device can pack the zero index 520A, the zero index 520B, the (n-1)th index 522A, and the (n-1)th index 522B into a bit stream 526. The bit stream 526 can be sent to another hardware component within the device, or the bit stream can be sent to another device.
[0079] Figure 6 Figure 600 illustrates an example of native decompression 602 according to one or more techniques of this disclosure. Native decompression 602 may also be referred to as "native decoding". In one example, native decompression 602 may be performed by device 104.
[0080] The device can obtain input bit stream 604. In one example, the device can receive input bit stream 604 from another device. In another example, a hardware component of the device can receive input bit stream 604 from another hardware component of the device. In one example, input bit stream 604 can be or includes bit stream 428.
[0081] The bitstream extractor 606 of the device can extract the zero-th index 420, the first index 422, and the nth index 424 from the input bitstream 604. The zero-th pixel recovery unit 608 can process the zero-th index 420, the first pixel recovery unit 610 can process the first index 422, and the (n-1)th pixel recovery unit 612 can process the nth index 424. For example, the zero-th pixel recovery unit 608 can decompress the zero-th index 420, the first pixel recovery unit 610 can decompress the first index 422, and the (n-1)th pixel recovery unit 612 can decompress the nth index 424. The aforementioned decompression can reproduce the zero-th bit component 408, the first bit component 410, and the nth bit component 412.
[0082] The device's tile packer 614 can pack the zeroth component 408, the first component 410, and the nth component 412 into a tile 616. In one example, tile 616 may be, include, or correspond to input tile 404. In one example, tile 616 may be a portion of an image that can be displayed on a monitor.
[0083] Figure 7 Figure 700 illustrates an example of non-native decompression 702 according to one or more techniques of this disclosure. Non-native decompression 702 may also be referred to as "native decoding". In one example, non-native decompression 702 may be performed by device 104.
[0084] The device can obtain input bit stream 704. In one example, the device can receive input bit stream 704 from another device. In another example, a hardware component of the device can receive input bit stream 704 from another hardware component of the device. In one example, input bit stream 704 can be or includes bit stream 526.
[0085] Bitstream extractor 706 can extract zero-index 520A, zero-index 520B, (n-1)-index 522A, and (n-1)-index 522B from input bitstream 704. Zero-pixel recovery unit 708 can process zero-index 520A, first pixel recovery unit 710 can process zero-index 520B, (2n-2)-pixel recovery unit 712 can process (n-1)-index 522A, and (2n-1)-pixel recovery unit 714 can process (n-1)-index 522B. For example, zero-pixel recovery unit 708 can decompress zero-index 520A, first pixel recovery unit 710 can decompress zero-index 520B, (2n-2)-pixel recovery unit 712 can decompress (n-1)-index 522A, and (2n-1)-pixel recovery unit 714 can decompress (n-1)-index 522B. The aforementioned decompression can reproduce the zeroth bit component 508A, the zeroth bit component 508B, the (n-1)th bit component 510A, and the (n-1)th bit component 510B.
[0086] Tile packer 716 can pack the zero-bit component 508A, the zero-bit component 508B, the (n-1)th bit component 510A, and the (n-1)th bit component 510B into tile 718. In one example, tile 718 may be, include, or correspond to input tile 504. In one example, tile 718 may be a portion of an image that can be displayed on a monitor.
[0087] Figure 8 This is a diagram 800 illustrating a first plot 802 of the most significant bit (MSB) and native bit component values and a second plot 804 of the least significant bit (LSB) and native bit component values according to one or more techniques of this disclosure. The first plot 802 and the second plot 804 may correspond to the non-native compression 502 and non-native decompression 702 described above.
[0088] When bit components (i.e., pixel bit components) are predicted non-natively (i.e., by splitting bit components into individual bit components and compressing each individual bit component independently), the prediction error associated with compression / decompression can be relatively high due to the relatively large discontinuities in the LSB values. Discontinuities can refer to breaks, jumps, or gaps in the plotting of the LSB versus the native bit component values, or breaks, jumps, or gaps in the plotting of the MSB versus the native bit component values. This can lead to a reduced compression ratio. For example, first plot 802 illustrates a first discontinuity 806 (illustrated by dashed lines) associated with MSB values, and second plot 804 illustrates a second discontinuity 808 (illustrated by dashed lines) associated with LSB values. As illustrated in Figure 800, the second discontinuity 808 can be greater than the first discontinuity 806.
[0089] In some respects, as the hardware complexity (e.g., timing, area, etc.) of compressing higher bit-width components increases, one approach to compressing higher bit-width pixel components is to split the components and compress them independently using the same hardware (e.g., with a smaller bit width). However, this can lead to very high prediction errors in certain ranges. The aspects presented in this paper can remove discontinuities in LSB components using preconditioning and LSB folding techniques on the split LSB components. In some cases, the MSB can remain unchanged.
[0090] Figure 9 Figure 900 illustrates an example aspect of pixel preconditioning processing according to one or more techniques of this disclosure. As discussed above, non-native compression / decompression (e.g., non-native compression 502, non-native decompression 702) can result in relatively large discontinuities in LSB values, which may reduce the compression ratio. To address this problem, the aspects presented herein can remove discontinuities relative to LSB bit components (e.g., the zeroth bit component 508B, the (n-1)th bit component 510B, etc.). Such techniques for removing discontinuities may be referred to as "preconditioning processing," "pixel preconditioning processing," or "LSB folding." A device (e.g., device 104) can remove discontinuities relative to LSB bit components according to the following equation (I).
[0091] (I)
[0092] In equation (I) above, "bit component `[k-1:0]" can refer to the LSB bit component after removing the discontinuities associated with the LSB, "bit component [k-1:0]" can refer to the LSB bit component of the bit component (e.g., the zeroth bit component 508B, the (n-1)th bit component 510B), "k" can be a positive integer as mentioned above, "N" can be a positive integer as mentioned above, "~" can refer to a bitwise NOT operation, and "bit component [N-1:k]" can refer to the MSB bit of the bit component (e.g., the zeroth bit component 508A, the (n-1)th bit component 510A). A bitwise NOT operation means switching "1" bits in a bit sequence to "0" bits and switching "0" bits in a bit sequence to "1" bits. In one example, performing a bitwise NOT operation on the bit sequence "0011" produces the resulting bit sequence "1100".
[0093] In one example, the compression / decompression format may have a bit width of 16, such as UBWC_P016. In this example, non-native compression / decompression can be performed by dividing the bit width of 16 by 2, resulting in 8-bit LSB and 8-bit MSB components, each of which can be compressed / decompressed independently. In this example, the prediction error may be relatively high for some pixels with small value differences when adjacent pixel values are approximately multiples of 256 (e.g., when some pixel values are higher than but relatively close to multiples of 256, and when other pixel values are lower than but relatively close to multiples of 256). In this example, "k" can be 7 and "N" can be 16, resulting in the following equation (II).
[0094] (II)
[0095] In other words, and as reflected in Equation (I) above, during the non-native compression process, the device hardware can extract bit components from the input tile. Since the compression process is non-native, the device hardware can further decompose the bit components into a first bit component (e.g., the zeroth bit component 508B) and a second bit component (the zeroth bit component 508A), where the first bit component may include the LSB of the bit component, and where the second bit component may include the MSB of the bit component. The device can determine the parity of the second bit component, i.e., whether the second bit component represents / corresponds to an even or odd number. When the second bit component is odd, the device can perform a bitwise NOT operation on the first bit component, as reflected in Equation (I). When the second bit component is even, the device can retain the first bit component as the first bit component. In either case, the device can retain the second bit component as the second bit component. Preconditioning the first bit component according to Equation (I) minimizes the prediction error when the actual prediction factor is relatively close to the actual pixel value. The device can then compress the first and second components as described above (e.g., via the zero compression unit 512 and the first compression unit 514).
[0096] In one aspect, a device (e.g., device 104) may perform pixel preconditioning processing according to the following equation (III).
[0097] (III)
[0098] During a non-native decompression process (e.g., non-native decompression 702), after the first and second components have been recovered (e.g., via pixel recovery units such as zero-pixel recovery unit 708 and first pixel recovery unit 710), the device hardware can again apply the above equation (I) to the first and second components.
[0099] Figure 900 depicts a second plot 804 of the LSB versus the original bit component values described above. Figure 900 also depicts a plot 902 of the new LSB versus the original bit component values. Plot 902 corresponds to the application of equation (I) above. As illustrated in plot 902, the second discontinuity 808 is removed by applying preconditioning to the LSB. Therefore, preconditioning improves the compression ratio, which allows for more efficient use of the device's computational and power resources.
[0100] Figure 10 Figure 1000 illustrates an example aspect of encoding 1002 related to one or more techniques according to this disclosure. Figure 1000 illustrates modification of non-native compression 502 using preconditioning processing 1004. More specifically, the device may apply preconditioning processing 1004 to the zero-bit component 508B and the (n-1)-th bit component 510B, that is, the device may apply equation (I) to the zero-bit component 508B and the (n-1)-th bit component 510B. Subsequently, the zero-bit compression unit 512 may compress the zero-bit component 508A, the first compression unit 514 may compress the (preconditioned) zero-bit component 508B, the (2n-2)-th compression unit 516 may compress the (n-1)-th bit component 510A, and the (2n-1)-th compression unit 518 may compress the (preconditioned) n-1-th bit component 510B.
[0101] Figure 11 Figure 1100 illustrates an example aspect of decoding 1102 according to one or more techniques of this disclosure. Figure 1100 illustrates modification of non-native decompression 702 using inverse conditioning 1104. More specifically, the device may apply inverse conditioning to the zeroth bit component 508B (after preconditioning) and the (n-1)th bit component 510B (after preconditioning), that is, the device may apply equation (I) to the zeroth bit component 508B (after preconditioning) and the (n-1)th bit component 510B (after preconditioning). Subsequently, the tile packer 716 can pack the zero-th component 508A, the (inverse conditionalized) zero-th component 508B (i.e., the zero-th component 508B), the (n-1)th component 510A, and the (inverse conditionalized) (n-1)th component 510B (i.e., the (n-1)th component 510B) into the tile 718.
[0102] As described above, preconditioning / LSB folding can improve the compression ratio of non-natively compressed tiles with a small increase in area in the compression / decompression hardware. In one example, when preconditioning / LSB folding is applied to the P016 tile (16-bit component), an improvement in compression ratio (CR) can be observed compared to non-native compression methods without preconditioning / LSB folding. Tables 1 and 2 below detail the aspects related to the improvement in CR provided by preconditioning / LSB folding.
[0103]
[0104] Table 1: Improvements to CR Preconditioning
[0105]
[0106] Table 2: Differences in CR Improvement After Preconditioning
[0107] As illustrated in Tables 1 and 2 above, a CR improvement of approximately 0.93% can be observed for bit components split into 8 MSBs and 8 LSBs and subjected to the preconditioning process described above. For bit components split into 6 MSBs and 10 LSBs, a CR improvement of approximately 2.3% can be observed. Preconditioning minimizes the area increase in compression / decompression hardware (e.g., approximately 0.3%).
[0108] Figure 12 This is a call flowchart 1200 illustrating example communication between encoder 1202 and decoder 1204 according to one or more techniques of this disclosure. Encoder 1202 may be implemented in hardware and / or software. Decoder 1204 may be implemented in hardware and / or software. In this example, encoder 1202 and / or decoder 1204 may be included in device 104. In another example, encoder 1202 and / or decoder 1204 may be included in different devices.
[0109] At 1206, encoder 1202 may obtain a bit set including a first subset and a second subset, wherein the bit width of the bit set is greater than the configuration bit width of the compression hardware. At 1208, encoder 1202 may process the first subset based on the parity corresponding to the second subset. At 1212, encoder 1202 may output a bit set including the processed first subset and second subset to the decoder of the compression hardware (e.g., decoder 1204). For example, at 1212A, encoder 1202 may send a bit set including the processed first subset and second subset to the decoder. At 1210, encoder 1202 may compress the bit set including the processed first subset and second subset, wherein outputting the bit set including the processed first subset and second subset at 1212 may include: outputting the compressed bit set to the decoder of the compression hardware (e.g., decoder 1204).
[0110] At 1214, decoder 1204 may obtain a set of bits including a first subset and a second subset from the encoder (e.g., encoder 1202) of the compression hardware, wherein the bit width of the bit set is greater than the configured bit width of the compression hardware. At 1218, decoder 1204 may process the first subset based on the parity corresponding to the second subset. At 1220, decoder 1204 may output a set of bits including the processed first subset and the second subset. In one aspect, obtaining the set of bits including the processed first subset and the second subset may include: at 1612A, receiving a compressed set of bits including the processed first subset and the second subset, and at 1216, decoder 1204 may decompress the compressed set of bits including the processed first subset and the second subset, wherein processing the first subset based on the parity corresponding to the second subset at 1218 may include: processing the decompressed first subset based on the parity corresponding to the decompressed second subset.
[0111] Figure 13 This is a flowchart 1300 of an example method for graphics processing according to one or more techniques of this disclosure. The method can be performed by devices such as: a graphics processing apparatus, a GPU, a CPU, device 104, an encoder 1202, a wireless communication device, etc., in combination with... Figures 1 to 12 This method is used in various aspects. It can be associated with various advantages at the device level, such as improved compression ratios and reduced computational resources. In one example, the method can be performed by compressor 198.
[0112] At 1302, the device (e.g., an encoder) obtains a set of bits including a first subset and a second subset, wherein the bit width of the bit set is greater than the configured bit width of the compression hardware. For example, Figure 12At 1206, encoder 1202 is shown to obtain a bit set including a first subset and a second subset, wherein the bit width of the bit set may be greater than the configured bit width of the compression hardware. In one example, the bit set may include zero-th bit component 508A, zero-th bit component 508B, (n-1)-th bit component 510A, and (n-1)-th bit component 510B. In one example, the first subset may correspond to the zero-th bit component 508B, and the second bit subset may correspond to the zero-th bit component 508A. In one example, the compression hardware may be or include a zero-th compression unit 512 or a first compression unit 514. In one example, 1302 may be performed by compressor 198.
[0113] At 1304, the device (e.g., an encoder) processes the first subset based on the parity corresponding to the second subset. For example, Figure 12 At 1208, encoder 1202 is shown to process the first subset based on the parity corresponding to the second subset. In one example, processing the first subset based on the parity corresponding to the second subset may correspond to preconditioning process 1004. In another example, processing the first subset based on the parity corresponding to the second subset may correspond to equation (I) or equation (II) above. In one example, 1304 may be performed by compressor 198.
[0114] At 1306, the decoder output of the device (e.g., an encoder) for compression hardware includes a set of bits from the processed first and second bit subsets. For example, Figure 12 At 1212, encoder 1202 is shown to output a set of bits, including the processed first and second bit subsets, to a decoder (e.g., decoder 1204) of compression hardware. In one example, the output bit set may correspond to bit stream 526. In one example, 1306 may be performed by compressor 198.
[0115] Figure 14 This is a flowchart 1400 of an example method for graphics processing according to one or more techniques of this disclosure. The method can be performed by devices such as: a graphics processing apparatus, a GPU, a CPU, device 104, an encoder 1202, a wireless communication device, etc., in combination with... Figures 1 to 12 This method is used in various aspects. It can be associated with various advantages at the device level, such as improved compression ratios and reduced computational resources. In one example, this method (including the various aspects detailed below) can be performed by compressor 198.
[0116] At 1402, the device (e.g., an encoder) obtains a set of bits including a first subset and a second subset, wherein the bit width of the bit set is greater than the configured bit width of the compression hardware. In one example, the bit set may include a zeroth bit component 508A, a zeroth bit component 508B, an (n-1)th bit component 510A, and an (n-1)th bit component 510B. In one example, the first subset may correspond to the zeroth bit component 508B, and the second bit subset may correspond to the zeroth bit component 508A. In one example, the compression hardware may be or include a zeroth compression unit 512 or a first compression unit 514. In one example, 1402 may be performed by a compressor 198.
[0117] At 1404, the device (e.g., an encoder) processes the first subset based on the parity corresponding to the second subset. For example, Figure 12 At 1208, encoder 1202 is shown to process the first subset based on the parity corresponding to the second subset. In one example, processing the first subset based on the parity corresponding to the second subset may correspond to preconditioning process 1004. In another example, processing the first subset based on the parity corresponding to the second subset may correspond to equation (I) or equation (II) above. In one example, 1404 may be performed by compressor 198.
[0118] At 1408, the decoder output of the device (e.g., an encoder) for compression hardware includes a set of bits from the processed first and second bit subsets. For example, Figure 12 At 1212, encoder 1202 is shown to output a set of bits, including the processed first and second bit subsets, to a decoder (e.g., decoder 1204) of compression hardware. In one example, the output bit set may correspond to bit stream 526. In one example, 1408 may be performed by compressor 198.
[0119] In one aspect, the parity corresponding to the second subset can indicate an odd number, and processing the first subset can include performing a bitwise NOT operation on the first subset. For example, the above aspect can correspond to equation (I) in... Or in equation (II) In one example, processing the first subset at 1208 could include performing a bitwise NOT operation on the first subset.
[0120] In one aspect, the parity corresponding to the second subset can indicate an even number, and processing the first subset can include avoiding adjustments to the first subset. For example, the above aspect can correspond to equation (I) in... Or in equation (II) In one example, processing the first subset at 1208 could include avoiding adjusting the first subset.
[0121] In one aspect, the first subset may correspond to the least significant bit (LSB) component of the bit set, and the second subset may correspond to the most significant bit (MSB) component of the bit set. For example, Figure 10 It is shown that the zeroth bit component 508B can correspond to the LSB component of the bit set, and the zeroth bit component 508A can correspond to the MSB component of the bit set.
[0122] In one aspect, processing the first subset can remove discontinuities associated with the LSB component. For example, processing the first subset can remove discontinuities associated with the LSB component at 1208. In one example, the discontinuity can be or include a second discontinuity 808.
[0123] In one respect, the bit set can correspond to the set of pixels in a frame. For example, the bit set obtained at 1206 can correspond to the set of pixels in a frame.
[0124] In one aspect, the output, comprising a set of bits including the processed first and second subsets, may include sending the set of bits including the processed first and second subsets to a decoder of the compression hardware. For example, Figure 12 At 1612A, it is shown that the output, which includes the processed first and second bit subsets, may include sending the processed first and second bit subsets to a decoder of the compression hardware (e.g., decoder 1204).
[0125] In one aspect, at 1406, the device (e.g., an encoder) can compress a set of bits including the processed first and second bit subsets, wherein outputting a set of bits including the processed first and second bit subsets may include: outputting the compressed set of bits to a decoder of the compression hardware. For example, Figure 12 At 1210, encoder 1202 is shown to compress a set of bits including the processed first and second bit subsets, wherein outputting the set of bits including the processed first and second bit subsets at 1212 may include: outputting the compressed set of bits to the decoder of the compression hardware. In one example, the compressed set of bits may be... Figure 10 The zeroth compression unit 512 is associated with the first compression unit 514. In one example, 1406 may be performed by compressor 198.
[0126] In one respect, the bit width of the bit set can be 10 bits, 12 bits, 14 bits, or 16 bits. For example, the bit width of the bit set obtained at 1206 can be 10 bits, 12 bits, 14 bits, or 16 bits.
[0127] Figure 15This is a flowchart 1500 of an example method for graphics processing according to one or more techniques of this disclosure. The method can be performed by devices such as: a graphics processing apparatus, a GPU, a CPU, device 104, a decoder 1204, a wireless communication device, etc., in combination with... Figures 1 to 12 This method is used in various aspects. It can be associated with various advantages at the device level, such as improved compression ratios and reduced computational resources. In one example, the method can be performed by decompressor 199.
[0128] At 1502, the device (e.g., a decoder) obtains a set of bits from the encoder of the compression hardware, including a first subset and a second subset, wherein the bit width of the bit set is greater than the configured bit width of the compression hardware. For example, Figure 12 At 1214, it is shown that decoder 1204 can obtain a set of bits, including a first subset and a second subset, from the encoder (e.g., encoder 1202) of compression hardware, wherein the bit width of the bit set is greater than the configured bit width of the compression hardware. In one example, the bit set may correspond to zero index 520A, zero index 520B, (n-1)th index 522A, or (n-1)th index 522B. In one example, the first subset may correspond to zero index 520B, and the second subset may correspond to zero index 520A. In one example, the compression hardware may be or include zero pixel recovery unit 708 and / or first pixel recovery unit 710. In one example, 1502 may be performed by decompressor 199.
[0129] At 1504, the device (e.g., the decoder) processes the first subset based on the parity corresponding to the second subset. For example, Figure 12 At 1218, it is shown that decoder 1204 can process the first subset based on the parity corresponding to the second subset. In one example, processing the first subset based on the parity corresponding to the second subset corresponds to inverse conditionalization process 1104. In another example, processing the first subset based on the parity corresponding to the second subset corresponds to equation (I) or equation (II) above. In one example, processing the first subset based on the parity corresponding to the second subset can generate zeroth component 508A and zeroth component 508B. In one example, 1504 can be performed by decompressor 199.
[0130] At position 1506, the device (e.g., a decoder) outputs a set of bits that includes the processed first and second bit subsets. For example, Figure 12 At 1220, it is shown that decoder 1204 can output a set of bits including the processed first and second bit subsets. In one example, the output bit set may correspond to block 718. In one example, 1506 may be performed by decompressor 199.
[0131] Figure 16 This is a flowchart 1600 of an example method for graphics processing according to one or more techniques of this disclosure. The method can be performed by devices such as: a graphics processing apparatus, a GPU, a CPU, device 104, a decoder 1204, a wireless communication device, etc., in combination with... Figures 1 to 12 This method is used in various aspects. It can be associated with various advantages at the device level, such as improved compression ratios and reduced computational resources. In one example, the method (including the various aspects detailed below) can be performed by decompressor 199.
[0132] At 1602, the device (e.g., a decoder) obtains a set of bits from the encoder of the compression hardware, including a first subset and a second subset, wherein the bit width of the bit set is greater than the configured bit width of the compression hardware. For example, Figure 12 At 1214, it is shown that decoder 1204 can obtain a set of bits, including a first subset and a second subset, from the encoder (e.g., encoder 1202) of compression hardware, wherein the bit width of the bit set is greater than the configured bit width of the compression hardware. In one example, the bit set may correspond to zero index 520A, zero index 520B, (n-1)th index 522A, or (n-1)th index 522B. In one example, the first subset may correspond to zero index 520B, and the second subset may correspond to zero index 520A. In one example, the compression hardware may be or include zero pixel recovery unit 708 and / or first pixel recovery unit 710. In one example, 1602 may be performed by decompressor 199.
[0133] At 1606, the device (e.g., the decoder) processes the first subset based on the parity corresponding to the second subset. For example, Figure 12 At 1218, it is shown that decoder 1204 can process the first subset based on the parity corresponding to the second subset. In one example, processing the first subset based on the parity corresponding to the second subset corresponds to inverse conditionalization process 1104. In another example, processing the first subset based on the parity corresponding to the second subset corresponds to equation (I) or equation (II) above. In one example, processing the first subset based on the parity corresponding to the second subset can generate zeroth bit component 508A and zeroth bit component 508B. In one example, 1606 can be performed by decompressor 199.
[0134] At 1608, the device (e.g., a decoder) outputs a set of bits that includes the processed first and second bit subsets. For example, Figure 12At 1220, it is shown that decoder 1204 can output a set of bits including the processed first and second bit subsets. In one example, the output bit set may correspond to block 718. In one example, 1608 may be performed by decompressor 199.
[0135] In one aspect, the parity corresponding to the second subset can indicate an odd number, and processing the first subset can include performing a bitwise NOT operation on the first subset. For example, the above aspect can correspond to equation (I) in... Or in equation (II) In one example, processing the first subset at 1218 could include performing a bitwise NOT operation on the first subset.
[0136] In one aspect, the parity corresponding to the second subset can indicate an even number, and processing the first subset can include avoiding adjustments to the first subset. For example, the above aspect can correspond to equation (I) in... Or in equation (II) In one example, processing the first subset at 1218 could include avoiding adjusting the first subset.
[0137] In one aspect, the first subset may correspond to the least significant bit (LSB) component of the bit set, and the second subset may correspond to the most significant bit (MSB) component of the bit set. For example, Figure 10 It is shown that the zero index 520B can correspond to the LSB component of the bit set, and the zero index 520A can correspond to the MSB component of the bit set.
[0138] In one respect, the bit set can correspond to the set of pixels in a frame. For example, the bit set obtained at 1214 can correspond to the set of pixels in a frame.
[0139] In one aspect, outputting a set of bits comprising the processed first and second subsets may include: sending the set of bits to a display; or storing the set of bits in at least one of memory, a buffer, or a cache. For example, outputting a set of bits comprising the processed first and second subsets at 1220 may include: sending the set of bits to a display (e.g., display 131); or storing the set of bits in at least one of memory, a buffer, or a cache.
[0140] In one aspect, obtaining a bit set including the processed first and second bit subsets may include: receiving a compressed bit set including the processed first and second bit subsets, and at 1604, a device (e.g., a decoder) may decompress the compressed bit set including the processed first and second bit subsets, wherein processing the first subset based on the parity corresponding to the second bit subset may include: processing the decompressed first subset based on the parity corresponding to the decompressed second bit subset. For example, obtaining a bit set including the processed first and second bit subsets at 1214 may include: (e.g., from encoder 1202) receiving a compressed bit set including the processed first and second bit subsets. Figure 12 At 1216, decoder 1204 is shown to decompress a compressed bit set including the processed first and second bit subsets, wherein processing the first subset based on the parity corresponding to the second bit subset at 1218 may include processing the decompressed first subset based on the parity corresponding to the decompressed second bit subset. In one example, the foregoing aspect may correspond to zero-pixel recovery unit 708 and / or first pixel recovery unit 710. In one example, 1604 may be performed by decompressor 199.
[0141] In one respect, the bit width of the bit set can be 10 bits, 12 bits, 14 bits, or 16 bits. For example, the bit width of the bit set obtained at 1214 can be 10 bits, 12 bits, 14 bits, or 16 bits.
[0142] In the configuration, a method or apparatus for graphics processing is provided. The apparatus may be a GPU, a CPU, or some other processor capable of performing graphics processing. In various aspects, the apparatus may be a processing unit 120 within device 104, or it may be other hardware within device 104 or another device. The apparatus may include means for obtaining a set of bits including a first subset and a second subset, wherein the bit width of the bit set is greater than the configuration bit width of the compression hardware. The apparatus may also include means for processing the first subset based on parity corresponding to the second subset. The apparatus may further include means for outputting the processed first subset and second subset of bits to the decoder of the compression hardware. The apparatus may also include means for compressing the processed first subset and second subset of bits, wherein outputting the processed first subset and second subset of bits includes: outputting the compressed bit set to the decoder of the compression hardware.
[0143] In the configuration, a method or apparatus for graphics processing is provided. The apparatus may be a GPU, a CPU, or some other processor capable of performing graphics processing. In various aspects, the apparatus may be a processing unit 120 within device 104, or it may be other hardware within device 104 or another device. The apparatus may include means for obtaining a set of bits including a first subset and a second subset from an encoder of compression hardware, wherein the bit width of the bit set is greater than the configuration bit width of the compression hardware. The apparatus may also include means for processing the first subset based on parity corresponding to the second subset. The apparatus may also include means for outputting the bit set including the processed first subset and the second subset. The apparatus may also include means for decompressing the compressed bit set including the processed first subset and the second subset, wherein processing the first subset based on parity corresponding to the second subset includes processing the decompressed first subset based on parity corresponding to the decompressed second subset.
[0144] It should be understood that the specific order or hierarchy of boxes / steps in the processes, flowcharts, and / or call flowcharts disclosed herein are merely illustrative of example methods. It should be understood that the specific order or hierarchy of boxes / steps in these processes, flowcharts, and / or call flowcharts may be rearranged based on design preferences. Furthermore, some boxes / steps may be combined or omitted. Other boxes / steps may also be added. The appended method claims provide the elements of various boxes / steps in an exemplary order, but are not intended to limit one to the given specific order or hierarchy.
[0145] The foregoing description is provided to enable any person skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects. Therefore, the claims are not intended to be limited to the aspects shown herein, but should be given the full scope consistent with the language of the claims, wherein, unless specifically stated otherwise, references to elements in the singular are not intended to mean “one and only one,” but rather “one or more.” The word “exemplary” is used herein to mean “serving as an example, instance, or illustration.” Any aspect described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other aspects.
[0146] Unless otherwise specified, the term "some" refers to one or more, and unless otherwise specified in the context, the term "or" may be interpreted as "and / or". Combinations such as "at least one of A, B, or C", "one or more of A, B, or C", "at least one of A, B, and C", "one or more of A, B, and C", and "A, B, C, or any combination thereof" include any combination of A, B, and / or C, and may include multiple A, multiple B, or multiple C. Specifically, combinations such as "at least one of A, B, or C", "one or more of A, B, or C", "at least one of A, B, and C", "one or more of A, B, and C", and "A, B, C, or any combination thereof" may be only A, only B, only C, A and B, A and C, B and C, or A and B and C, wherein any such combination may include one or more members of A, B, or C. All structural and functional equivalents of the elements throughout the various aspects described herein that are known to or will later be known to a person skilled in the art are expressly incorporated herein by reference and are intended to be covered by the claims. Furthermore, nothing disclosed herein is intended to be offered to the public, whether or not such disclosure is explicitly recited in the claims. The words “module,” “mechanism,” “element,” “device,” etc., cannot replace the word “component.” Therefore, no claim element will be construed as a functional component unless the element is explicitly recited using the phrase “component for…”. Unless otherwise stated, the phrase “processor” may refer to “any processor in one or more processors” (e.g., one processor in one or more processors, a plurality (more than one) of one or more processors, or all processors in one or more processors), and the phrase “memory” may refer to “any memory in one or more memories” (e.g., one memory in one or more memories, a plurality (more than one) of one or more memories, or all memories in one or more memories).
[0147] In one or more examples, the functionality described herein may be implemented in hardware, software, firmware, or any combination thereof. For example, although the term "processing unit" is used throughout this disclosure, such a processing unit may be implemented in hardware, software, firmware, or any combination thereof. If any functionality, processing unit, technique, or other module described herein is implemented in software, then such functionality, processing unit, technique, or other module may be stored on or transmitted on a computer-readable medium as one or more instructions or code.
[0148] Computer-readable media may include computer data storage media and communication media, including any media that facilitates the transfer of computer programs from one place to another. In this way, computer-readable media may generally correspond to: (1) a non-transitory tangible computer-readable storage medium; or (2) a communication medium, such as a signal or carrier wave. Data storage media may be any available medium that can be accessed by one or more computers or one or more processors to extract instructions, code, and / or data structures for implementing the techniques described in this disclosure. By way of example and not limitation, such computer-readable media may include RAM, ROM, EEPROM, compressed optical disc read-only memory (CD-ROM) or other optical disc storage devices, magnetic disk storage devices, or other magnetic storage devices. As used herein, magnetic disks and optical discs include compressed optical discs (CD), laser optical discs, optical discs, digital versatile optical discs (DVD), floppy disks, and Blu-ray discs, wherein magnetic disks typically magnetically copy data, while optical discs optically copy data using lasers. Combinations of the above should also be included within the scope of computer-readable media. Computer program products may include computer-readable media.
[0149] The techniques disclosed herein can be implemented in a wide variety of devices or apparatuses, including wireless mobile phones, integrated circuits (ICs), or IC sets (e.g., chipsets). Various components, modules, or units are described in this disclosure to emphasize functional aspects of a device configured to perform the disclosed techniques, but they do not necessarily need to be implemented by different hardware units. Rather, as described above, various units can be combined in any hardware unit or provided by a collection of interoperable hardware units (including one or more processors as described above) combined with suitable software and / or firmware. Therefore, the term "processor" as used herein can refer to any of the above-described structures or any other structure suitable for implementing the techniques described herein. Furthermore, these techniques can be fully implemented in one or more circuit or logic elements.
[0150] The following aspects are merely illustrative and may be combined with other aspects or teachings described herein without limitation.
[0151] Aspect 1 is a method for data processing, the method comprising: obtaining a bit set including a first subset and a second subset, wherein the bit width of the bit set is greater than the configuration bit width of compression hardware; processing the first subset based on parity corresponding to the second subset; and outputting the bit set to a decoder of the compression hardware, including the processed first subset and the second subset.
[0152] Aspect 2 can be combined with aspect 1, wherein the parity indicator corresponding to the second bit subset is odd, and wherein processing the first bit subset includes performing a bitwise NOT operation on the first bit subset.
[0153] Aspect 3 can be combined with aspect 1, wherein the parity indicator corresponding to the second bit subset is even, and wherein processing the first bit subset includes avoiding adjusting the first bit subset.
[0154] Aspect 4 may be combined with any one of aspects 1 to 3, wherein the first subset corresponds to the least significant bit (LSB) component of the bit set, and wherein the second subset corresponds to the most significant bit (MSB) component of the bit set.
[0155] Aspect 5 can be combined with aspect 4, wherein processing the first subset of bits removes the discontinuities associated with the LSB components.
[0156] Aspect 6 can be combined with any of aspects 1 to 5, wherein the bit set corresponds to the pixel set in the frame.
[0157] Aspect 7 can be combined with any of aspects 1 to 6, wherein the output comprising the bit set of the processed first subset and the second subset comprises: sending the bit set comprising the processed first subset and the second subset to the decoder of the compression hardware.
[0158] Aspect 8 may be combined with any one of aspects 1 to 7, and the method further includes: compressing the bit set comprising the processed first subset and the second subset, wherein outputting the bit set comprising the processed first subset and the second subset comprises: outputting the compressed bit set to the decoder of the compression hardware.
[0159] Aspect 9 may be combined with any one of aspects 1 to 8, wherein the bit width of the bit set is 10 bits, 12 bits, 14 bits or 16 bits.
[0160] Aspect 10 is an apparatus for data processing, the apparatus comprising: a processor coupled to a memory, and configured to implement the method according to any one of aspects 1 to 9 based on information stored in the memory.
[0161] Aspect 11 may be combined with aspect 10 and includes: the device is a wireless communication device, the wireless communication device including at least one of a transceiver or an antenna coupled to the processor, wherein in order to output the bit set, the processor is configured to output the bit set via at least one of the transceiver or the antenna.
[0162] Aspect 12 is an apparatus for data processing, the apparatus comprising components for implementing the method according to any one of aspects 1 to 9.
[0163] Aspect 13 is a computer-readable medium (e.g., a non-transitory computer-readable medium) storing computer-executable code that, when executed by a processor, causes the processor to implement the method according to any one of aspects 1 to 9.
[0164] Aspect 14 is a data processing method, the method comprising: obtaining from an encoder of compression hardware a bit set including a first subset and a second subset, wherein the bit width of the bit set is greater than the configuration bit width of the compression hardware; processing the first subset based on parity corresponding to the second subset; and outputting the bit set including the processed first subset and the second subset.
[0165] Aspect 15 can be combined with aspect 14, wherein the parity indicator corresponding to the second bit subset is odd, and wherein processing the first bit subset includes performing a bitwise NOT operation on the first bit subset.
[0166] Aspect 16 can be combined with aspect 14, wherein the parity indicator corresponding to the second bit subset is even, and wherein processing the first bit subset includes avoiding adjusting the first bit subset.
[0167] Aspect 17 may be combined with any of aspects 14 to 16, wherein the first subset corresponds to the least significant bit (LSB) component of the bit set, and wherein the second subset corresponds to the most significant bit (MSB) component of the bit set.
[0168] Aspect 18 may be combined with any of aspects 14 to 17, wherein the bit set corresponds to the pixel set in the frame.
[0169] Aspect 19 may be combined with any of aspects 14 to 18, wherein the output comprising the bit set of the processed first subset and second subset includes: sending the bit set to a display; or storing the bit set in at least one of a memory, a buffer, or a cache.
[0170] Aspect 20 may be combined with any of aspects 14 to 19, wherein obtaining the bit set including the processed first subset and the second subset includes: receiving a compressed bit set including the processed first subset and the second subset, the method further comprising: decompressing the compressed bit set including the processed first subset and the second subset, wherein processing the first subset based on the parity corresponding to the second subset includes: processing the decompressed first subset based on the parity corresponding to the decompressed second subset.
[0171] Aspect 21 may be combined with any of aspects 14 to 20, wherein the bit width of the bit set is 10 bits, 12 bits, 14 bits or 16 bits.
[0172] Aspect 22 is an apparatus for data processing, the apparatus comprising: a processor coupled to a memory, and configured to implement the method according to any one of aspects 14 to 21 based on information stored in the memory.
[0173] Aspect 23 may be combined with aspect 22 and includes: the device is a wireless communication device, the wireless communication device including at least one of a transceiver or an antenna coupled to the processor, wherein in order to obtain the bit set, the processor is configured to obtain the bit set via at least one of the transceiver or the antenna.
[0174] Aspect 24 is an apparatus for data processing, the apparatus comprising components for implementing the method according to any one of aspects 14 to 21.
[0175] Aspect 25 is a computer-readable medium (e.g., a non-transitory computer-readable medium) storing computer-executable code that, when executed by a processor, causes the processor to implement the method according to any one of aspects 14 to 21.
[0176] Various aspects have been described herein. These and other aspects are within the scope of the following claims.
Claims
1. An apparatus for data processing, the apparatus comprising: Memory; and A processor, coupled to the memory, and configured based on information stored in the memory, to: Obtain a set of bits including the first subset and the second subset, wherein the bit width of the set of bits is greater than the configuration bit width of the compression hardware; The first position subset is processed based on the parity corresponding to the second position subset; as well as The decoder output for the compression hardware includes the bit set of the first and second bit subsets processed.
2. The apparatus of claim 1, wherein the parity indicator corresponding to the second bit subset is odd, and wherein, in order to process the first bit subset, the processor is configured to perform a bitwise NOT operation on the first bit subset.
3. The apparatus of claim 1, wherein the parity indicator corresponding to the second bit subset is an even number, and wherein, in order to process the first bit subset, the processor is configured to avoid adjusting the first bit subset.
4. The apparatus of claim 1, wherein the first subset corresponds to the least significant bit (LSB) component of the bit set, and wherein the second subset corresponds to the most significant bit (MSB) component of the bit set.
5. The apparatus of claim 4, wherein, in order to process the first bit subset, the processor is configured to process the first bit subset to remove discontinuities associated with the LSB component.
6. The apparatus of claim 1, wherein the bit set corresponds to the pixel set in a frame.
7. The apparatus of claim 1, wherein, in order to output the bit set comprising the processed first subset and the second subset of bits, the processor is configured to send the bit set comprising the processed first subset and the second subset of bits to the decoder of the compression hardware.
8. The apparatus of claim 1, wherein the processor is configured to: Compression includes the set of bits comprising the processed first subset and the second subset, wherein, in order to output the set of bits comprising the processed first subset and the second subset, the processor is configured to output the compressed set of bits to the decoder of the compression hardware.
9. The apparatus of claim 1, wherein the bit width of the bit set is 10 bits, 12 bits, 14 bits, or 16 bits.
10. The apparatus of claim 1, wherein the apparatus is a wireless communication device, the wireless communication device including at least one of a transceiver or an antenna coupled to the processor, wherein in order to output the bit set, the processor is configured to output the bit set via at least one of the transceiver or the antenna.
11. An apparatus for data processing, the apparatus comprising: Memory; and A processor, coupled to the memory, and configured based on information stored in the memory, to: A set of bits, including a first subset and a second subset, is obtained from the encoder of the compression hardware, wherein the bit width of the bit set is greater than the configured bit width of the compression hardware. The first position subset is processed based on the parity corresponding to the second position subset; and The output includes the set of bits of the first and second bit subsets processed.
12. The apparatus of claim 11, wherein the parity indicator corresponding to the second bit subset is odd, and wherein, in order to process the first bit subset, the processor is configured to perform a bitwise NOT operation on the first bit subset.
13. The apparatus of claim 11, wherein the parity indicator corresponding to the second bit subset is an even number, and wherein, in order to process the first bit subset, the processor is configured to avoid adjusting the first bit subset.
14. The apparatus of claim 11, wherein the first subset corresponds to the least significant bit (LSB) component of the bit set, and wherein the second subset corresponds to the most significant bit (MSB) component of the bit set.
15. The apparatus of claim 11, wherein the bit set corresponds to the pixel set in a frame.
16. The apparatus of claim 11, wherein, in order to output the bit set including the processed first subset and the second subset of bits, the processor is configured to: Send the bit set to the display; or The bit set is stored in at least one of the memory, buffer, or cache.
17. The apparatus of claim 11, wherein, in order to obtain the bit set including the processed first subset and the second subset of bits, the processor is configured to receive a compressed bit set including the processed first subset and the second subset of bits, and wherein the processor is further configured to: Decompression includes the compressed bit set of the processed first bit subset and the second bit subset, wherein in order to process the first bit subset based on the parity corresponding to the second bit subset, the processor is configured to process the decompressed first bit subset based on the parity corresponding to the decompressed second bit subset.
18. The apparatus of claim 11, wherein the bit width of the bit set is 10 bits, 12 bits, 14 bits, or 16 bits.
19. The apparatus of claim 11, wherein the apparatus is a wireless communication device, the wireless communication device including at least one of a transceiver or an antenna coupled to the processor, wherein in order to obtain the bit set, the processor is configured to obtain the bit set via at least one of the transceiver or the antenna.
20. A method for data processing, the method comprising: Obtain a set of bits including the first subset and the second subset, wherein the bit width of the set of bits is greater than the configuration bit width of the compression hardware; The first position subset is processed based on the parity corresponding to the second position subset; as well as The decoder output for the compression hardware includes the bit set of the first and second bit subsets processed.
21. The method of claim 20, wherein the parity indicator corresponding to the second bit subset is odd, and wherein processing the first bit subset includes performing a bitwise NOT operation on the first bit subset.
22. The method of claim 20, wherein the parity indicator corresponding to the second bit subset is even, and wherein processing the first bit subset includes avoiding adjusting the first bit subset.
23. The method of claim 20, wherein the first subset corresponds to the least significant bit (LSB) component of the bit set, and wherein the second subset corresponds to the most significant bit (MSB) component of the bit set.
24. The method of claim 23, wherein processing the first subset of bits removes the discontinuities associated with the LSB components.
25. The method of claim 20, wherein the bit set corresponds to the pixel set in the frame.
26. The method of claim 20, wherein the output comprising the bit set including the processed first bit subset and the second bit subset comprises: The set of bits, comprising the processed first subset and the second subset, is sent to the decoder of the compression hardware.
27. The method of claim 20, further comprising: Compression includes the set of bits of the processed first subset and second subset, wherein the output including the set of bits of the processed first subset and second subset includes: outputting the compressed set of bits to the decoder of the compression hardware.
28. The method of claim 20, wherein the bit width of the bit set is 10 bits, 12 bits, 14 bits, or 16 bits.
29. A method for data processing, the method comprising: A set of bits, including a first subset and a second subset, is obtained from the encoder of the compression hardware, wherein the bit width of the bit set is greater than the configured bit width of the compression hardware. The first position subset is processed based on the parity corresponding to the second position subset; and The output includes the set of bits of the first and second bit subsets processed.
30. The method of claim 29, wherein the parity indicator corresponding to the second bit subset is odd, and wherein processing the first bit subset includes performing a bitwise NOT operation on the first bit subset.