Apparatus, method, program, and device for a framework for GPU-driven data loading
The GPU-driven data loading framework addresses CPU bottlenecks in GPU data loading by shifting data loading and decoding to the GPU, resulting in improved performance, reduced power consumption, and adaptable algorithm customization.
Patent Information
- Application Number
- JP2024563945
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-20
- Filing Date
- 2023-03-16
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2043-03-16
AI Technical Summary
Current GPU data loading processes are bottlenecked by the CPU's involvement in loading and decoding data, leading to slowed performance and increased power consumption, especially in complex rendering scenarios.
A GPU-driven data loading framework that offloads data loading and decoding processes from the CPU to the GPU, allowing the GPU to load data directly from storage to video memory and decode it in parallel using multiple GPU thread groups.
This approach reduces CPU involvement, leading to faster processing, lower power consumption, and reduced bandwidth usage, while also providing a flexible framework for customizing encoding and decoding algorithms.
Smart Images

Figure 2025516249000001_ABST
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Patent Application No. 18 / 085,367, filed December 20, 2022, the entire contents of which are incorporated herein by reference.
[0002] The present disclosure generally relates to graphics processing units (GPUs) that include one or more techniques for loading data in a computing device. [Background technology]
[0003] Computing devices often utilize a graphics processing unit (GPU) in combination with a central processing unit (CPU) to render graphical data for display or to perform non-graphics related functions that utilize the massive processing parallelism provided by the GPU. Such computing devices may include, for example, computer workstations, mobile phones such as smartphones, embedded systems, personal computers, tablet computers, and video game consoles. The GPU processes instructions and / or data in a processing pipeline that includes one or more processing stages that work together to perform processing commands for graphically related and non-graphically related functions. The CPU may control the operation of the GPU by issuing one or more processing commands to the GPU. Modern CPUs are typically capable of simultaneously executing multiple applications, each of which may need to utilize the GPU during execution.
[0004] While GPUs were initially intended to improve graphics rendering, the parallel computing nature of GPUs has proven beneficial in accelerating a variety of processing applications. The ability to perform separate tasks in parallel, as well as the modular architecture of modern GPUs, means that there can be a variety of ways in which a solution can be designed for any graphical or non-graphical need. This has increased GPU power by making GPUs more adaptable and programmable for purposes other than rendering. Today, GPU parallel computing is used for a wide range of different applications.
[0005] Architecturally, a CPU consists of one or more cores with cache memory that can process several software threads at a time. This makes a CPU suitable for sequential processing, since it can quickly execute a series of operations at once. In contrast, a GPU may contain hundreds of cores that may be capable of processing thousands of threads simultaneously. This makes a GPU suitable for parallel processing, since it can process thousands of operations at once. The hundreds of cores in a GPU are low power and are more suitable for performing simple simultaneous calculations, such as arithmetic. Thus, GPU parallel computing allows a GPU to split a complex problem into thousands or millions of separate tasks and process them all at once, instead of one at a time like a CPU. This also makes a GPU more powerful than a CPU, since it has more cores, more computing power, and therefore more potential for parallelism in calculations. Summary of the Invention [Problem to be solved by the invention]
[0006] Typically, the GPU waits for the CPU to load data from storage, decode the data, and transfer the decoded data to video memory where the GPU can perform its processing. However, as the complexity of the content being rendered increases and CPU performance constraints increase, the need for improved graphics processing or computer processing increases.
[0007] The following presents a simplified summary of one or more aspects in order to provide a basic understanding of such aspects. This summary is not an extensive overview of all possible aspects, and is not intended to identify key elements of all aspects or to delineate the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simplified form as a prelude to the more detailed description that is presented later. [Means for solving the problem]
[0008] The present disclosure relates to a method and apparatus for data loading in a computing device. One aspect of the subject matter described in this disclosure is embodied in a method for loading data in a computing device. The method includes identifying data to load in a GPU based on execution of an application program. The method also includes loading, via the GPU, data chunks of the identified data in encoded form from a data storage device to a video memory associated with the GPU. The method further includes decoding the data chunks in parallel by using multiple GPU thread groups in parallel to decode the data chunks. Each of the data chunks is decoded independently of the other data chunks.
[0009] Another further aspect of the subject matter described in this disclosure may be embodied in an apparatus for data loading in a computing device. The apparatus includes a graphics processing unit (GPU) configured to identify data to load based on execution of an application program. The GPU is also configured to load data chunks of the identified data in encoded form from a data storage device to a video memory associated with the GPU. The GPU is further configured to decode the data chunks in parallel by using multiple GPU thread groups in parallel to decode the data chunks. Each of the data chunks is decoded independently of the other data chunks.
[0010] Other further aspects of the subject matter described in this disclosure may be embodied in a non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the processor to identify data to load in a graphics processing unit (GPU) based on execution of an application program. The processor is also configured to load, via the GPU, data chunks of the identified data in encoded form from a data storage device to a video memory associated with the GPU. The processor is further configured to decode the data chunks in parallel by using multiple GPU thread groups in parallel to decode the data chunks. Each of the data chunks is decoded independently of the other data chunks.
[0011] Yet another further aspect of the subject matter described in this disclosure may be implemented in a device. The device includes a controller configured to identify data to load based on execution of an application program. The controller is also configured to load data chunks of the identified data in encoded form from a data storage device to a video memory associated with the GPU. The controller is further configured to decode the data chunks in parallel by using multiple GPU thread groups in parallel to decode the data chunks. Each of the data chunks is decoded independently of the other data chunks.
[0012] To the accomplishment of the foregoing and related ends, the one or more aspects comprise the features hereinafter fully described and particularly pointed out in the claims. The following description and the annexed drawings set forth in detail certain illustrative features of the one or more aspects. These features are indicative, however, of but a few of the various ways in which the principles of the various aspects may be employed and the description is intended to include all such aspects and their equivalents.
[0013] Details of one or more aspects of the subject matter described in this disclosure are set forth in the accompanying drawings and the following description. However, the accompanying drawings only illustrate some typical aspects of the disclosure and therefore should not be considered as limiting its scope. Other features, aspects, and advantages will become apparent from the description, drawings, and claims. [Brief description of the drawings]
[0014] [Figure 1A] FIG. 1 is a block diagram illustrating an example of a data loading system in accordance with one or more techniques of the present disclosure. [Figure 1B] FIG. 1 is a block diagram illustrating an example of a data loading system in accordance with one or more techniques of the present disclosure. [Figure 2A]FIG. 2 is a diagram of an example of a data loading pipeline handled by multiple processors. [Figure 2B] FIG. 1 is a diagram of an example GPU-driven rendering pipeline. [Diagram 3] FIG. 2 is a diagram of an example data loading pipeline handled by a GPU in accordance with one or more techniques of this disclosure. [Figure 4] 1 is a flowchart of an example method of data loading in a computing device in accordance with one or more techniques of this disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0015] Like reference numbers and designations in the various drawings indicate like elements.
[0016] The following description is directed to several exemplary embodiments for the purposes of illustrating the innovative aspects of the present disclosure, however, those skilled in the art will readily recognize that the teachings herein can be applied in many different ways.
[0017] Related systems have a framework for loading data using multiple processors. As a first issue, the data may be any block of data consumed in a runtime process. For example, the data may be video data, mesh data, texture data, machine learning training data, text data, etc. In other words, the data may be any type of data that may benefit from parallel processing. In related systems, the GPU may load data from storage and request the CPU to decode the data into a form that the GPU can consume before the decoded or uncompressed data can be transferred to video memory for GPU use. However, this process creates a CPU bottleneck that slows down the process due to the swapping of the GPU and the CPU. Furthermore, the CPU bottleneck is exacerbated on mobile devices because the CPU consumes a significant amount of processing power on mobile devices.
[0018] In a related GPU-driven rendering pipeline, the GPU receives the data required for rendering by loading the video data from storage and requesting the CPU to decode the video data. However, this process stalls the GPU pipeline, slowing down rendering. There are many related systems that have data loading frameworks used to render graphics. Typically, compressed data is loaded into system memory and decompressed by the CPU before being sent to the GPU, which increases the loading time.
[0019] As an example, a first related system may load a file into memory with or without decoding. As another example, in a second related system, the video data must be compressed by the Kraken algorithm, which provides a special API for loading the data. However, a dedicated chip must be used to decode the video data from the Kraken algorithm into GPU consumable content. As yet another example, in a third related system, only asynchronous loading is available, and there is no encoding or decoding part. The third related system may also have a new API for mapping storage into memory, which requires changes to the drive and OS, and is not practical for mobile devices.
[0020] Another fourth related system may include fast resource loading to stream data into textures and buffer directly from storage using an asynchronous input / output (I / O) application programming interface (API). Yet another fifth related system may transfer data between the GPU and other devices in the data center. However, both the fourth related system and the fifth related system cannot run on multiple platforms and multiple hardware. Furthermore, the above-mentioned related systems generally have one fixed decompression algorithm or no decompression algorithm at all. This means that the related framework cannot reduce bandwidth by customizing the decompression algorithm to different scenarios or by employing different decoding for different types of data. In real-world applications, different scenarios may have their best suitable compression techniques.
[0021] Aspects of the present disclosure utilize a cross-platform GPU-driven data loading framework that can be used on desktops, mobiles, consoles, servers, and the like. Furthermore, no new operating systems (OS), new application programming interfaces (APIs), or new hardware need to be modified or used to implement the disclosed techniques. The framework offloads processes typically performed by the CPU to be performed on the GPU, reducing power consumption and improving performance by bypassing the CPU bottleneck. Data may also be kept in a compressed form until it is ready to be consumed by the GPU, which also reduces bandwidth when transferring data, further reducing power consumption.
[0022] Aspects of the present disclosure utilize an adaptable framework so that they do not need to focus on a specific algorithm. Instead, the GPU-driven data loading framework provides the adaptability to customize encoding / decoding algorithms according to different scenarios. This allows developers to create WARP-based parallel building blocks utilizing customized decoding algorithms, rather than starting from scratch. Unlike related GPU programs that operate within blocks, these WARP-based building blocks for compression or decompression algorithms are optimized to operate within WARP to modify sequential compression or sequential decompression so that those algorithms can be used in parallel by the GPU.
[0023] Various aspects of the system, device, computer program product, and method are described more fully below with reference to the accompanying drawings. However, the present disclosure may be embodied in many different forms and should not be construed as limited to a specific structure or specific functions presented throughout the present disclosure. Rather, these aspects are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art. Based on the teachings herein, one skilled in the art should understand that the scope of the present disclosure is intended to cover any aspect of the system, device, computer program product, and method disclosed herein, whether implemented independently of or in combination with other aspects of the present disclosure. For example, an apparatus may be implemented or a method may be performed using any number of the aspects described herein. Moreover, the scope of the present disclosure is intended to cover such an apparatus or method that is implemented using other structures, functions, or that is implemented using other structures, functions in addition to or other than the various aspects of the present disclosure described herein. Any aspect disclosed herein may be embodied by one or more elements of a claim.
[0024] Although various aspects are described herein, many variations and permutations of these aspects are within the scope of the disclosure. Although some potential benefits and advantages of the aspects of the disclosure are mentioned, the scope of the disclosure is not intended to be limited to any particular benefit, use, or purpose. Rather, the aspects of the disclosure are intended to be broadly applicable to different wireless technologies, system configurations, networks, and transmission protocols, some of which are shown as examples in the drawings and the following description. The detailed description and drawings are merely illustrative of the disclosure rather than limiting, and the scope of the disclosure is defined by the appended claims and their equivalents.
[0025] Several aspects are presented with reference to various apparatus and methods, which are described in the following detailed description and illustrated in the accompanying drawings by various blocks, components, circuits, processes, algorithms, etc. (collectively referred to as "elements"). These elements may be implemented using electronic hardware, computer software, or any combination thereof. Whether such elements are implemented as hardware or software depends on the particular application and design constraints imposed on the overall system.
[0026] As an example, an element, or any portion of an element, or any combination of elements, may be implemented as a "processing system" including one or more processors (which may be referred to as processing circuits). One or more processors in the processing system may execute software. Software may be broadly construed to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software components, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, functions, and the like, whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise. The term application may refer to software. As described herein, one or more techniques may refer to an application, i.e., software, configured to perform one or more functions. In such an example, the application may be stored in memory, e.g., the on-chip memory of a processor, system memory, or any other memory. Hardware described herein, such as a processor, may be configured to execute the application. For example, an application may be described as including code that, when executed by the hardware, causes the hardware to perform one or more techniques described herein. As an example, hardware may access code from memory and execute the code accessed from memory to perform one or more techniques described herein. In some examples, components are identified in this disclosure. In such examples, the components may be hardware, software, or a combination thereof. The components may be separate components or subcomponents of a single component.
[0027] Thus, in one or more examples described herein, the functions described may be implemented in hardware, software, or any combination thereof. If implemented in software, the functions may be stored or encoded as one or more instructions or code on a computer-readable medium. Computer-readable media includes computer storage media. A storage medium may be any available medium that can be accessed by a computer. By way of example and not limitation, such computer-readable media may include random access memory (RAM), read-only memory (ROM), electrically erasable programmable ROM (EEPROM), optical disk storage, magnetic disk storage, other magnetic storage, combinations of the aforementioned types of computer-readable media, or any other medium that can be used to store computer-executable code in the form of instructions or data structures that can be accessed by a computer.
[0028] As a first issue, it should be noted that the term "GPU" as used herein does not necessarily mean a processing device used exclusively for graphics processing. Conversely, the GPUs described herein are parallel processing accelerators. While a CPU typically consists of several cores optimized for a series of sequential processes, a GPU typically has a massively parallel architecture that may contain thousands of smaller, more efficient computing cores designed to handle multiple tasks simultaneously. This allows GPUs to be used for many purposes other than graphics, including accelerating high performance computing, deep learning and artificial intelligence, analytics and other processing applications.
[0029] The parallel architecture also makes GPUs ideal for deep learning and neural networks because GPUs perform many simultaneous calculations, thereby reducing the time it takes to train a neural network using traditional CPU techniques from days to hours. As described herein, each GPU may be used for any advanced processing task, and is particularly useful for complex tasks that benefit from massively parallel processing.
[0030] The present disclosure includes techniques for data loading in a computing device by utilizing a framework to load GPU-driven data with minimal CPU involvement. Aspects of the present disclosure offload the loader and decoder processes, which are typically performed sequentially by the CPU, to be performed in parallel by the GPU. With this framework, the GPU may load data directly from storage to video memory and decode them on the GPU. In addition, data is kept in compressed form until consumed by the GPU, reducing bandwidth to transfer and further reducing power consumption. The encoding and decoding algorithms may also be customizable without the need to modify existing operating systems (OS), APIs, or hardware. Because the CPU is bypassed, the disclosed data loading framework requires less power consumption and provides faster processing. Other example benefits are described throughout the present disclosure.
[0031] Examples of the term "content" as used herein may refer to "graphical content," "images," and vice versa. This is true regardless of whether the terms are used as adjectives, nouns, or other parts of speech. In some examples, the term "graphical content" as used herein may refer to content generated by one or more processes of a graphics processing pipeline. In some examples, the term "graphical content" as used herein may refer to content generated by a processing device configured to perform graphics processing. In some examples, the term "graphical content" as used herein may refer to content generated by a graphics processing device.
[0032] As used herein, the term "display content" may refer to content generated by a processing device configured to perform display processing. In some examples, as used herein, the term "display content" may refer to content generated by a display processing device. Graphical content may be processed to become display content. For example, a graphics processing device may output graphical content, such as a frame, to a buffer (sometimes referred to as a frame buffer). The display processing device may read graphical content, such as one or more frames, from the buffer and perform one or more display processing techniques to generate the display content. For example, a display processing device may be configured to perform compositing on one or more rendered layers to generate a frame. As another example, a display processing device may be configured to compose, blend, or otherwise combine two or more layers together into a single frame. A display processing device may be configured to perform scaling, e.g., upscaling or downscaling, on a frame. In some examples, a frame may refer to a layer. In other examples, a frame may refer to two or more layers that have already been fused together to form a frame, i.e., a frame includes two or more layers, and a frame including two or more layers may be subsequently fused.
[0033] A related system loads data in a computing device that utilizes processes performed by both the CPU and the GPU. The CPU first loads the data from a disk or storage device, decodes it, and optionally decompresses the loaded data into a form that the GPU can consume, and then sends the decoded or decompressed data to the GPU memory. However, in mobile devices, the CPU consumes a significant amount of power and can be slow due to CPU processing bottlenecks. If the entire loading process can be performed directly by the GPU with minimal CPU involvement, the processing pipeline is much faster and requires less power consumption.
[0034] Accordingly, embodiments of the present disclosure include a method of data loading in a computing device and an apparatus for implementing a framework for data loading in a computing device. It is noted that while rendering video data for visual representation is used as an example, aspects of the present disclosure may be applied to loading any data used by a computing device in a runtime process that may benefit from parallel processing. With this framework, a GPU may load data from storage to video memory and decode the data with minimal CPU involvement. The subject matter described herein may be implemented to achieve one or more benefits or advantages. For example, by moving roles traditionally performed by the CPU, such as loading and decoding, to the GPU, the embodiments may maximize performance, minimize bandwidth, and reduce power consumption. Additionally, the embodiments may be implemented without modifying existing OS, APIs, or hardware. The framework also utilizes common compression and decompression algorithms such that the framework may employ different decoding for different types of data.
[0035] 1A is a block diagram illustrating an example of a data loading system 100 configured to implement one or more techniques of the present disclosure. The data loading system 100 includes a CPU 128, a GPU 120, and a system memory 124 configured to load data according to an exemplary embodiment. The CPU 128 may execute a software application 111, an OS 113, and a graphics driver 115. In addition, the system memory 124 may include an indirect buffer that stores command streams for rendering primitives, as well as secondary commands to be executed by the GPU 120. The GPU 120 may include a video memory 121, which may be "on-board" with the GPU 120, including a decider function 123, a loader function 125, and a decoder function 127. As described in more detail in connection with FIG. 1B, the components of the data loading system 100 may be part of devices including, but not limited to, video devices, media players, set-top boxes, wireless handsets such as mobile phones and so-called smartphones, personal digital assistants (PDAs), desktop computers, laptop computers, gaming consoles, video conferencing units, tablet computing devices, and the like.
[0036] CPU 128 may be coupled to one or more GPUs. GPU 120 may include a processing unit configured to perform graphics-related functions, such as generating and outputting graphics data for presentation on a display, as well as non-graphics-related functions that exploit the processing parallelism provided by GPU 120. Because GPU 120 may provide general-purpose processing capabilities in addition to graphics processing capabilities, GPU 120 may be referred to as a general-purpose GPU (GP-GPU). Examples of CPU 128 and GPU 120 include, but are not limited to, digital signal processors (DSPs), general-purpose microprocessors, application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. In some examples, GPU 120 may be a microprocessor designed for a specific application, such as providing massively parallel processing for processing graphics as well as for running non-graphics-related applications. Additionally, while CPU 128 and GPU 120 are shown as separate components, aspects of the present disclosure are not so limited and may be implemented, for example, in a common integrated circuit (IC).
[0037] Software applications 111 executing on CPU 128 may include one or more graphics rendering instructions that instruct CPU 128 to render graphics data to a display (not shown in FIG. 1A). In some examples, the graphics rendering instructions may include software instructions that may conform to a graphics API. To process the graphics rendering instructions, CPU 128 may issue one or more graphics rendering commands to GPU 120 (e.g., via graphics driver 115) to cause GPU 120 to perform some or all of the rendering of the graphics data. In some examples, the graphics data to be rendered may include a list of graphics primitives, such as, for example, points, lines, triangles, quadrilaterals, triangle strips, etc.
[0038] GPU 120 may be configured to perform graphics operations to render one or more graphics primitives to a display. Thus, when one of the software applications executing on CPU 128 requires graphics processing, CPU 128 may provide graphics commands and graphics data to GPU 120 for rendering to a display. The graphics data may include, for example, drawing commands, state information, primitive information, texture information, etc. GPU 120 may be built with a highly parallel structure that, in some cases, provides more efficient processing of complex graphics-related operations than CPU 128. For example, GPU 120 may include multiple processing elements configured to operate in parallel on multiple vertices or pixels.
[0039] GPU 120 may process data locally using local storage (i.e., video memory 121) instead of host or system memory. This allows GPU 120 to operate in a more efficient manner, for example, by eliminating the need for GPU 120 to read and write data over a shared bus that may experience large amounts of bus traffic. Video memory 121 may include one or more volatile or non-volatile memory or storage devices, such as, for example, random access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), and one or more registers.
[0040] The video memory 121 may also be used directly by the decider function 123, the loader function 125, and the decoder function 127. The decider function 123 may be configured to identify data to load based on the execution of an application program. The loader function 125 may be configured to load, via the GPU, data chunks of the identified data in encoded form from a data storage device to a video memory associated with the GPU (e.g., video memory 121). The decoder function 127 may be configured to decode the data chunks in parallel by using multiple GPU thread groups in parallel to decode the data chunks. In some aspects, the processor performing the above-mentioned functions may be a general-purpose processor (e.g., a CPU).
[0041] The CPU 128 and / or GPU 120 may store the rendered image data in a frame buffer 138, which may be a separate memory or may be allocated within the system memory 124. A display processor may retrieve the rendered image data from the frame buffer 138 and display the rendered image data on a display.
[0042] System memory 124 may be memory within the device or may be external to CPU 128 and GPU 120, i.e., off-chip to CPU 128 and off-chip to GPU 120. System memory 124 may store applications executed by CPU 128 and GPU 120. Additionally, system memory 124 may store data on which the executed applications operate, as well as data resulting from the applications.
[0043] The system memory 124 may store program modules, instructions, or both accessible for execution by the CPU 128, and / or data for use by programs executing on the CPU 128. For example, the system memory 124 may store a window manager application used by the CPU 128 to present a graphical user interface (GUI) on a display. Additionally, the system memory 124 may store user applications and application surface data associated with the applications. As described in more detail below, the system memory 124 may act as device memory for the GPU 120 and store data on which the GPU 120 should operate as well as data resulting from operations performed by the GPU 120. For example, the system memory 124 may store any combination of texture buffers, depth buffers, stencil buffers, vertex buffers, frame buffers, and the like.
[0044] Examples of system memory 124 include, but are not limited to, random access memory (RAM), read only memory (ROM), or electrically erasable programmable read only memory (EEPROM), or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer or processor. As an example, system memory 124 may be removed from the device and moved to another device. As another example, a storage device substantially similar to system memory 124 may be inserted into the device.
[0045] FIG 1B is a detailed block diagram illustrating an example of a data loading system 100 configured to implement one or more techniques of the present disclosure. It should be noted that the data loading system 100 illustrated in FIG 1B may correspond to the data loading system of FIG 1A. In this regard, the data loading system 100 of FIG 1B includes a CPU 128, a GPU 120, and a system memory 124.
[0046] As further shown, the data loading system 100 includes a device 104 that may include one or more components configured to perform one or more techniques of the present disclosure. In the example shown, the device 104 may include a GPU 120, a content encoder / decoder 137, and a system memory 124. In some aspects, the device 104 may include several additional components, such as a communication interface 126, a transceiver 132, a receiver 133, and a transmitter 130, as well as one or more displays 131. References to a display 131 may refer to one or more displays 131. For example, the display 131 may include a single display or multiple displays. The display 131 may include a first display and a second display. In a further example, the results of the graphics processing may not be displayed on the device, e.g., the display 131 may not receive frames for presentation thereon. Instead, the results of the processing of the frames or graphics may be transferred to another device. In some aspects, this may be referred to as hybrid rendering.
[0047] The GPU 120 includes a video memory 121. The GPU 120 may be configured to perform graphics processing or non-graphics processing. The GPU 120 may be configured to identify data to load based on the execution of an application program, load data chunks of the identified data in encoded form from a data store video to a video memory associated with the GPU (e.g., the video memory 121), and decode the data chunks in parallel by using multiple GPU thread groups in parallel to decode the data chunks. The content encoder / decoder 137 may include an internal memory 135. In some examples, the device 104 may include a display processor, such as a CPU 128, to perform one or more display processing techniques on one or more frames generated by the GPU 120 prior to presentation by one or more displays 131, as described above. The CPU 128 may be configured to perform display processing. The one or more displays 131 may be configured to display or present the frames processed by the CPU 128. In some examples, the one or more displays 131 may include one or more of a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, a projection display device, an augmented reality display device, a virtual reality display device, a head-mounted display, or any other type of display device.
[0048] Memory external to the GPU 120 and the content encoder / decoder 137, such as the system memory 124 as described above, may be accessible to the GPU 120 and the content encoder / decoder 137. For example, the GPU 120 and the content encoder / decoder 137 may be configured to read from and / or write to an external memory, such as the system memory 124. The GPU 120 and the content encoder / decoder 137 may be communicatively coupled to the system memory 124 via a bus. In some examples, the GPU 120 and the content encoder / decoder 137 may be communicatively coupled to each other via a bus or a different connection.
[0049] The content encoder / decoder 137 may be configured to receive graphical content or data from any source, such as the system memory 124 and / or the communication interface 126. The system memory 124 may be configured to store the received encoded or decoded graphical content or data. The content encoder / decoder 137 may be configured to receive encoded or decoded graphical content or data in the form of encoded pixel data or encoded data, for example, from the system memory 124 and / or the communication interface 126. The content encoder / decoder 137 may be configured to encode or decode any graphical content or data.
[0050] According to some examples, video memory 121 or system memory 124 may be a non-transitory computer-readable storage medium. The term "non-transitory" may indicate that the storage medium is not embodied in a carrier wave or propagating signal. However, the term "non-transitory" should not be interpreted to mean that video memory 121 or system memory 124 are non-movable or that their contents are static. As an example, system memory 124 may be removed from device 104 and moved to another device. As another example, system memory 124 may not be removable from device 104.
[0051] The GPU (or processing circuitry) may be configured to perform graphics or non-graphics processing according to the exemplary techniques as described herein. In some examples, the GPU 120 may be integrated into the motherboard of the device 104. In some examples, the GPU 120 may be on a graphics card installed in a port on the motherboard of the device 104 or may be incorporated into a peripheral device configured to interoperate with the device 104. The GPU 120 may include one or more processors, such as one or more microprocessors, GPUs, ASICs, FPGAs, ALUs, DSPs, discrete logic, software, hardware, firmware, other equivalent integrated circuits or discrete logic circuits, or any combination thereof. If the techniques are implemented in part in software, the GPU 120 may store instructions for the software in a suitable non-transitory computer-readable storage medium and execute these instructions in hardware using one or more processors to perform the techniques of this disclosure. Any of the above, including hardware, software, a combination of hardware and software, etc., may be considered to be one or more processors.
[0052] In some aspects, the data loading system 100 may include a communication interface 126. The communication interface 126 may include a receiver 133 and a transmitter 130. The receiver 133 may be configured to perform any receiving function described herein with respect to the device 104. Additionally, the receiver 133 may be configured to receive information from another device, such as, for example, eye or head position information, rendering commands, or position information. The transmitter 130 may be configured to perform any transmitting function described herein with respect to the device 104. For example, the transmitter 130 may be configured to transmit information to the other device, which may include a request for content. The receiver 133 and the transmitter 130 may be combined into a transceiver 132. In such an example, the transceiver 132 may be configured to perform any receiving and / or transmitting functions described herein with respect to the device 104.
[0053] 1B , in certain aspects, the device 104 may include a control component 198 configured to control a processor (e.g., GPU 120) to identify data to load based on execution of an application program. Further, the control component 198 may also be configured to control the processor to load data chunks of the identified data in encoded form from a data storage device into a video memory (e.g., 121) associated with the GPU. In addition, the control component 198 may be further configured to control the processor to decode the data chunks in parallel by using multiple GPU thread groups in parallel to decode the data chunks.
[0054] As described herein, a device, such as device 104, may refer to any device, apparatus, or system configured to perform one or more of the techniques described herein. For example, a device may be a server, a client device, a computer (e.g., a personal computer), a desktop computer, a laptop computer, a tablet computer, a computer workstation, or a mainframe computer, a telephone, a smartphone, a video game platform or console, a handheld device (e.g., a portable video game device or a personal digital assistant (PDA)), a wearable computing device, such as a smart watch, an augmented reality device, a virtual reality device, a display or display device, a television, a television set-top box, a network device, a digital media player, a video streaming device, a content streaming device, an in-vehicle computer, or any other device configured to perform one or more of the techniques described herein. Although the processes herein may be described as being performed by a particular component, such as a GPU, in further embodiments, they may be performed using other processing components configured to perform the described processes.
[0055] FIG. 2A illustrates a diagram of an example of a data loading pipeline handled by multiple processors. Specifically, FIG. 2A illustrates the roles of and interactions between multiple processors during the life cycle of the data. As mentioned above, the data may be any block of data consumed in a runtime process. As shown in FIG. 2A, the processing pipeline 200a includes at least a hard drive 201, a host memory 203, a CPU 205, a video memory 207, and a GPU 209.
[0056] In the related system, the CPU 205 performs decider function 211, loader function 213, and decoder function 215, and the GPU 209 performs consuming function 217. The decider function 211 is configured to determine which data to load. The loader function 213 is configured to read and load the data into memory. The decoder function 215 is configured to decode the data. The consuming function 217 is configured to consume the data.
[0057] In processing pipeline 200a, GPU 209 requests CPU 205 to load data from hard drive 201. The data is decoded by CPU 205 and sent to host memory 203. CPU 205 and video memory 207 may then access the decoded data from host memory 203. The process of the GPU requesting the CPU to load data and then waiting for the CPU to decode the data creates a bottleneck while the GPU is waiting for the data.
[0058] FIG. 2B illustrates a diagram of an example of a GPU-driven rendering pipeline operated by multiple processors. Specifically, FIG. 2B illustrates the roles of multiple processors and the interactions between the processors during the life cycle of video data. As shown in FIG. 2B, the processing pipeline 200b includes at least a hard drive 201, a host memory 203, a CPU 205, a video memory 207, and a GPU 209.
[0059] In a GPU-driven rendering pipeline, the GPU 209 determines the data required for rendering, loads the data from storage, and requests the CPU 205 to decode the data. In contrast to the processing pipeline 200a shown in FIG. 2A, the processing pipeline 200b moves the role of the decider function 211 to the GPU 209. This allows the GPU 209 to make decisions about what to load based on the scene and viewing angle. The decision is then read back to the CPU 205 to perform the loader function 213 and the decoder function 215. Finally, the decoded data is forwarded to the GPU to perform the consume function 219.
[0060] Generally, when the GPU 209 requests the CPU 205 to load data, processing the request introduces a latency of 1-3 frames because both the CPU 205 and the GPU 209 operate asynchronously and the data must be read back to the CPU 205 each time.
[0061] The processing pipelines 200a and 200b have several drawbacks. First, the GPU 209 must wait for the CPU to load and decode the data, which slows down the processing pipeline. Second, the performance-to-power ratio of the CPU 205 is lower than that of the GPU 209, so requesting the CPU 205 to perform the decoding consumes significant processing power. In addition, in a non-unified memory architecture (NUMA), the bandwidth to transfer the determined data from the host memory 203 to the video memory 207 is also high, which increases the power consumption.
[0062] Finally, GPU decompression algorithms focus on one type of data. For example, mesh compression can only compress a mesh and then decompress it into another mesh. Similarly, texture compression can only focus on texture data. There is a need for GPUs that can perform compression and decompression that are applicable to different types of data.
[0063] 3 illustrates a diagram of an example of a data loading pipeline handled by a GPU in accordance with one or more techniques of this disclosure. Specifically, FIG. 3 shows the entire processing pipeline handled by a GPU 309. In contrast to processing pipeline 200a from FIG. 2A and processing pipeline 200b from FIG. 2B, processing pipeline 300 shows loader function 313, decoder function 315, decider function 317, and consumer function 319, all performed by GPU 309.
[0064] Aspects of the present disclosure offload the loader function 213 and the decoder function 215 from the CPU 205 to the GPU 309. This allows the entire processing pipeline to run on the GPU 309 with minimal CPU involvement. Thus, there may be no exchange between the GPU and the CPU during the process of loading data for use in the GPU 309. Instead, the GPU may more efficiently decode the data in parallel and transfer the data in compressed form. In addition, compressed data is read directly into the GPU memory and the GPU 209 decodes the data into uncompressed data and consumes it, so there is no need to go back to the host memory 303 as processing occurs in the video memory 307. Bypassing the CPU bottleneck results in higher performance, less power consumption, and less bandwidth usage compared to other related systems involving a CPU.
[0065] Another benefit of the processing pipeline 300 is that it does not require modification of the current OS, APIs, and hardware. The framework introduced in the processing pipeline 300 may mainly include three parts. First, a file mapping function provided by the OS is used to swap the contents of the file into a memory block. Using the file mapping function means that no memory needs to be allocated since the file is mapped into the memory space. After the file is mapped into the memory block, the OS may load the contents asynchronously. Second, an API extension may be used to correlate the memory block to a graphics API buffer so that the GPU 309 can indirectly access the file through the memory block. The benefit of this is that no buffers need to be created. Instead, data on storage may be bound to the graphics API buffer. Third, building blocks are provided to build GPU decoding algorithms so that generic decompression (e.g., unzip) can be done in parallel on the GPU 309 to convert the encoded data into GPU consumable data. The benefit of this is to provide a flexible framework without having to stick to a specific algorithm for a specific file type.
[0066] Because the GPU 309 is not optimized for rendering from the mapped host memory, rendering from the mapped host memory may be slower than reading from the video memory 307. Therefore, the third part (e.g., building blocks for building GPU decoding algorithms) is used to decode the content from the mapped host memory 303 to the video memory 307. Here, the GP GPU technology is configured to convert the sequential CPU decoding algorithms into parallel GPU algorithms. For example, building blocks such as prefix sum scan, sorting, and matching may be converted into parallel GPU algorithms. After the data is decoded to the video memory 307, the GPU 209 may use the decoded data.
[0067] A non-limiting example of a tool that may be used to create a buffer from an existing memory address or block is the Vulkan extension group. Vulkan is a graphics API used in mobile and personal computer / console game development. Vulkan extensions are supported by various platforms and may map existing memory to a Vulkan buffer. Specifically, in this non-limiting example of using Vulkan extensions, and with respect to the second portion (e.g., correlating the memory block to the graphics API buffer), a Vulkan extension (e.g., VK_KHR_external_memory) may be used to associate the memory block with a Vulkan buffer. This allows the GPU to access the file indirectly through the Vulkan buffer. In Windows, a second Vulkan extension called memory host (e.g., VK_EXT_external_memory_host) may be used to create a buffer from an existing memory address. On Linux and Android, another Vulkan extension called Memory FD (e.g., VK_KHR_external_memory_fd) may create a buffer from a file descriptor.
[0068] These extensions allow file contents to be bound to a Vulkan buffer without explicit copying. Vulkan buffers are agnostic to the memory allocation method. Instead, by default, space is created within the Vulkan extension group as a buffer. This is in contrast to related systems where a buffer must be created before data can be delivered to the GPU. In these related systems, when data is delivered to the GPU, the data must first be placed in a buffer, copied, and then sent to the GPU. With this extension, an existing allocated buffer may be bound to a Vulkan buffer such that data can be sent directly to the GPU for consumption.
[0069] Continuing with the non-limiting example of using Vulkan extensions, the Vulkan extensions (e.g., VK_KHR_external_memory) used to correlate memory blocks to Vulkan buffers may include additional conditions for memory. The first condition is that the memory is aligned to VkPhysicalDeviceExternalMemoryHostPropertiesEXT::minImportedHostPointerAlignment. The second condition is that the size of the memory is a multiple of the alignment value. The third condition is that the memory status is in READ and WRITE modes even if the memory is only being read from it. In Windows, memory addresses from file mappings automatically follow the first and second conditions. The third condition may be met when creating the file, creating the mapping, and getting the address.
[0070] In addition, the associated decompression algorithm is designed without considering parallelism. Thus, data is decompressed serially. To increase the parallelism of processing to take advantage of the parallel processing architecture of the GPU 309, data may be reorganized into data chunks. For example, data may be reorganized into 4Kb per data chunk. Each data chunk is encoded and decoded independently with customizable algorithms. Although a GPU may have millions of threads, they are organized into groups such that each group always executes the same instruction simultaneously in a single instruction / multiple data (SIMD) manner. A group is called a WARP. A WARP is a collection of threads that are executed simultaneously by a streaming multiprocessor (SM). Each SM has a set of execution units, a set of registers, and a chunk of shared memory. Multiple warps may run on a SM at one time. In a GPU, a WARP may contain 16 to 128 threads depending on the hardware implementation.
[0071] Simply allocating threads to decode chunks would slow down the decoding due to branch deviations. Instead, WARPs are allocated to decode chunks and rely on data parallelism for some decoding steps. To minimize interference between decoding chunks, the decoding algorithm may be restricted to use only intra-WARP operations. Therefore, a WARP-based building block is provided in the framework for this purpose, which can also help to implement parallel decoding procedures. In some aspects, a parallelized unzip algorithm may be used as a default solution.
[0072] Some examples of WARP-based building block operations, such as prefix sum scanning, sorting, and matching, are described in more detail below. Each of these WARP-based building blocks plays an important role in the decoding algorithm and may be applicable to all types of compression and decompression scenarios.
[0073] The first WARP-based building block is the prefix-sum scan operation. In this framework, instead of adding block-based shared memory, the WARP shuffle function is used for inter-thread data exchange within WARP. The scan operation is important for transforming sequential algorithms into parallel algorithms, especially when the output length of each thread is variable.
[0074] The second WARP-based building block is the sorting operation. There may be two sorting algorithms. The first sorting algorithm is a radix sort on 32-bit integers. In a radix sort, the input data length may be less than or equal to the width of WARP. The second sorting algorithm is a merge sort on any data type, which may be optimized for data lengths less than 512 elements. This sorting algorithm is based on a merge pass and is adapted to the WARP scenario. Each thread in WARP holds N elements, so that the second sorting algorithm can sort (N*WARP-sized) elements in total. Synchronization operations and communication between threads may be modified to stay within WARP.
[0075] Another WARP-based building block is the matching operation. Many decompression algorithms use a matching process to copy from previously decompressed data. For example, "Abcd" may be encoded as "Abcd a[D=5, L=9]", where D is the distance and L is the length. For example, [D=5, L=9] corresponds to rewinding the pointer by 5 bytes and copying 9 bytes. In sequential decompression, this is a trivial task, since the bytes are copied one by one. However, when parallelizing, more data may be copied in parallel at once, and some bytes are generated during the copy. In this example, the first 5 bytes "bcd a" are copied because they are already there, but the next 4 bytes must be copied later. These 9 bytes cannot be copied at once in parallel. Therefore, the algorithm is iterative. At each iteration, the [D, L] region is divided into two, namely the non-overlapping region [D, min(L, D)] and the overlapping region [D+min(L, D), max(0, LD)]. Non-overlapping regions can be directly copied to become new [D,L] regions, and splitting continues until L is less than or equal to D. In the above example, [D=5,L=9] is split into [D=5,L=5] and [D=10,L=4]. Non-overlapping regions can be copied in parallel. Then [D=10,L=4] becomes a non-overlapping region, and the other copies take care of it. The non-overlapping length doubles with each iteration. Thus, parallelism doubles with each iteration.
[0076] The above-mentioned decoding building block operations are general and lossless. The framework may handle any data, such as video data, texture data, mesh data, neural network data, text data, etc. The framework may also be used on top of domain-specific compression, such as Adaptive Scalable Texture Compression (ASTC) texture compression or Draco geometry compression, to provide higher compression ratios. ASTC is a form of texture compression that uses variable block sizes rather than a single fixed size. ASTC is designed to essentially make most traditional compression formats obsolete by providing one format in addition to all the features of other compression formats. Draco is a library for compressing and decompressing 3D geometric meshes and point clouds, and is intended to improve the storage and transmission of 3D graphics. Additionally, instead of one piece of data per file, all game data may be put into a large file, and a virtual file system may be used to manage the data. This allows applications to correlate a large file to the GPU and then have the GPU load any data within it.
[0077] Thus, in processing pipeline 300, GPU 309 can decode data in parallel more efficiently than processing pipelines 200a and 200b because there is no communication with the CPU. In addition, GPU 309 can also transfer data in a compressed form, resulting in higher performance, lower power consumption, and less bandwidth usage than processing pipelines 200a and 200b.
[0078] 4 illustrates a flowchart of various example methods of data loading in a computing device according to one or more techniques of the present disclosure. The method 400 may be performed by an apparatus such as the control component 198, as described above. In some implementations, the method 400 is performed by processing logic including hardware, firmware, software, or a combination thereof. In some implementations, the method 400 is performed by a processor executing code stored in a non-transitory computer-readable medium (e.g., a memory). The method 400 includes loading data in the computing device.
[0079] At block 402, the method 400 includes identifying, at the GPU, data to load based on execution of the application program. In some aspects, the data may correspond to video data, texture data, mesh data, neural network data, or text data. For example, referring back to FIG. 3, the GPU 309 is configured to determine which data to load based on execution of the application program in the decider function 317.
[0080] At block 404, the method 400 includes loading, via the GPU, the data chunks of the identified data in encoded form from the data storage device into a video memory associated with the GPU. The identified data is organized into chunks during an offline (i.e., previously performed) compression phase. This organization of the data into data chunks allows the GPU to load the data chunks during runtime. In some aspects, loading the data into the video memory includes mapping a file of data (e.g., one or more files of data chunks) into memory blocks and associating the memory blocks mapped to the file of data with a buffer using an application program interface (API) extension group. For example, referring back to FIG. 3, the GPU 309 is configured in a loader function 313 to load the data in encoded form from the hard drive 301 into the video memory 307.
[0081] At block 406, the method 400 includes decoding the data chunks in parallel by using multiple GPU thread groups in parallel to decode the data chunks. Each of the data chunks is decoded independently of the other data chunks. For example, referring back to FIG. 3, the GPU 309 is configured to decode the data chunks in parallel at the decoder function 315.
[0082] In some aspects, each of the multiple GPU thread groups may decode a respective one of the data chunks independently from the other data chunks by employing an intra-group decoding function limited to using data from the respective thread group. In some aspects, decoding the data chunks in parallel may include decoding a first data chunk using a first GPU thread group of the multiple GPU thread groups in parallel and independently with decoding a second data chunk using a second GPU thread group of the multiple thread groups. In some aspects, the intra-group decoding function may include a shuffle function that enables inter-thread data exchange within one of the multiple GPU thread groups. In some aspects, the intra-group decoding function may include a radix sort of one or more threads in one of the multiple GPU thread groups, the radix sort having an input bit length that is less than or equal to the number of threads in one of the multiple GPU thread groups. In some aspects, the intra-group decoding function may include a merge sort applicable to a maximum number of elements equal to the product of the number of elements in one thread and the number of threads in the GPU thread group. In some aspects, the intra-group decoding function may include a matching operation that includes iteratively (i) copying a non-overlapping region split from an overlapping region of data, and (ii) splitting the overlapping region into a subsequent non-overlapping region and a subsequent overlapping region until the length of the subsequent overlapping region is less than the distance of the subsequent overlapping region.
[0083] The method 400 may include rendering the decoded data for visual presentation. For example, referring back to FIG. 3, the GPU 309 is configured to consume the decoded data by rendering the decoded data for visual presentation in a consuming function 319.
[0084] The subject matter described herein can be implemented to realize one or more benefits or advantages. For example, the techniques disclosed herein enable a method of data loading in a computing device, where the identification, reading, and decoding process is performed by the GPU instead of the CPU. As a result, the CPU bottleneck is bypassed to provide faster performance and lower latency. In addition, no modification of current OS, APIs, hardware, drivers, or existing components is required to utilize the method. Furthermore, the techniques disclosed herein provide a flexible framework such that it is not necessary to rely on specific algorithms for specific file types.
[0085] The subject matter described herein can be implemented to realize one or more benefits or advantages. For example, the described graphics processing and non-graphics processing techniques can be used by a server, client, GPU, or any other processor capable of computer processing or graphics processing to implement the sharing techniques described herein. This can also be accomplished at a low cost compared to other computer or graphics processing techniques. Furthermore, the computer processing or graphics processing techniques herein can improve or speed up data processing or execution. Furthermore, the computer processing or graphics processing techniques herein can improve resource or data utilization and / or resource efficiency.
[0086] In accordance with the present disclosure, the term "or" may be interpreted as "and / or" unless the context dictates otherwise. Furthermore, phrases such as "one or more" or "at least one" may be used with some features disclosed herein but not with other features, and features without such language may be interpreted as having such an implied meaning unless the context dictates otherwise.
[0087] In one or more examples, the functions described herein may be implemented in hardware, software, firmware, or any combination thereof. For example, although the term "processor" is used throughout this disclosure, such a processor may be implemented in hardware, software, firmware, or any combination thereof (e.g., by a processing circuit). If any function, processor, technique described herein, or other module is implemented in software, the function, processor, technique described herein, or other module may be stored or transferred as one or more instructions or code on a computer-readable medium. Computer-readable media may include computer data storage media or communication media, including any medium that facilitates transfer of a computer program from one place to another. In this manner, computer-readable media may generally correspond to (1) a tangible computer-readable storage medium that is non-transitory, or (2) a communication medium, such as a signal or carrier wave. Data storage media may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. By way of example and not limitation, such computer readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage. Disk and disc as used herein include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc, where disks typically reproduce data magnetically and discs reproduce data optically using a laser. Combinations of the above should also be included within the scope of computer readable media. A computer program product may include a computer readable medium.
[0088] The code may be executed by one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, application specific integrated circuits (ASICs), arithmetic logic units (ALUs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Thus, the term "processor" as used herein may refer to any of the foregoing structures, or any other structure suitable for implementing the techniques described herein. Also, the techniques may be implemented entirely in one or more circuits or logic elements.
[0089] The techniques of the present disclosure may be implemented in a wide variety of devices or apparatuses, including wireless handsets, integrated circuits (ICs), or sets of ICs, such as chipsets. In this disclosure, various components, modules, or units are described to highlight functional aspects of devices configured to perform the disclosed techniques, but do not necessarily require realization by different hardware units. Rather, as described above, the various units may be combined into any hardware unit, along with appropriate software and / or firmware, or may be provided by a collection of interoperating hardware units including one or more processors as described above. [Explanation of symbols]
[0090] 100 Data Loading System 104 Devices 111 Software Applications 113 OS 115 Graphics Driver 120 GPU 121 Video Memory 123 Decider function 124 system memory 125 Loader Function 126 Communication Interface 127 Decoder Function 128 CPU 130 Transmitter 131 Display 132 Transceiver 133 Receiver 135 Internal Memory 137 Content Encoder / Decoder 138 Frame Buffer 198 Control Components 200a, 200b, 300 Processing Pipeline 201 Hard Drive 203 Host Memory 205 CPU 207 Video Memory 209 GPU 211 Decider function 213 Loader Function 215 Decoder Function 217, 219 consumption function 301 Hard Drive 303 Host Memory 307 Video Memory 309 GPU 313 Loader Function 315 Decoder Function 317 Decider function 319 Consumption function 400 ways 402 Block 404 Block 406 Block
Claims
1. A data loading method in a computing device, comprising: in a graphics processing unit (GPU), identifying data to be loaded based on execution of an application program; loading, via the GPU, data chunks of the identified data in an encoded form from a data storage device to a video memory associated with the GPU; decoding the data chunks in parallel by using a plurality of GPU thread groups in parallel to decode the data chunks, each of the data chunks being decoded independently of other data chunks; A data loading method comprising the above steps.
2. The method according to claim 1, wherein each of the plurality of GPU thread groups decodes one of the data chunks independently of other data chunks by adopting an in-group decoding function limited to using data from the respective thread group.
3. The method according to claim 2, wherein the step of decoding the data chunks in parallel includes decoding a first data chunk using a first GPU thread group among the plurality of GPU thread groups and decoding a second data chunk using a second GPU thread group among the plurality of thread groups in parallel and independently.
4. The method according to claim 2, wherein the in-group decoding function includes a shuffle function that enables data exchange between threads within one of the plurality of GPU thread groups.
5. The method according to claim 2, wherein the in-group decoding function includes a radix sort of one or more threads within one of the plurality of GPU thread groups, the radix sort having an input bit length less than or equal to the number of threads within one of the plurality of GPU thread groups.
6. The method according to claim 2, wherein the in-group decoding function includes a merge sort applicable to a maximum number of elements equal to the product of the number of elements in one thread and the number of threads in a GPU thread group.
7. The in-group decoding function in claim 2 includes a matching operation that repeatedly includes: (i) copying a non-overlapping region divided from an overlapping region of the data; and (ii) dividing the overlapping region into a subsequent non-overlapping region and the subsequent overlapping region until the length of the subsequent overlapping region becomes smaller than the distance of the subsequent overlapping region.
8. The step of loading the data into the video memory includes: mapping a file of the data to a memory block; associating the memory block mapped to the file of the data with a buffer using an application program interface (API) extension group The method according to claim 1.
9. The data according to claim 1 corresponds to video data, texture data, mesh data, neural network data, or text data.
10. The method according to claim 1 further includes the step of rendering the decoded data for visual representation.
11. An apparatus for data loading in a computing device, the apparatus comprising: a graphics processing unit (GPU), identifying data to be loaded based on the execution of an application program; loading a data chunk of the identified data in an encoded form from a data storage device into a video memory associated with the GPU; parallel decoding of the data chunk by using a plurality of GPU thread groups in parallel to decode the data chunk, wherein each of the data chunks is decoded independently of other data chunks; A graphics processing unit (GPU) configured as An apparatus comprising.
12. Each of the plurality of GPU thread groups decodes each one of the data chunks independently of other data chunks by adopting an in-group decoding function limited to using data from the respective thread group. The apparatus according to claim 11.
13. Decoding the data chunks in parallel includes decoding a first data chunk using a first GPU thread group among the plurality of GPU thread groups and decoding a second data chunk using a second GPU thread group among the plurality of thread groups in parallel and independently, the apparatus according to claim 12.
14. The in-group decoding function includes a shuffle function that enables data exchange between threads within one of the plurality of GPU thread groups, the apparatus according to claim 12.
15. The in-group decoding function includes a radix sort of one or more threads within one of the plurality of GPU thread groups, the radix sort having an input bit length that is less than or equal to the number of threads within one of the plurality of GPU thread groups, the apparatus according to claim 12.
16. The in-group decoding function includes a merge sort applicable to a maximum number of elements equal to the product of the number of elements within one thread and the number of threads within a GPU thread group, the apparatus according to claim 12.
17. The in-group decoding function includes an equality operation that repeatedly includes (i) copying a non-overlapping region divided from an overlapping region of the data and (ii) dividing the overlapping region into a subsequent non-overlapping region and the subsequent overlapping region until the length of the subsequent overlapping region becomes smaller than the distance of the subsequent overlapping region, the apparatus according to claim 12.
18. Loading the data into the video memory includes mapping the file of the data to a memory block and associating the memory block mapped to the file of the data with a buffer using an application program interface (API) extension group the apparatus according to claim 11.
19. The data corresponds to video data, texture data, mesh data, neural network data, or text data, the apparatus according to claim 11.
20. When executed by at least one processor, the processor identifies data to be loaded based on the execution of an application program in a graphics processing unit (GPU) Load, via the GPU, data chunks of the identified data in encoded form from a data storage device into a video memory associated with the GPU; Parallelly decode the data chunks by using a plurality of GPU thread groups in parallel to decode the data chunks, each of the data chunks being decoded independently of other data chunks; A non-transitory computer-readable storage medium storing instructions for causing the above to be performed.
21. Identify data to be loaded based on the execution of an application program; Load data chunks of the identified data in encoded form from a data storage device into a video memory associated with a GPU; Parallelly decode the data chunks by using a plurality of GPU thread groups in parallel to decode the data chunks, each of the data chunks being decoded independently of other data chunks; A controller configured to: A device comprising the same.
22. The device according to claim 21, wherein each of the plurality of GPU thread groups decodes one of the data chunks independently of other data chunks by adopting an in-group decoding function limited to using data from the respective thread group.
23. The device according to claim 22, wherein parallelly decoding the data chunks includes decoding a first data chunk using a first GPU thread group among the plurality of GPU thread groups and decoding a second data chunk using a second GPU thread group among the plurality of thread groups in parallel and independently.
24. The device according to claim 22, wherein the in-group decoding function includes a shuffle function that enables data exchange between threads within one of the plurality of GPU thread groups.
25. The in-group decoding function includes a radix sort of one or more threads within one of the plurality of GPU thread groups, the radix sort having an input bit length that is less than or equal to the number of threads within one of the plurality of GPU thread groups. The device according to claim 22.
26. The in-group decoding function includes a merge sort applicable to a maximum number of elements equal to the product of the number of elements within one thread and the number of threads within the GPU thread group. The device according to claim 22.
27. The in-group decoding function includes a matching operation that iteratively includes (i) copying a non-overlapping region split from the overlapping region of the data, and (ii) splitting the overlapping region into a subsequent non-overlapping region and the subsequent overlapping region until the length of the subsequent overlapping region is less than the distance of the subsequent overlapping region. The device according to claim 22.
28. Loading the data into the video memory includes mapping the file of the data to a memory block, associating the memory block mapped to the file of the data with a buffer using an application program interface (API) extension group The device according to claim 22.
29. The data corresponds to video data, texture data, mesh data, neural network data, or text data. The device according to claim 22.
30. The controller is further configured to render the decoded data for visual representation. The device according to claim 21.
Citation Information
Patent Citations
Memory controller interface for micro-tiled memory access
CN102981961A
Image processor, image processing method, information processor, information processing system, semiconductor device and computer program
JP2004213641A
Memory controller interface for accessing microtiled memory
JP2008544426A
Graphics processor and graphics processing method
JP2015176492A
Rendering of graphics data using visibility information
JP2016509718A