Graphics processing unit and device having the same

The GPU optimizes graphics processing by selecting models based on device attributes and complexity, reducing tessellation load in graphics pipelines.

DE102015117768B4Active Publication Date: 2025-07-10SAMSUNG ELECTRONICS CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
DE102015117768
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2014-11-27
Filing Date
2015-10-19
Publication Date
2025-07-10
Estimated Expiration
2035-10-19

AI Technical Summary

Technical Problem

The operational load of graphics pipelines, particularly due to tessellation stages, is increased by the computational requirements of handling models with varying complexities in graphics processing.

Method used

A graphics processing unit (GPU) selects models based on the computing device's attributes and determines whether to perform tessellation by comparing model complexity with a reference, reducing redundant computations.

Benefits of technology

This approach reduces the operational load of the graphics pipeline by selectively performing tessellation only on models that require it, thereby optimizing resource utilization and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A graphics processing unit (GPU) (260) for determining whether to perform a tessellation operation on a first model of a plurality of provided models according to a control of a central processing unit (CPU) (210), comprising: an access circuit (252) configured to read the first model from a memory (310-1, 310-2), the memory (310-1, 310-2) storing the plurality of provided models having different complexities, to calculate a complexity of the first model, to compare the calculated complexity with a reference complexity, and to determine whether to perform a tessellation operation on the first model according to the comparison result, wherein the GPU (260) is configured to receive one or more addresses corresponding to memory areas of the memory (310-1, 310-2) storing the models from the CPU (210), to estimate at least one of a bandwidth between the GPU (260) and the memory (310-1, 310-2), a computing power of the GPU (260), a maximum power consumption of the GPU (260), a voltage and a frequency determined depending on dynamic voltage frequency scaling (DVFS) of the GPU (260), and a temperature of the GPU (260), to select a first address among the addresses based on an estimation result, to read the first model from a first memory area among the memory areas using the selected first address, to calculate the complexity of the first model using geometry information of the first model, and to determinewhether the tessellier operation should be performed on the first model, according to the result of comparing the calculated complexity with the reference complexity.
Need to check novelty before this filing date? Find Prior Art

Description

REGIONEmbodiments of the inventive concept relate to graphics processing, and more particularly, to a graphics processing unit (GPU= Graph Processing Unit= Graphik processing Unit) for reading one of predetermined models having different redundancies and determining whether to perform a tessellier operation on the model that has been read in response to a comparison of the complexity of the read model and a reference complexity.BACKGROUNDIn computer graphics, a level of detail (LOD=Level of Detail= Detail) involves setting detail based on geometry information such as depth values of control points or control points, respectively, or a curvature defined by the control points. In other words, LOD involves reducing the complexity of a three-dimensional (3D) object representation as it moves away from a viewer. LOD techniques increase rendering efficiency by reducing workload on graphics pipeline stages (e.g., vertex transforms).Among tessellier stages in a graphics pipeline, a tessellator is expected to perform a tessellier operation on an object. As a result, the operational load of the tessellator increases due to computational requirements.US 2012 / 0 169 728 A1 describes a method for determining the intersection points between a set of beams and a set of triangles. The method processes arbitrary rays and arbitrary primitives that provide lower complexity typical of ray tracing algorithms without using a spatial subdivision data structure.US 2013 / 0 265 309 A1 relates generally to a process for rendering graphics that includes performing vertex shading operations with a hardware shading unit of a vertex shading processing unit (GPU) to shading input vertices to output vertex shading output vertices, wherein the hardware unit is configured to receive a single vertex as input and generate a single vertex as output. The process also includes performing, with the hardware shader of the GPU, a geometry shading operation to generate one or more new vertices based on one or more of the vertex shaded vertices, the geometry shading operation processing at least one of the one or more vertex shaded vertices to output the one or more new vertices.U.S. Pat. No. 7,295,204 B2 discloses a method and a system for presenting bicubic surfaces of an object on a computer system. Each bicube surface is defined by sixteen control points and bounded by four boundary curves each corresponding to an edge, and each boundary curve is formed by a boundary frame of line segments formed between four of the control points. The method and system include transforming only those control points of the surface which represent a view of the object, rather than points across the entire bicubic surface, and using the four boundary edges for purposes of subdivision. Next, a pair of orthogonal boundary curves to be processed is selected. After the boundary curves are selected, each of the curves is iteratively divided and the pair of orthogonal internal curves, with two new curves being generated at each division. The division of each of the curves is ended when the curves meet a flatness threshold value expressed in screen coordinates, thereby minimizing the number of computations required for rendering the object.SUMMARYSome embodiments of the inventive concept provide a graphics processing unit (GPU) for selecting a model among models provided to have different redundancies before a rendering operation according to attributes or features of a computing device and determining whether to perform the tessellation on the selected model, thereby reducing an operating load of a graphics pipeline (e.g., tessellation), and devices having the same.Some embodiments of the inventive concept are set forth in the appended claims.BRIEF DESCRIPTION OF THE DRAWINGSThe above and other features and advantages of the inventive concept will become more apparent by describing exemplary embodiments thereof in detail with reference to the accompanying drawings, in which: FIG. 1 is a block diagram of a computing device according to some embodiments of the inventive concept; FIG. 2 is a block diagram of a central processing unit (CPU) and a graphics processing unit (GPU) illustrated in FIG. 1. FIG. 3 is a conceptual diagram for explaining a graphics pipeline of the GPU illustrated in FIG. 1, according to some embodiments of the inventive concept; FIG. 4 is a diagram of the structure of a first memory illustrated in FIG. 1 ; FIG. 5 is a conceptual diagram of models having different redundancies; FIG. 6 is a diagram of a method for calculating a complexity of a selected model according to some embodiments of the inventive concept; FIG. 7 is a diagram of a method for calculating a complexity of a selected model according to other embodiments of the inventive concept; FIG. 8 is a conceptual diagram of models having different redundancies; FIGS. 9A and 9B are diagrams illustrating the results of performing a tessellating operation on models having different redundancies; FIG. 10 is a flowchart of an operation of the computing device illustrated in FIG. 1 ; FIG. 11 is a diagram of a method for selecting a model among models having different redundancies according to some embodiments of the inventive concept; FIG. 12 is a diagram of a method for selecting a model among models having different redundancies according to other embodiments of the inventive concept; and FIG. 13 is a block diagram of a computing device including a graphics card according to some embodiments of the inventive concept.DETAILED DESCRIPTION OF THE EMBODIMENTSThe inventive concept will now be described more fully hereinafter with reference to the accompanying drawings, in which embodiments of the invention are shown. This invention may, however, be embodied in many different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the invention to those skilled in the art. In the drawings, the size and relative sizes of layers and regions may be exaggerated for clarity. Like numerals refer to like elements throughout.It will be understood that when an element is referred to as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or intervening elements may be present. In contrast, when an element is referred to as being "directly connected" or "directly coupled" to another element, there are no intervening elements present. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items and may be abbreviated " / ".It will be understood that although the terms first / first / first, second / second / second, etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are used only to distinguish one element from another. For example, a first signal could be termed a second signal, and similarly, a second signal could be termed a first signal without departing from the teachings of the disclosure.The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. As used herein, the singular forms "a / an" and "the / s" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising" or "includes" and / or "including," when used in this specification, specify the presence of stated features, regions, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, regions, integers, steps, operations, elements, components, and / or groups thereof.Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and / or the present application, and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.In a brief overview, embodiments of the present inventive concept have a coarse level of detail (LOD=level of detail=level of detail) geometry scheme, whereby the nearest LOD geometry is selected in a coarse LOD geometry map, and a calculation is made that requires only one additional geometry based on the nearest LOD geometry. In this way, the same LOD geometry can be reused without redundant computations. Also, a tessellier render operating load, which is widely used in current configurations, can be reduced by applying the closest LOD geometry.FIG. 1 is a block diagram of a computing device 100 according to some embodiments of the inventive concept. The computing device 100 may be implemented in an electronic device as, for example, a television (TV) (for example, a digital TV or a smart TV), a personal computer (PC), a desktop computer, a laptop computer, a computer workstation, a tablet PC, a video game platform (or a video game console), a server, or a portable electronic device. A portable electronic device implementing embodiments may include, but is not limited to, a mobile phone, a smart phone, a personal digital assistant (PDA= Personal Digital Assistant= Persönliche Digital Assistant), an enterprise digital assistant (EDA), a digital still camera, a digital video camera, a portable multimedia player (PMP= Portable Multimedia Player), a personal navigation device or a portable navigation device (PND= Portable Navigation Device= Tragbare Navigation Device), a mobile internet device (MID= Mobile Internet Device= Mobile Internet Device), a portable computer, an internet of things (IoT) device, an Internet of Everything (IoE) device or an E-book.The computing device 100 may include various types of devices capable of processing and displaying two-dimensional (2D) or three-dimensional (3D) graphics data. The computing device 100 includes a system-on-chip (SoC=system-on-chip=system-on-chip) 200, one or more memories 310- 1 and 310- 2, and a display 400.The SoC 200 may function as a host of the computing device 100. The SoC 200 may control the overall operation of the computing device 100. For example, the SoC 200 may be replaced with an integrated circuit (IC= Integr Circuit= Integrierte), an application processor (AP= Appli Processor= Anwendungs Processor), or a mobile AP that may perform an operation described hereinafter in the embodiments of the inventive concept, that is, an operation for determining whether to perform a tessellating-related operation on a model.The SoC 200 may include a central processing unit (CPU= C Processing Unit= Zentrale Processing Unit) 210, one or more memory controllers 220- 1 and 220- 2, a user interface 230, a display controller 240, and a graphics processing unit (GPU= Graph Processing Unit= Graphik Processing Unit) 260, which is also referred to as a graphics processor, which may communicate with each other via a bus 201. The bus 201 may be implemented as a peripheral component interconnect (PCI) bus, a PCI express (PCIe) bus, an advanced microcontroller bus architecture (AMBA™), an advanced high-performance bus (AHB), an advanced peripheral bus (APB), an advanced extensible interface (AXI), or a combination thereof.The CPU 210 may control the operation of the SoC 200. According to some embodiments, the CPU 210 may estimate (i.e., calculate or measure) at least one of the attributes or features of the computing device 100, select one of addresses of storage areas included in the first memory 310- 1 storing a plurality of provided models based on the result of the estimation (i.e., calculation or measurement), and transmit the selected address to the GPU 260. The attributes or features may include at least one of a bandwidth between the GPU 260 and the memory 310- 1 or 310- 2, a calculation power of the GPU 260, a maximum power consumption of the GPU 260, a voltage and a frequency determined depending on a dynamic voltage frequency scaling (DVFS=dynamic voltage frequency scaling) of the GPU 260, and a temperature of the GPU 260.The SoC 200 may include a software (firmware) component and / or a hardware component 205 that estimates (i.e., calculates or measures) at least one of the bandwidth between the GPU 260 and the memory 310- 1 or 310- 2, the calculation power of the GPU 260, the maximum power consumption of the GPU 260, the voltage and the frequency determined depending on a DVFS of the GPU 260, and the temperature of the GPU 260.The software component may be stored in a hardware memory and executed by one or more processors of the SoC 200 and / or one or more processors external to the SoC 200. For example, the software component may be executed by the CPU 210. At least one hardware component 205 can have at least one computer, detector and / or sensor. The software component (e.g., an application 211 in FIG. 2 ) executing within the CPU 210 may estimate (i.e., calculate or measure) at least one of the attributes or features of the computing device 100 using an output signal output by the at least one hardware component 205.The GPU 260 may receive an address from the CPU 210, read a model from one of the memory areas using the address, for example, calculate a complexity of the model using geometry information of the model using the access circuit 252 or another processor, compare the complexity with a reference complexity, and determine whether or not to perform tessellation on the model based on the comparison result.Alternatively, the CPU 210 may transmit a plurality of addresses of a plurality of storage areas in the memory 310- 1 or 310- 2 in which a plurality of models are stored to the GPU 260. The GPU 260 may receive the plurality of addresses from the CPU 210, estimate (i.e., calculate or measure) at least one of the attributes or features of the computing device 100, and select one of the addresses based on the estimation (i.e., calculation or measurement) result. The GPU 260 may read a model from one of the memory regions using the selected address, calculate a complexity of the model using geometry information of the model, compare the complexity with a reference complexity, and determine whether or not to perform tessellation on the model based on the comparison result.In embodiments where the computing device 100 is a portable electronic device, the computing device 100 may also include a battery 203. The at least one hardware component 205 and / or the software component (e.g., application 211 in FIG. 2 ) executed in CPU 210 may estimate (i.e., calculate or measure) a remaining value of battery 203. In this case, at least one attribute or feature of the computing device 100 may include the remaining value of the battery 203 or information about the remaining value.A user may input an input to the SoC 200 via the user interface 230 such that the CPU 210 executes at least one application (e.g., the software application 211 in FIG. 2 ). The at least one application executed by the CPU 210 may include an operating system (OS=Oper System= Betriebssystem), a word processor application, a media player application, a video game application, and / or a graphical user interface (GUI= Graph User Interface= Grafische User Interface) application.A user may input an input to the SoC 200 via an input device (not shown) connected to the user interface 230. For example, the input device may be implemented as a keyboard, a mouse, a microphone, or a touchpad. An application (e.g., 211 in FIG. 2 ) executed by the CPU 210 may include graphics rendering commands that may be related to a graphics application programming interface (API= Appli Programming Interface= Anwendungs Programming Interface).A graphics API may include one or more of an open graphics library (OpenGL®) API, open graphics library for embedded systems (open GL ES) API, DirectX API, renderscript API, WebGL API, or open VG® API. To process graphics render commands, CPU 210 may transmit a graphics render command to GPU 260 through bus 201. The GPU 260 may process (or render) graphics data in response to the graphics render command.The graphics data may include points, lines, triangles, quadrilaterals, patches, and / or primitives. The graphic data may also include line segments, elliptical arcs, quadratic Bezier curves, and / or cubic Bezier curves.The one or more memory controllers 220- 1 and 220- 2 may read data (e.g., graphics data) from one or more memories 310- 1 and 310- 2 in response to a read request from the CPU 210 or the GPU 260, and may transmit the data (e.g., the graphics data) to a corresponding element (e.g., 210, 240, or 260). The one or more memory controllers 220- 1 and 220- 2 may write data (e.g., graphics data) from the corresponding element (e.g., 210, 230, or 240) to the one or more memories 310- 1 and 310- 2 in response to a write request from the CPU 210 or the GPU 260.The one or more memory controllers 220- 1 and 220- 2 are separate from the CPU 210 or the GPU 260 in the embodiments illustrated in FIG. 1. However, the one or more memory controllers 220- 1 and 220- 2 may be in the CPU 210, the GPU 260, or the one or more memories 310- 1 and 310- 2.When the first memory 310- 1 is formed with a volatile memory and the second memory 310- 2 is formed with a nonvolatile memory, the first memory controller 220- 1 may be implemented to communicate with the first memory 310- 1, and the second memory controller 220- 2 may be implemented to be capable of communicating with the second memory 310- 2. The volatile memory may be a random access memory (RAM= Random Access Memory= Direkt), a static RAM (SRAM=Static RAM=Static RAM), a dynamic RAM (DRAM=Dynamic RAM=Dynamic RAM), a synchronous DRAM (SDRAM=Synchronous DRAM=synchronous DRAM), a thyristor RAM (T-RAM= Thyristor RAM= Thyristor RAM), a zero-capacitance RAM (Z-RAM= Zer Capacitor RAM= Null-capacitance RAM), or a twin transistor RAM (TTRAM=Win Transistor RAM=Win transistor RAM). The nonvolatile memory may be an electrically erasable programmable read-only memory (EEPROM= El Erasable Programmable read-only memory), a flash memory, a magnetic RAM (MRAM= Magnet RAM= Magnetischer RAM), a spin transfer torque MRAM, a ferroelectric RAM (FeRAM= Ferro RAM= Ferroelektrisch RAM), a phase transition RAM (PRAM=Phase-Change RAM=Phase-Transition RAM), or a resistive RAM (RRAM=Resisive RAM=resissive RAM). The non-volatile memory may also be implemented as a multimedia card (MMC= Multi Card= Multimedia card), an embedded MMC (eMMC= Em MMC= Eingebettete MMC), a universal flash memory (UFS= Universal Flash Storage= Universell flash memory), a solid state drive (SSD=Solid State Drive=Solid state Drive) or a universal serial bus (USB= Universal Serial Bus= Universal serial bus) flash drive.The one or more memory controllers 220- 1 and 220- 2 may store a program (or an application) or instructions that may be executed by the CPU 210. In addition, the one or more memory controllers 220- 1 and 220- 2 may store data to be used by a program executed by the CPU 210. The one or more memory controllers 220- 1 and 220- 2 may also store a user application and graphics data related to the user application, and may store data (or information) to be used or generated by the components included in the SoC 200. The one or more memory controllers 220- 1 and 220- 2 may store data used for operation of the GPU 260 and / or data generated by operation of the GPU 260. The one or more memory controllers 220- 1 and 220- 2 may store command streams (command streams) for the process of the GPU 260.The display controller 240 may transmit data processed by the CPU 210 or data (for example, graphics data) processed by the GPU 260 to the display 400. The display 400 may be implemented as a monitor, a TV monitor, a projection device, a thin film transistor liquid crystal display (TFT LCD=thin film transistor liquid crystal display=thin film transistor liquid crystal display), a light emitting diode (LED=light emitting diode=light emitting diode) display, an organic LED (OLED) display, an active matrix OLED (AMOLED) display, or a flexible display.The display 400 may be integrated with (or embedded in) the computing device 100. The display 400 may be a screen of a portable electronic device or a stand-alone device connected to the computing device 100 via a wireless or wired communication link. Alternatively, the display 400 may be a computer monitor connected to a PC via a cable or a wired connection.The GPU 260 may receive commands from the CPU 210 and execute the commands. Instructions executed by the GPU 260 may include a graphics instruction, a memory transfer instruction, a kernel execute instruction, a tessellier instruction, or a texturing instruction. The GPU 260 may perform graphics operations to render graphics data.When an application executed by the CPU 210 requests to perform graphics processing, the CPU 210 may transmit graphics data and a graphics command to the GPU 260 such that the graphics data is rendered on the display 400. The graphics command may include a tessellating command and / or a texturing command. The graphics data may include vertex data, texture data, or surface data. A surface may include a parametric surface, a partitioning surface, a triangle mesh, or a curve.The CPU 210 may transmit a graphics command and graphics data to the GPU 260 in some embodiments. In other embodiments, when the CPU 210 writes a graphics command and graphics data to the one or more memories 310- 1 and 310- 2, the GPU 260 may read the graphics command and the graphics data from the one or more memories 310- 1 and 310- 2.The GPU 260 may directly access a GPU cache 290. Here, the GPU 260 may write or read graphics data to or from the GPU cache 290 without using the bus 201. GPU cache 290 is an example of GPU memory that can be accessed by GPU 260.The GPU 260 and the GPU cache 290 are separate from each other in the embodiments illustrated in FIG. 1. In other embodiments, the GPU 260 may include the GPU cache 290. The GPU cache 290 may be formed including the DRAM or SRAM or the like. The CPU 210 or the GPU 260 may store processed (or rendered) graphics data in a frame buffer included in the one or more memories 310- 1 and 310- 2, respectively.FIG. 2 is a block diagram of the CPU 210 and the GPU 260 illustrated in FIG. 1. Referring to FIG. 2, the hardware component 205, the CPU 210, and the GPU 260 may communicate with each other via the bus 201. In some embodiments, the hardware component 205, the CPU 210, and the GPU 260 may be integrated into a motherboard or an SoC, or may be implemented in a graphics card installed in a motherboard. In a computing device 100A illustrated in FIG. 13, the hardware component 205 may be implemented in a motherboard or graphics card.The CPU 210 may include one or more applications (e.g., software applications) 211, a graphics API 213, a GPU driver 215, and an OS 217. The CPU 210 may operate or execute the components 211, 213, 215, and 217.The application 211 may include instructions for displaying graphics data and / or instructions to be executed in the GPU 260. For example, the application 211 may estimate (i.e., calculate or measure) at least one of the at least one attribute or feature of the computing device 100, and may transmit an address associated with one of the plurality of models or addresses associated with the plurality of models to the GPU 260 based on the estimation (i.e., the calculation or measurement result). At this time, the application 211 may use, process, and / or refer to an output signal of the hardware component 205.In some embodiments, application 211 issues the commands to graphics API 213. The graphics API 213 may convert the commands received from the application 211 into a format used by the GPU driver 215.The GPU driver 215 may receive the commands through the graphics API 213 and may control the operation of the GPU 260 such that the commands are executed by the GPU 260. For example, GPU driver 215 may transmit commands to GPU 260 via OS 217 or may transmit the commands to the one or more memories 310- 1 and 310- 2 that may be accessed by GPU 260. The GPU 260 may include a command decoder (or a command engine) 251, an access circuit 252, and one or more processing units 253.The command decoder 251 may receive a command from the CPU 210 or a command received via the one or more memories 310- 1 and 310- 2, and may control the GPU 260 to execute the command or according to the command. For example, the command decoder 251 may receive an address associated with one of a plurality of models or a plurality of addresses connected to the plurality of models from the CPU 210, and may transmit the address or the plurality of addresses to the access circuit 252. The access circuit 252 may read a model from the first memory 310- 1 using the address transmitted from the CPU 210 or an address transmitted from one of the processing units 253, and may transmit the model read to a graphics pipeline (for example, 260A in FIG. 3 ).The processing units 253 may include a programmable processing unit, a fixed function processing unit, and an estimation unit that estimates (i.e., calculates or measures) at least one of the at least one attribute or feature of the computing device 100 and transmits the estimation result to the access circuit 252. For example, the programmable processing unit may be a shader programmable unit that can execute at least one shader program. The programmable shader unit may be downloaded from the CPU 210 to the GPU 260. Programmable shader units operating in the processing units 253 may include one or more of a vertex shader unit, a hull shader unit, a domain shader unit, a geometry shader unit, a pixel shader unit (or fragment shader unit), and / or a unified shader unit.The fixed function processing unit may include a hardware component, for example, hardware component 205 described herein. The hardware component may be hardwired to perform certain functions. For example, the fixed function processing unit may include processing units that perform raster operations among the processing units 253. For example, processing units 253 may form a 3D graphics pipeline. For example, the 3D graphics pipeline may conform to OpenGL® API, Open GL ES API, DirectX API, Renderscript API, WebGL API, or Open VG® API.FIG. 3 is a conceptual diagram for explaining a graphics pipeline 260A of the GPU 260 illustrated in FIG. 1, according to some embodiments of the inventive concept. The graphics pipeline 260A, which may be executed in the GPU 260, may correspond to a graphics pipeline in Microsoft® DirectX 11.The graphics pipeline 260A may include a plurality of processing stages performed at one or more processing units 253 illustrated in FIG. 2 and a resource block 263. The processing stages (or processing units 253) may include some or all of, but are not limited to, an input assembler 261- 1, a vertex shader 261- 2, a hull shader 261- 3, a tessellator 261- 4, a domain shader 261- 5, a geometry shader 261- 6, a rasterizer 261- 7, a pixel shader 261- 8, and an output combiner (output messenger) 261- 9.The hull shader 261- 3, the tessellator 261- 4, and the domain shader 261- 5 may form tessellating stages of the graphics pipeline 260A. Accordingly, the tessellating stages may perform a tessellating (or a tessellating operation). The pixel shader 261- 8 may be referred to as a fragment shader. For example, the input assembler 261- 1, the tessellator 261- 4, the rasterizer 261- 7, and the output assembler 261- 9 are fixed function stages. The vertex shader 261- 2, the hull shader 261- 3, the domain shader 261- 5, the geometry shader 261- 6, and the pixel shader 261- 8 are programmable stages.The programmable stages have a structure in which a certain type of shader program can be performed. For example, the vertex shader 261- 2 may execute a vertex shader program, the hull shader 261- 3 may execute a hull shader program, the domain shader 261- 5 may execute a domain shader program, the geometry shader 261- 6 may execute a geometry shader program, and the pixel shader 261- 8 may execute a pixel shader program. Each shader program may be executed in a shader unit of the GPU 260 at an appropriate timing.Different shader programs may be executed in a common shader (or shared shader unit) of the GPU 260. For example, the common shader may be a unified shader. In other embodiments, at least one dedicated shader can exclusively execute at least one particular type of shader program.The input assembler 261- 1, the vertex shader 261- 2, the hull shader 261- 3, the domain shader 261- 5, the geometry shader 261- 6, the pixel shader 261- 8, and the output assembler 261- 9 may directly or indirectly communicate data with the resource block 263 via one or more connectors or interfaces. Accordingly, the input assembler 261- 1, the vertex shader 261- 2, the hull shader 261- 3, the domain shader 261- 5, the geometry shader 261- 6, the pixel shader 261- 8, and the output combiner 261- 9 may retrieve or receive input data from the resource block 263. The geometry shader 261- 6 and the output combiner 261- 9 may each have an output for writing output data to the resource block 263.The communication between each of the components 261- 1 to 261- 9 and the resource block 263 illustrated in FIG. 3 is merely an example and may be modified in various ways. Thus, except for the operation of the components (e.g., 261- 1 and 261- 4) for the embodiments of the inventive concept, the operation of the components 261- 1 to 261- 9 is substantially the same as or similar to that of a graphics pipeline defined in Microsoft® DirectX 11. Thus, detailed descriptions thereof will be omitted.The input assembler 261- 1 provides graphics-related data (e.g., triangles, lines, and / or points) to the graphics pipeline 260A. A model read by the access circuit 252 may be provided for the input assembler 261- 1. The input assembler 261- 1 reads data (e.g., triangles, lines, and / or points) from the resource block 263 and assembles the data into primitives that may be used at other processing stages. The input assembler 261- 1 assembles vertices into different types of primitives (e.g., line lists, triangle strips, or primitives with neighborhood of the data in memory).The vertex shader 261- 2 processes (e.g., performs transformation, skinning, morphing, and per-vertex operation such as per-vertex lighting) the assembled vertices output from the input assembler 261- 1. Vertex shader 261- 2 takes or operates a single input vertex and generates a single output vertex.The hull shader 261- 3 transforms input control points, which are output from the vertex shader 261- 2 and define a low-order surface, into output control points that form a patch. The hull shader 261- 3 may perform per-patch computations to provide data to the tessellator 261- 4 and domain shader 261- 5.For example, the hull shader 261- 3 may receive input control points from the vertex shader 261- 2, generate output control points (which may be the same as the input control points), patch constant data and tessellation factors regardless of the number of tessellation factors, output the output control points and tessellation factors to the tessellator 261- 4, and output the patch constant data and tessellation factors to the domain shader 261- 5 for processing.The tessellator 261- 4 may subdivide a domain (e.g., a quadrangle, a triangle, or a lint) into smaller objects (e.g., triangles, points, or lines) using the output control points and the tessellating factors. The domain shader 261- 5 calculates vertex positions of the output control points output from the hull shader 261- 3 and vertex positions of the divided points of a patch output from the tessellator 261- 4.FIG. 4 is a diagram of the structure of the first memory 310- 1 of the computing device 100 illustrated in FIG. 1. Referring to FIGS. 1 and 4, the first memory 310- 1 may be accessed by the GPU 260, and may include a plurality of memory regions MEM 1 to MEM 5. Models MOD 1 to MOD 5 each having different redundancies are stored in the memory areas MEM 1 to MEM 5 respectively in advance before a rendering operation. The memory areas MEM1 to MEM5 may be selected or defined by addresses ADD1 to ADD5, respectively. Here, the term "model" may be a superordinate concept for an object.FIG. 5 is a conceptual diagram of the models MOD 1 to MOD 5 having different redundancies. Referring to FIGS. 4 and 5, it is assumed that the first model MOD 1 has the lowest complexity and the fifth model MOD 5 has the highest complexity. In other words, the complexity of the second model MOD 2 is higher than that of the first model MOD 1, the complexity of the third model MOD 3 is higher than that of the second model MOD 2, the complexity of the fourth model MOD 4 is higher than that of the third model MOD 3, and the complexity of the fifth model MOD 5 is higher than that of the fourth model MOD 4. Models MOD1 to MOD5 may require different amounts of computational power, respectively, for example needed for graphics processing and depending on complexity.Although five models MOD 1 to MOD 5 having different redundancies are illustrated in FIGS. 4 and 5, these are merely examples. In other embodiments, at least one model may exist between corresponding two models MOD1 and MOD2, MOD2 and MOD3, MOD3 and MOD4, or MOD4 and MOD5.Referring to FIGS. 1 to 5, the CPU 210 may estimate (i.e., calculate or measure) at least one of one or more attributes or features of the computing device 100, select one of the addresses ADD 1 to ADD 5 corresponding to the storage areas MEM 1 to MEM 5 of the first memory 310- 1, respectively, storing the models MOD 1 to MOD 5 according to the estimation (i.e., calculation or measurement) result, and transmit the selected address to the GPU 260.As described above, the one or more attributes or features may include the bandwidth between the GPU 260 and the memory 310- 1, the calculation power of the GPU 260, the maximum power consumption of the GPU 260, the voltage and frequency determined depending on DVFS of the GPU 260, and / or the temperature of the GPU 260, etc.For example, in embodiments where the computing device 100 is a desktop computer or a computer workstation, the CPU 210 may transmit the fifth address ADD 5 corresponding to the fifth memory area MEM 5 storing the fifth model MOD 5 having the highest complexity to the GPU 260 according to the estimation (i.e., calculation or measurement) result.However, when the computing device 100 is a portable electronic device such as a smartphone, the CPU 210 may transmit the first address ADD 1 corresponding to the first storage area MEM 1 storing the first model MOD 1 having the lowest complexity to the GPU 260 according to the estimation (i.e., calculation or measurement) result. In embodiments where computing device 100 is a portable electronic device that includes a battery 203, the attributes or features of computing device 100 may include the remaining value of battery 203, respectively, which may be displayed in some embodiments.Accordingly, the CPU 210 may transmit the second address ADD 2 corresponding to the second model MOD 2 to the GPU 260 when the remaining value of the battery 203 is a first value, and may transmit the first address ADD 1 corresponding to the first model MOD 1 to the GPU 260 when the remaining value of the battery 203 is a second value. At this time, the first value may be higher than the second value.The GPU 260 may receive the address from the CPU 210, read a model from one of the memory areas MEM 1 to MEM 5 using the address, calculate a complexity of the model using geometry information of the model, compare the complexity with a reference complexity, and determine whether or not to perform tessellation on the model according to the comparison result.Alternatively, the CPU 210 may transmit a list of addresses ADD 1 to ADD 5 corresponding to the memory areas MEM 1 to MEM 5, respectively, storing the models MOD 1 to MOD 5, respectively, to the GPU 260. The GPU 260 may receive the addresses ADD 1 to ADD 5 from the CPU 210, estimate (i.e., calculate or measure) at least one of one or more attributes or features of the computing device 100, and may select one of the addresses ADD 1 to ADD 5 related to the models MOD 1 to MOD 5 according to the estimation (i.e., calculation or measurement) result.As described above, the selected address may be transmitted to the access circuit 252. The GPU 260, and more specifically, the access circuit 252 may read a model from one of the memory areas MEM 1 to MEM 5 using the selected address, calculate a complexity of the model using geometry information of the model, compare the complexity with a reference complexity, and determine whether or not to perform tessellation on the model according to the comparison result.A model read using the resource block 263 may be transmitted to the input assembler 261- 1. For example, the access circuit 252 of the GPU 260 may read the model stored in a storage area corresponding to an address received from the CPU 210 or selected by the GPU 260 from the first memory 310- 1 or the resource block 263 using the address.The geometry information may include a depth value of one or more control points or a curvature defined by the control points included in the model. For example, the hull shader 261- 3, which is part of the tessellating stage of the graphics pipeline 260A, may calculate a complexity of the model using the geometry information of the model, compare the complexity with a reference complexity, and transmit comparison information INF corresponding to the comparison result to the tessellator 261- 4.The tessellator 261- 4 may receive data regarding the model and the comparison information INF, and determine whether to perform tessellation on the model based on the comparison information INF. In other words, the tessellator 261- 4 may perform tessellation on the model or alternatively pass the model on to the domain shader 261- 5 as it is.FIG. 6 is a diagram of a method for calculating a complexity of a selected model according to some embodiments of the inventive concept. The GPU 260, and more specifically the hull shader 261- 3, described herein in some embodiments may calculate the complexity of the model using geometry information, i.e., depth values of control points of the model, and perform a comparison.The CPU 210 or the GPU 260 may generate an address that activates (or reads) a model closest to a desired one by the computing device 100 to be read based on the attributes of the computing device 100. Accordingly, the GPU 260 reads the model closest to the one desired by the computing device 100 from the first memory 310- 1 using the address.FIG. 6 is an exemplary conceptual diagram provided for convenience in description. Referring to a first case CASE 1, when a model read by the GPU 260 has a square patch including four control points P 11, P 12, P 13, and P 14, and the depth values of the control points P 11, P 12, P 13, and P 14 are less than a reference depth value RDEP (for example, when the patch is relatively close to a viewer), the hull shader 261- 3 may transmit the comparison information INF that commands to perform tessellation on the model to the tessellator 261- 4. Accordingly, the tessellator 261- 4 may perform a tessellating operation on the model received from the hull shader 261- 3.However, referring to a third case CASE 3, when a model read by the GPU 260 is a quadrangular patch having four control points P 31, P 32, P 33, and P 34, and the depth values of the control points P 31, P 32, P 33, and P 34 are larger than the reference depth value RSEP (for example, when the patch is relatively far from the viewer), the hull shader 261- 3 may transmit the comparison information INF that commands to pass the model to the tessellator 261- 4.Accordingly, the tessellator 261- 4 does not perform the tessellating on the model received from the hull shader 261- 3. Since the tessellator 261- 4 does not perform a tessellating operation, a processing-intensive operating load that is otherwise generated at the tessellator 261- 4 is reduced. As a result, the operational load of the graphics pipeline 260A is reduced.Referring to a second case CASE 2, when a model read by the GPU 260 is a quadrangular patch having four control points P 21, P 22, P 23, and P 24 and the depth value of only the control point P 21 among the control points P 21, P 22, P 23, and P 24 is less than the reference depth value RDEP, then the hull shader 261- 3 may transmit comparison information INF to the tessellator 261- 4 that commands the tessellator 261- 4 to perform a tessellating operation on the model or comparison information INF that commands to pass the model according to a program that has been set.As an alternative, the GPU 260 may read a model that is of higher complexity than that desired by the computing device 100. For example, if the model desired by the computing device 100 is the first model MOD 1 and the model read by the GPU 260 is the third, fourth, or fifth model MOD 3, MOD 4, or MOD 5, then the hull shader 261- 3 may transmit to the tessellator 261- 4 the comparison information INF that commands to pass one or more models MOD 3, MOD 4, or MOD 5, even in any of the cases CASE 1, CASE 2, and CASE 3.As another alternative, the GPU 260 may read a model that is less complex than the one desired by the computing device 100. For example, if the model desired by the computing device 100 is the fifth model MOD 5 and the model read by the GPU 260 is the first, third, or fourth model MOD 1, MOD 3, or MOD 4, then the hull shader 261- 3 may transmit to the tessellator 261- 4 the comparison information INF that commands to perform a tessellating operation on the model MOD 1, MOD 3, or MOD 4 even in any of the above-mentioned cases CASE 1, CASE 2, and CASE 3.FIG. 7 is a diagram of a method for calculating a complexity of a selected model according to other embodiments of the inventive concept. Referring to FIGS. 3-5 and 7, the GPU 260, and more specifically the hull shader 261- 3, may calculate the complexity of the model using geometry information of the model, for example, a curvature defined by a set of control points P 41, P 42, P 43, and P 44, and perform a comparison.It is assumed that the CPU 210 or the GPU 260 generates an address that enables a model closest to one desired by the computing device 100 to be read (or read) on the attributes of the computing device 100. Accordingly, the GPU 260 reads the model closest to the one desired by the computing device 100 from the first memory 310- 1 using the address.When a curvature CV 1 defined by control points P 41, P 42, P 43, and P 44 included in the model read by the GPU 260 is larger than a reference curvature RCV, the hull shader 261- 3 may transmit, to the tessellator 261- 4, the comparison information INF that commands to perform a tessellating operation on the model. However, if a curvature CV 2 defined by the control points P 41, P 42, P 43, and P 44 included in the model read by the GPU 260 is less than the reference curvature RCV, then the hull shader 261- 3 may transmit, to the tessellator 261- 4, the comparison information INF that instructs to pass the model.Alternatively, the GPU 260 may read a model that is of higher complexity than the one desired by the device 100. For example, if the model desired by computing device 100 is first model MOD 1 and the model read by GPU 260 is third, fourth, or fifth model MOD 3, MOD 4, or MOD 5, then hull shader 261- 3 may transmit to tessellator 261- 4 comparison information INF that commands to pass model MOD 3, MOD 4, or MOD 5 regardless of the calculated curvature.Alternatively, the GPU 260 may read a model that is less complex than the one desired by the computing device 100. For example, if the model desired by the computing device 100 is the fifth model MOD 5 and the model read by the GPU 260 is the first, third, or fourth model MOD 1, MOD 3, or MOD 4, then the hull shader 261- 3 may transmit to the tessellator 261- 4 the comparison information INF that commands to perform a tessellating operation on at least one model MOD 1, MOD 3, or MOD 4 regardless of the calculated curvature.FIG. 8 is a conceptual diagram of models MOD 11, MOD 12, MOD 13, and MOD 14 having different redundancies. Referring to FIGS. 4 and 8, the models MOD 11, MOD 12, MOD 13, and MOD 14 having different redundancies may be stored in the memory areas MEM 1 to MEM 4, respectively. A model, for example, MOD1 or MOD11 is stored in the first memory area MEM1. A model, for example, MOD2 or MOD12 is stored in the second memory area MEM2. A model such as MOD3 or MOD13 is stored in the third memory area MEM3. A model such as MOD4 or MOD14 is stored in the fourth memory area MEM4.The number M 4 of patches (e.g., triangles) included in the ninth model MOD 14 is larger than the number M 3 of patches (e.g., triangles) included in the eighth model MOD 13. The number M 3 of patches (e.g., triangles) included in the eighth model MOD 13 is larger than the number M 2 of patches (e.g., triangles) included in the seventh model MOD 12. The number M 2 of patches (e.g., triangles) included in the seventh model MOD 12 is larger than the number M 1 of patches (e.g., triangles) included in the sixth model MOD 11.Although a patch is a triangle in the embodiments illustrated in FIG. 8, the patch may be a quadrangle in other embodiments. The number of patches included in a model is linked to the complexity of the model. In detail, the greater the number of patches included in a model, the higher the complexity of the model. Each of the models MOD 11, MOD 12, MOD 13, and MOD 14 is transmitted to the hull shader 261- 3 via the input assembler 261- 1 and the vertex shader 261- 2. The GPU 260, and more specifically, the access circuit 252, may read any one of the models MOD 11, MOD 12, MOD 13, and MOD 14 stored in the first memory 310- 1 based on an address selected by the CPU 210 or the GPU 260.FIGS. 9A and 9B are diagrams illustrating the results of performing a tessellier operation on models having different redundancies. An image TES 1 illustrated in FIG. 9A may correspond to the ninth model MOD 14 illustrated in FIG. 8. An image TES 2 illustrated in FIG. 9B may correspond to the sixth model MOD 11 illustrated in FIG. 8.When the model MOD 11, MOD 12, or MOD 13 is read from the first memory 310- 1 by the GPU 260, the tessellator 261- 4 performs a tessellating operation on the model MOD 11, MOD 12, or MOD 13 according to the comparison information INF received from the hull shader 261- 3. At this time, the image TES1 of FIG. 9A may correspond to the model MOD11, MOD12, or MOD13 that has been tessellated.However, when the model MOD 14 is read from the first memory 310- 1 by the GPU 260, the tessellator 261- 4 does not perform tessellation on the model MOD 14 according to the comparison information INF received from the hull shader 261- 3. At this time, the image TES1 of FIG. 9A may correspond to the model MOD14 that has not been tessellated.When the model MOD 11 is read from the first memory 310- 1 by the GPU 260, the tessellator 261- 4 performs a tessellating operation on the model MOD 11 according to the comparison information INF received from the hull shader 261- 3. At this time, the image TES 2 of FIG. 9B may correspond to the model MOD 11 that has been tessellated.FIG. 10 is a flowchart of an operation of the computing device 100 illustrated in FIG. 1. Referring to FIGS. 1 to 10, prior to a rendering operation, a combination of the CPU 210, the GPU 260, and / or a modeler generates a plurality of the models MOD 1 to MOD 5 or MOD 11 to MOD 14 having different redundancies, and stores the models MOD 1 to MOD 5 or MOD 11 to MOD 14 in a memory accessible by the GPU 260 in operation S 110. The CPU 210 or the GPU 260 estimates at least one of the attributes or features of the computing device 100 in operation S 120.The GPU 260, and more specifically, the access circuit 252, reads one of the models MOD 1 to MOD 5 or MOD 11 to MOD 14 from the memory using an address (i.e., the estimation result) selected by the CPU 210 or the GPU 260 in operation S 130. The model that has been read is provided to graphics pipeline 260A.The graphics pipeline 260A, and more specifically a tessellier stage (e.g., the hull shader 261- 3), calculates a complexity of the model based on geometry information of the model and compares the calculated complexity with a reference complexity in operation S 140. The calculated complexity and the reference complexity may be defined by depth values of control points of a patch of, for example, a triangle, polygon, geometric primitive, and so forth, included in the model, or a curvature defined by the control points. The hull shader 261- 3 may calculate the complexity of the model in a unit of an object, primitive, patches, border, vertex, or control point.When the calculated complexity is determined to be less than the reference complexity, that is, when the tessellation of the model is needed, then the hull shader 261- 3 transmits to the tessellator 261- 4 the comparison information INF that instructs to perform a tessellating operation. At this time, the hull shader 261- 3 may transmit the model and tessellation factors for tessellation to the tessellator 261- 4. The tessellator 261- 4 performs the tessellating operation on the model in operation S 150.However, if the calculated complexity is greater than the reference complexity, that is, if the tessellation of the model is not needed, then the hull shader 261- 3 transmits to the tessellator 261- 4 the comparison information INF that instructs to pass the model without tessellation. The tessellator 261- 4 does not perform a tessellating operation on the model in operation S 160.FIG. 11 is a diagram of a method for selecting a model among models having different redundancies according to some embodiments of the inventive concept. Some or all of the method may be performed on a hardware device, for example, the CPU 210 and / or the GPU 260 as described herein. In the embodiments illustrated in FIG. 11, the CPU 210 estimates the attributes or features of the computing device 100. Referring to FIGS. 1 to 7 and 11, before a rendering operation in operation S 210, the CPU 210 generates the GPU 260 or a modeler a plurality of the models MOD 1 to MOD 5 having different redundancies, and stores the models MOD 1 to MOD 5 in the first memory 310- 1 accessible by the GPU 260.The CPU 210 estimates (i.e., calculates or measures) at least one of the attributes or features of the computing device 100 in operation S 220. As shown in FIGS. 2, 4, and 5, the CPU 210 transmits an address (for example, ADDi where 1≤i≤5) among the addresses ADD 1 to ADD 5 corresponding to the storage areas MEM 1 to MEM 5 storing the models MOD 1 to MOD 5, respectively, to the GPU 260 according to the estimation (i.e., calculation or measurement) result in operation S 230.The GPU 260 receives the address ADDi from the CPU 210 and reads a model MODi from the one of the memory areas MEM 1 to MEM 5 using the address ADDi in operation S 240. The GPU 260 calculates a complexity of the model MODi using geometry information related to the model MODi, and compares the calculated complexity with a reference complexity in operation S 250.The GPU 260 may determine whether to perform a tessellier operation on the model MODi according to the comparison result in operation S 260. In other words, the tessellator 261- 4 may or may not perform the tessellating operation on the model MODi according to the comparison information INF in operation S 260.FIG. 12 is a diagram of a method for selecting a model among models having different redundancies according to other embodiments of the inventive concept. In the embodiments illustrated in FIG. 12, the GPU 260 estimates the attributes or features of the computing device 100. Referring to FIGS. 1 to 7 and 12, before a rendering operation in operation S 310, the CPU 210 generates the GPU 260 or a modeler a plurality of the models MOD 1 to MOD 5 having different redundancies, and stores the models MOD 1 to MOD 5 in the first memory 310- 1 accessible by the GPU 260.The CPU 210 may transmit a plurality of the addresses ADD 1 to ADD 5 corresponding to the memory areas MEM 1 to MEM 5 storing the models MOD 1 to MOD 5 to the GPU 260 in operation S 320. The GPU 260 receives the addresses ADD 1 to ADD 5 from the CPU 210 and estimates (i.e., calculates or measures) at least one of the attributes or features of the computing device 100 in operation S 330.The GPU 260 selects an address (for example, ADDj, where 1≤j≤5) among the addresses ADD 1 to ADD 5 according to the estimation (i.e., calculation or measurement) result in operation S 340, and reads a model MODj from one of the storage areas MEM 1 to MEM 5 using the selected address ADDj in operation S 350. The GPU 260 calculates a complexity of the model MODj using geometry information related to the model MODj, and compares the calculated complexity with a reference complexity in operation S 360.The GPU 260 may determine whether to perform a tessellier operation on the model MODj according to the comparison result in operation S 370. In other words, the tessellator 261- 4 may or may not perform the tessellating operation on the model MODj according to the comparison information INF in operation S 370.FIG. 13 is a block diagram of a computing device 100A including a graphics card 500 according to some embodiments of the inventive concept. Referring to FIG. 13, the computing device 100A may include the CPU 210, the graphics card 500, and the display 400. The computing device 100A may be a TV (for example, a digital TV or a smart TV), a PC, a desktop computer, a laptop computer, a computer workstation, or a computing device using the graphics card 500.The graphics card 500 includes the GPU 260, an interface 510, a graphics memory 520, a digital-to-analog converter (DAC= Digital-to-analog converter= Digital-to-analog converter) 530, an output port 540, and a card connector 550. The interface 510 may transmit a command and / or data from the CPU 210 to the graphics memory 520, or may transmit information to the CPU 210 via the graphics card 500. The graphics memory 520 may store data generated by the GPU 260 and may transmit a command from the CPU 210 to the GPU 260.In conjunction with CPU 210, GPU 260 performs one or more of the operations described with reference to FIGS. 1-12. Data generated in the GPU 260 is transferred to the graphics memory 520.The DAC 530 converts digital signals into analog signals. The output terminal 540 transmits the analog signals, i.e., image signals from the DAC 530 to the display 400. The card connector 550 is inserted into a slot of a mainboard having the CPU 210.According to some embodiments of the inventive concept, a GPU determines whether or not to perform tessellation on a model selected from provided models, thereby reducing the operational load of a graphics pipeline (i.e., tessellation). In addition, the GPU stores models having different redundancies in memory in advance of a rendering operation, selects one of the models according to the feature of the computing device, and determines whether or not to perform tessellation on the selected model, thereby reducing the operational load of the graphics pipeline (i.e., tessellation).

Claims

A graphics processing unit (GPU) (260) for determining whether to perform a tessellating operation on a first model of a plurality of provided models according to a control of a central processing unit (CPU) (210), comprising: an access circuit (252) configured to read the first model from a memory (310-1, 310-2), wherein the memory (310-1, 310-2) stores the plurality of provided models having different redundancies, to calculate a complexity of the first model, to compare the calculated complexity with a reference complexity, and to determine whether to perform a tessellating operation on the first model according to the comparison result, wherein the GPU (260) is configured to calculate one or more addresses, which memory areas correspond to the memory (310-1, 310-2) storing the models, are received by the CPU (210) to estimate at least one of a bandwidth between the GPU (260) and the memory (310-1, 310-2), a calculation power of the GPU (260), a maximum power consumption of the GPU (260), a voltage and a frequency determined depending on a dynamic voltage frequency scaling (DVFS) of the GPU (260), and a temperature of the GPU (260) to select a first address among the addresses based on an estimation result to read the first model from a first memory area among the memory areas using the selected first address, to calculate the complexity of the first model using geometry information of the first model, and to determine whether to perform the tessellier operation on the first model according to the result of comparing the calculated complexity with the reference complexity.The GPU (260) of claim 1, wherein the GPU (260) is configured to calculate the complexity of the first model in a unit of an object, primitive, patch, edge, vertex, and a control point.The GPU (260) of claim 2, wherein the GPU (260) is configured to calculate the complexity of the first model based on depth values of vertices included in the primitive or a curvature defined by the vertices.The GPU (260) of claim 2, wherein the GPU (260) is configured to calculate the complexity of the first model based on depth values of control points included in the patch or a curvature defined by the control points.A system-on-chip (SoC) (200) comprising: a graphics processing unit (GPU) (260) comprising a plurality of tessellier stages; a central processing unit (CPU) (210); a memory (310-1, 310-2) accessed by the GPU (260); and a memory controller (220-1, 220-2) controlled by the CPU (210), wherein the GPU (260) is configured to read a first model from the memory (310-1, 310-2) storing a plurality of provided models having different redundancies via the memory controller (220-1, 220-2) to calculate a complexity of the first model, to compare the calculated complexity with a reference complexity, and to determine whether to perform tessellation on the first model according to the comparison result according to a control of the CPU (210), wherein the CPU (210) is configured to transmit one or more addresses corresponding to storage areas of the memory (310-1, 310-2) storing the models to the GPU (260), and the GPU (260) is configured to transmit at least one of a bandwidth between the GPU (260) and the memory (310-1, 310-2), a calculation power of the GPU (260), a maximum power consumption of the GPU (260), a voltage, and a frequency, which are determined depending on dynamic voltage frequency scaling (DVFS) of the GPU (260) and a temperature of the GPU (260), to select a first address among the addresses based on the estimation result, and to read the first model from a first memory area among the memory areas using the first address.The SoC (200) of claim 5, wherein the GPU (260) is configured to calculate the complexity using geometry information of control points of each of patches included in the first model.The SoC (200) of claim 6, wherein the geometry information is depth values of the control points or a curvature defined by the control points.The SoC (200) of claim 5, wherein the memory (310-1, 310-2) storing the models before a rendering operation is implemented outside the SoC (200).The SoC (200) of claim 5, wherein the GPU (260) is configured to calculate the complexity of the first model using a depth value of each of control points of each of patches included in the first model or a curvature defined by the control points.The SoC (200) of claim 9, wherein the SoC (200) is configured to perform tessellation on the first model when the complexity of the first model is less than the reference complexity.The SoC (200) of claim 5, wherein the tessellator stage comprises: a hull shader (261-3); a tessellator (261-4); and a domain shader (261-5), wherein the calculation and the comparison are performed by the hull shader (261-3), wherein the determination is performed by the tessellator (261-4), and wherein the tessellator (261-4) is configured to pass the first model to the domain shader (261-5) or perform the tessellation on the first model using tessellator factors transmitted from the tessellator (261-4) based on the result of the determination.A computing device (100) comprising: a memory (310-1, 310-2) configured to store a plurality of provided models having different redundancies; a display (400); and a system-on-chip (SoC) (200) configured to control the memory (310-1, 310-2) and the display (400), the SoC (200) comprising: a graphics processing unit (GPU) (260) having tessellier stages; a central processing unit (CPU) (210); a display controller (240) configured to control the display (400) according to control of the CPU (210); and a memory controller (220-1, 220-2) configured to:, to control the memory (310-1, 310-2) according to control of the CPU (210), wherein the GPU (260) is configured to read a first model from the memory (310-1, 310-2) storing the models by the memory controller (220-1, 220-2), calculate a complexity of the first model, compare the calculated complexity with a reference complexity, and determine whether to perform tessellation on the first model according to the comparison result according to control of the CPU (210), wherein the CPU (210) is configured to transmit one or more addresses corresponding to storage areas of the memory (310-1, 310-2) storing the models to the GPU (260), and the GPU (260) is configured to:, to estimate at least one of a bandwidth between the GPU (260) and the memory (310-1, 310-2), a calculation power of the GPU (260), a maximum power consumption of the GPU (260), a voltage and a frequency determined depending on a dynamic voltage frequency scaling (DVFS) of the GPU (260), and a temperature of the GPU (260), to select a first address among the addresses based on the estimation result, and to read the first model from a first storage area among the storage areas using the first address.The computing device (100) of claim 12, wherein the GPU (260) is configured to calculate the complexity of the first model in a unit of an object, primitive, patch, edge, vertex, and a control point.The computing device (100) of claim 12, wherein the GPU (260) is configured to perform the tessellier operation on the first model when the complexity of the first model is less than the reference complexity.A graphics processing unit (GPU) system, comprising: a memory (310-1, 310-2) configured to store a plurality of provided models having different redundancies; a graphics processor configured to read a model of the plurality of provided models from the memory (310-1, 310-2), calculate a complexity of the model, compare the calculated complexity with a reference complexity, and determine whether to perform a tessellating operation on the model in response to the comparison result, wherein the graphics processor is configured to receive one or more addresses corresponding to memory areas of the memory (310-1, 310-2) storing the models from the CPU (210), to estimate at least one of a bandwidth between the graphics processor and the memory (310-1, 310-2), a computing power of the graphics processor, a maximum power consumption of the graphics processor, a voltage and a frequency determined depending on a dynamic voltage frequency scaling (DVFS) of the graphics processor, and a temperature of the graphics processor, to select a first address among the addresses based on an estimation result, to read the first model from a first memory area among the memory areas using the selected first address, to calculate the complexity of the first model using geometry information of the first model, and to determine whether to perform the tessellating operation on the first model according to the result of comparing the calculated complexity with the reference complexity..The GPU system of claim 15, wherein the graphics processor includes a command decoder that controls the GPU system according to a command received from the CPU (210) or the memory (310-1, 310-2).The GPU system of claim 15, wherein the graphics processor includes access circuitry (252) configured to read the model from the memory region using the selected first address, calculate the complexity of the model using geometry information of the model, compare the complexity to the reference complexity, and determine whether to perform the tessellier operation on the model in response to the comparison result.The GPU system of claim 15, wherein the graphics processor includes one or more processing units to perform a raster operation.

Citation Information

Patent Citations

  • Direct Ray Tracing of 3D Scenes

    US20120169728A1

  • Patched shading in graphics processing

    US20130265309A1

  • Rapid zippering for real time tesselation of bicubic surfaces

    US7295204B2