Parallelization of GPU Synthesis and DPU Topology Selection
By parallelizing DPU topology selection and GPU composition selection on the CPU and adapting them to the application layer of GPU composition selection, the problem of excessive CPU workload during high refresh rate display composition in the prior art is solved, achieving higher refresh rates and faster display pipeline composition.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- QUALCOMM INC
- Filing Date
- 2022-09-19
- Publication Date
- 2026-05-26
AI Technical Summary
Existing technologies cannot effectively solve the problem of excessive CPU workload during high refresh rate display synthesis, which causes the display pipeline to be unable to complete the synthesis in time, affecting the refresh rate of the display.
By parallelizing DPU topology selection and GPU synthesis on the CPU, and utilizing batch prediction heuristics to select application layers suitable for synthesis on both the DPU and GPU, parallel processing of DPU and GPU is achieved.
The display pipeline's rendering speed has been improved, more stringent delay rendering deadlines have been met, and higher refresh rates have been achieved, allowing the display to scale to higher refresh rates.
Smart Images

Figure CN117980988B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims the benefit of U.S. Patent Application No. 17 / 449,630, filed on September 30, 2021, entitled “PARALLELIZATION OF GPU COMPOSITION WITH DPU TOPOLOGY SELECTION”, the entire contents of which are hereby expressly incorporated herein by reference. Technical Field
[0003] In general, this disclosure relates to processing systems; more specifically, this disclosure relates to one or more technologies for display processing. Background Technology
[0004] Computing devices frequently perform graphics and / or display processing (e.g., using a graphics processing unit (GPU), a central processing unit (CPU), a display processor, etc.) to render and display visual content. Such computing devices can include, for example, computer workstations, mobile phones such as smartphones, embedded systems, personal computers, tablet computers, and video game consoles. A GPU is configured to execute a graphics processing pipeline comprising one or more processing stages that work together to execute graphics processing commands and output frames. The CPU can control the operation of the GPU by issuing one or more graphics processing commands to it. Modern CPUs are typically capable of executing multiple applications simultaneously, each of which may require the GPU to be used during execution. A display processor can be configured to convert digital information received from the CPU into analog values and can issue commands to a display panel to display visual content. Devices that provide visually rendered content on a display can use a CPU, GPU, and / or display processor.
[0005] Current technology may not be able to handle the increasingly demanding CPU workloads associated with compositing on high refresh rate displays. Improvements to hybrid compositing techniques are needed to accommodate higher refresh rates without unduly straining CPU processing resources. Summary of the Invention
[0006] To provide a basic understanding of one or more aspects of the invention, a brief overview of these aspects is given below. This overview is not an exhaustive summary of all anticipated aspects, nor is it intended to identify key or essential elements of all aspects, or to describe the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simple form as a prelude to the detailed description that follows.
[0007] In one aspect of this disclosure, a method, computer-readable medium, and apparatus are provided. The apparatus can receive instructions for a plurality of application layers for synthesis at a first processor and a second processor. The apparatus can select one or more first application layers from the plurality of application layers for synthesis attempts on the first processor, and select one or more second application layers from the plurality of application layers for synthesis on the second processor. The apparatus can send each of the one or more first application layers to the first processor for synthesis, and send each of the one or more second application layers to the second processor for synthesis. The first processor and the second processor can synthesize at least some of the one or more first application layers and at least some of the one or more second application layers in parallel.
[0008] For the purposes described above and related, one or more aspects include the features detailed below and specifically pointed out in the claims. Certain exemplary features of one or more aspects are described in detail below and in the accompanying drawings. However, these features merely indicate some of the various methods that can employ the basic principles of these aspects, and this description is intended to include all such aspects and their equivalents. Attached Figure Description
[0009] Figure 1 This is a block diagram illustrating an example content generation system based on one or more techniques according to this disclosure.
[0010] Figure 2 This is a block diagram showing an example display frame.
[0011] Figure 3 This is a diagram illustrating a display composition in a display pipeline, based on one or more techniques according to the present disclosure.
[0012] Figure 4 This is a diagram illustrating a composite of one or more techniques according to this disclosure.
[0013] Figure 5 This is a diagram illustrating layer segmentation according to one or more techniques based on this disclosure.
[0014] Figure 6 This is a call flowchart illustrating example communication between a CPU, a display processing unit (DPU), and a GPU, based on one or more techniques described in this disclosure.
[0015] Figure 7 The flowchart illustrates a processing method as an example of one or more techniques according to this disclosure.
[0016] Figure 8The flowchart illustrates a processing method as an example of one or more techniques according to this disclosure. Detailed Implementation
[0017] The following description, with reference to the accompanying drawings, provides a more complete overview of various aspects of the systems, apparatuses, computer program products, and methods. However, this disclosure may be implemented in many different forms and should not be construed as being limited to any particular structure or function given throughout this disclosure. Rather, these aspects are provided only to make this disclosure thorough and complete, and to fully convey the scope of protection of this disclosure to those skilled in the art. Based on the teachings of this application, those skilled in the art should understand that the scope of this disclosure is intended to cover any aspect of the systems, apparatuses, computer program products, and methods disclosed herein, whether implemented independently or in combination with other aspects of this disclosure. For example, an apparatus or method may be implemented using any number of the aspects set forth herein. Furthermore, the scope of this disclosure is intended to cover such apparatus or methods that may be implemented using other structures, functions, or structures and functions other than those of the aspects of this disclosure set forth herein, or structures and functions different from those of the aspects of this disclosure set forth herein. Any aspect disclosed herein may be embodied in one or more components of the invention.
[0018] While various aspects have been described herein, numerous variations and arrangements of these aspects also fall within the scope of this disclosure. Although some potential benefits and advantages of the aspects of this disclosure have been mentioned, the scope of this disclosure is not intended to be limited to specific benefits, uses, or objectives. Rather, the aspects of this disclosure are intended to be broadly applicable to various wireless technologies, system configurations, processing systems, networks, and transport protocols, some of which are illustrated by way of example in the accompanying drawings and the following description. The detailed description and accompanying drawings are merely illustrative and not restrictive of this disclosure, and the scope of this disclosure is defined by the appended claims and their equivalents.
[0019] Some aspects are now described with reference to various apparatuses and methods. These apparatuses and methods will be described in the following detailed embodiments and depicted in the accompanying drawings by various blocks, components, circuits, processes, algorithms, etc. (collectively referred to as “elements”). Such elements can be implemented using electronic hardware, computer software, or any combination thereof. Whether these elements are implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system.
[0020] For example, an element, any part of an element, or any combination of elements can be implemented as a "processing system" including one or more processors (which may also be called processing units). Examples of processors include microprocessors, microcontrollers, graphics processing units (GPUs), general-purpose GPUs (GPGPUs), central processing units (CPUs), application processors, digital signal processors (DSPs), reduced instruction set computing (RISC) processors, system-on-a-chip (SoCs), baseband processors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), programmable logic devices (PLDs), state machines, gate logic, discrete hardware circuits, and other suitable hardware configured to perform the various functions described throughout this disclosure. One or more processors in a processing system can execute software. Software should be broadly interpreted as meaning instructions, instruction sets, code, code segments, program code, programs, subroutines, software components, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, functions, etc., regardless of whether it is referred to as software, firmware, middleware, microcode, hardware description languages, or other terms.
[0021] The term "application" can refer to software. As described herein, one or more technologies can refer to an application (e.g., software) configured to perform one or more functions. In such examples, the application may be stored in memory (e.g., on-chip memory of a processor, system memory, or any other memory). The hardware described herein (e.g., a processor) may be configured to execute the application. For example, an application may be described as including code that, when executed by the hardware, causes the hardware to perform one or more technologies described herein. For instance, the hardware may access code in memory and execute the code accessed from memory to perform one or more technologies described herein. In some instances, components are identified in this disclosure. In such instances, a component may be hardware, software, or a combination thereof. A component may be a single component or a subcomponent of a single component.
[0022] In one or more examples described herein, the functionality described herein can be implemented in hardware, software, or any combination thereof. When implemented in software, these functions can be stored or encoded as one or more instructions or code on a computer-readable medium. Computer-readable media include computer storage media. Storage media can be any available medium accessible to a computer. By way of example, and not limitation, such computer-readable media can include random access memory (RAM), read-only memory (ROM), electrically erasable programmable ROM (EEPROM), optical disc storage, magnetic disk storage, other magnetic storage devices, combinations of computer-readable media of the foregoing types, or any other medium capable of storing computer-executable code in the form of instructions or data structures and accessible to a computer.
[0023] As used herein, instances of the term "content" can refer to "graphic content," "image," etc., regardless of whether these terms are used as adjectives, nouns, or other parts of speech. In some instances, as used herein, the term "graphic content" can refer to the content produced by one or more processes in the graphics processing pipeline. In other instances, as used herein, the term "graphic content" can refer to the content produced by a processing unit configured to perform graphics processing. In still other instances, as used herein, the term "graphic content" can refer to the content produced by a graphics processing unit.
[0024] High refresh rate (HRR) displays (e.g., displays with a refresh rate of 120Hz or higher) may be associated with stringent latency constraints in the following ways: determining the application layers eligible for overlay compositing at the Display Processing Unit (DPU), scheduling GPU compositing for application layers not eligible for overlay compositing at the DPU, and configuring DPU hardware resources for the final compositing. Compositing of a first subset of application layers at the DPU and a second subset of application layers at the GPU can be termed hybrid-mode compositing. Latency constraints can be particularly stringent in the case of hybrid-mode compositing at high refresh rates. Although optimized for other parts of the display pipeline, DPU topology selection and GPU compositing are still serialized operations and may lie in the path that determines the longest total execution time.
[0025] The aspects described in this article relate to the parallelization of DPU topology selection and GPU shader programming at the CPU, thereby enabling faster compositing in the display pipeline to meet increasingly tight compositing deadlines and allowing displays to be scaled to higher refresh rates. In the various configurations described herein, DPU can refer to any type of display processor (e.g., a dedicated DPU, a CPU, or a GPU that performs display processing functions).
[0026] Figure 1This diagram illustrates a block diagram of an example content generation system 100 configured to implement one or more of the techniques described herein, based on one or more techniques of this disclosure. The content generation system 100 includes a device 104. Device 104 may include one or more components or circuitry for performing the various functions described herein. In some examples, one or more components of device 104 may be components of a System-on-a-Chip (SOC). Device 104 may include one or more components configured to perform one or more techniques of this disclosure. In the illustrated example, device 104 may include a processing unit 120, a content encoder / decoder 122, and a system memory 124. In some aspects, device 104 may include multiple optional components (e.g., a communication interface 126, a transceiver 132, a receiver 128, a transmitter 130, a display processor 127, and one or more displays 131). Display 131 may refer to one or more displays 131. For example, display 131 may include a single display or multiple displays, which may include a first display and a second display. The first display may be a left-eye display, and the second display may be a right-eye display. In some examples, the first and second displays may receive different frames to render on them. In other examples, the first and second displays may receive the same frames to render on them. In further examples, the results of graphics processing may not be displayed on the devices; for example, the first and second displays may not receive any frames to render on them. Instead, these frames or the results of graphics processing may be transmitted to another device. In some respects, this can be referred to as segmented rendering.
[0027] Processing unit 120 may include internal memory 121. Processing unit 120 may be configured to perform graphics processing using graphics processing pipeline 107. Content encoder / decoder 122 may include internal memory 123. In some examples, device 104 may include a processor that may be configured to perform one or more display processing techniques on one or more frames generated by processing unit 120 before displaying them on one or more displays 131. While the processor in the exemplary content generation system 100 is configured as display processor 127, it should be understood that display processor 127 is one example of a processor, and other types of processors, controllers, etc., may be used as alternatives to display processor 127. Display processor 127 may be configured to perform display processing. For example, display processor 127 may be configured to perform one or more display processing techniques on one or more frames generated by processing unit 120. One or more displays 131 may be configured to display or otherwise present the frames processed by display processor 127. In some examples, one or more displays 131 may include one or more of the following: liquid crystal display (LCD), plasma display, organic light-emitting diode (OLED) display, projection display device, augmented reality display device, virtual reality display device, head-mounted display, or any other type of display device.
[0028] External memory (e.g., system memory 124) of processing unit 120 and content encoder / decoder 122 can be accessed by processing unit 120 and content encoder / decoder 122. For example, processing unit 120 and content encoder / decoder 122 can be configured to read from and / or write to external memory (e.g., system memory 124). Processing unit 120 can be communicatively coupled to system memory 124 via a bus. In some examples, processing unit 120 and content encoder / decoder 122 can be communicatively coupled to internal memory 121 via a bus or via a different connection.
[0029] Content encoder / decoder 122 can be configured to receive graphical content from any source (e.g., system memory 124 and / or communication interface 126). System memory 124 can be configured to store the received encoded or decoded graphical content. Content encoder / decoder 122 can be configured to receive encoded or decoded graphical content, for example, from system memory 124 and / or communication interface 126, in the form of encoded pixel data. Content encoder / decoder 122 can be configured to encode or decode any graphical content.
[0030] Internal memory 121 or system memory 124 may include one or more volatile or non-volatile memories or storage devices. In some examples, internal memory 121 or system memory 124 may include RAM, static random access memory (SRAM), dynamic random access memory (DRAM), erasable programmable ROM (EPROM), EEPROM, flash memory, magnetic data media or optical storage media, or any other type of memory. According to some examples, internal memory 121 or system memory 124 may be a non-transitory storage medium. The term "non-transitory" may mean that the storage medium is not embodied in a carrier wave or propagating signal. However, the term "non-transitory" should not be construed as meaning that internal memory 121 or system memory 124 is not movable or that its contents are static. For example, system memory 124 can be removed from device 104 and moved to another device. As another example, system memory 124 may not be removable from device 104.
[0031] Processing unit 120 may be a CPU, GPU, GPGPU, or any other processing unit that can be configured to perform graphics processing. In some examples, processing unit 120 may be integrated into the motherboard of device 104. In further examples, processing unit 120 may reside on a graphics card mounted in a port on the motherboard of device 104, or may otherwise be incorporated into a peripheral device configured to interoperate with device 104. Processing unit 120 may include one or more processors, such as one or more microprocessors, GPUs, ASICs, FPGAs, arithmetic logic units (ALUs), DSPs, discrete logic, software, hardware, firmware, other equivalent integrated or discrete logic circuits, or any combination thereof. If these techniques are implemented in part in software, processing unit 120 may store instructions for software in a suitable, non-transitory computer-readable storage medium (e.g., internal memory 121), and these instructions may be executed in hardware using one or more processors to perform the techniques of this disclosure. Any of the foregoing (including hardware, software, combinations of hardware and software, etc.) may be considered as one or more processors.
[0032] The content encoder / decoder 122 can be any processing unit configured to perform content decoding. In some examples, the content encoder / decoder 122 can be integrated into the motherboard of device 104. The content encoder / decoder 122 may include one or more processors, such as one or more microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), arithmetic logic units (ALUs), digital signal processors (DSPs), video processors, discrete logic, software, hardware, firmware, other equivalent integrated or discrete logic circuits, or any combination thereof. If these techniques are implemented in part using software, the content encoder / decoder 122 may store software instructions in a suitable, non-transitory computer-readable storage medium (e.g., internal memory 123), and these instructions may be executed in hardware using one or more processors to perform the techniques of this disclosure. Any of the foregoing (including hardware, software, combinations of hardware and software, etc.) can be considered as one or more processors.
[0033] In some aspects, the content generation system 100 may include an optional communication interface 126. Communication interface 126 may include a receiver 128 and a transmitter 130. Receiver 128 may be configured to perform any of the receiving functions described herein with respect to device 104. Additionally, receiver 128 may be configured to receive information (e.g., eye or head position information, rendering commands, and / or location information) from another device. Transmitter 130 may be configured to perform any of the transmitting functions described herein with respect to device 104. For example, transmitter 130 may be configured to transmit information to another device, which may include a request for content. Receiver 128 and transmitter 130 may be combined to form transceiver 132. In such an example, transceiver 132 may be configured to perform any of the receiving and / or transmitting functions described herein with respect to device 104.
[0034] Refer again Figure 1In some aspects, processing unit 120 may include a layer predictor 198 configured to receive indications for a plurality of application layers to be synthesized at a first processor and a second processor. Layer predictor 198 may be further configured to: select one or more first application layers from the plurality of application layers for attempted synthesis at the first processor, and select one or more second application layers from the plurality of application layers for synthesis at the second processor. Layer predictor 198 may be further configured to send each of the one or more first application layers to the first processor for synthesis, and send each of the one or more second application layers to the second processor for synthesis. The first processor and the second processor may synthesize at least some of the one or more first application layers and at least some of the one or more second application layers in parallel.
[0035] Devices such as device 104 can refer to any device, apparatus, or system configured to perform one or more of the techniques described herein. For example, a device can be a server, base station, user equipment, client device, station, access point, computer (e.g., personal computer, desktop computer, laptop computer, tablet computer, computer workstation, or mainframe computer), terminal product, apparatus, telephone, smartphone, server, video game platform or console, handheld device (e.g., portable video game device or personal digital assistant (PDA)), wearable computing device (e.g., smartwatch, augmented reality device, or virtual reality device), non-wearable device, display or display device, television set, set-top box, intermediate network device, digital media player, video streaming device, content streaming device, in-vehicle computer, any mobile device, any device configured to generate graphical content, or any device configured to perform one or more of the techniques described herein. The processes described herein may be described as being performed by a specific component (e.g., GPU), but in other embodiments, other components consistent with the disclosed embodiments (e.g., CPU) may be used to perform them.
[0036] GPUs can process various types of data or data packets within the GPU pipeline. For example, in some aspects, a GPU can process two types of data or data packets (e.g., context register packets and fetch call data). Context register packets can be a set of global state information (e.g., information about global registers, shadow procedures, or constant data) that can regulate how the graphics context is processed. For example, a context register packet may include information about color formats. In some aspects of context register packets, there may be bits indicating which workload belongs to the context register. Furthermore, multiple functions or programs can run simultaneously and / or in parallel. For example, a function or program may describe an operation (e.g., a color mode or color format). Therefore, context registers can define multiple states of the GPU.
[0037] Context states can be used to determine the function of individual processing units, such as vertex extractors (VFDs), vertex shaders (VSs), shader processors, or geometry processors, and / or the operating mode of the processing unit. For this purpose, the GPU can use context registers and program data. In some aspects, the GPU can generate workloads (e.g., vertex or pixel workloads) in the pipeline based on the context register definitions of modes or states. Certain processing units (e.g., VFDs) can use these states to determine certain functions, such as how vertices are assembled. Because these modes or states can change, the GPU may need to modify the corresponding context. Furthermore, the workload corresponding to a mode or state can follow the changed mode or state.
[0038] Figure 2 This is a block diagram 200 illustrating an example display frame including a processing unit 120, a system memory 124, a display processor 127, and a display 131, as can be identified in conjunction with exemplary device 104.
[0039] GPUs are typically included in devices that provide content for visual presentation on a display. For example, processing unit 120 may include GPU 210, which is configured to render graphics data for display on a computing device (e.g., device 104), which may be a computer workstation, mobile phone, smartphone or other smart device, embedded system, personal computer, tablet computer, video game console, etc. The operation of GPU 210 can be controlled based on one or more graphics processing commands provided by CPU 215. CPU 215 can be configured to execute multiple applications simultaneously. In some cases, each of the multiple applications executing simultaneously can utilize GPU 210 concurrently. Processing techniques can be performed by outputting frames over physical or wireless communication channels via processing unit 120.
[0040] System memory 124, executable by processing unit 120, may include user space 220 and kernel space 225. User space 220 (sometimes referred to as "application space") may include software applications and / or application frameworks. For example, software applications may include operating systems, media applications, graphics applications, workspace applications, etc. Application frameworks may include frameworks used by one or more software applications, such as libraries, services (e.g., display services, input services, etc.), application interfaces (APIs), etc. Kernel space 225 may also include display driver 230. Display driver 230 may be configured to control display processor 127. For example, display driver 230 may cause display processor 127 to synthesize one or more application layers at overlay plane 237. Each of overlay planes 237 may refer to a unit of display processor / DPU hardware resources. For example, display processor 127 with four overlay planes may synthesize four layers simultaneously.
[0041] Display processor 127 includes display control block 235 and display interface 240. Display processor 127 can be configured to operate the functions of display 131 (e.g., based on input received from display driver 230). For example, display control block 235 can be configured to receive instructions from display driver 230 to cooperate with overlay plane (e.g., DPU overlay plane) 237 to synthesize one or more application layers. In some examples, display control block 235 can additionally or alternatively perform post-processing based on image data provided by system memory 124 executed by processing unit 120.
[0042] Display interface 240 can be configured to cause display 131 to display image frames. Display interface 240 can output image data to display 131 according to an interface protocol (e.g., MIPIDSI (Mobile Industrial Processor Interface, Display Serial Interface)). That is, display 131 can be configured according to the MIPIDSI standard. The MIPIDSI standard supports video mode and command mode. In the example where display 131 operates in video mode, display processor 127 can continuously refresh the graphic content of display 131. For example, the entire graphic content can be refreshed in each refresh cycle (e.g., progressive scan). In the example where display 131 operates in command mode, display processor 127 can write the graphic content of a frame into buffer 250.
[0043] Based on the display controller 245, display client 255, and buffer 250, frames are displayed on the monitor 131. The display controller 245 can receive image data from the display interface 240 and store the received image data in the buffer 250. In some examples, the display controller 245 can output the image data stored in the buffer 250 to the display client 255. Therefore, the buffer 250 can represent the local memory of the monitor 131. In some examples, the display controller 245 can directly output the image data received from the display interface 240 to the display client 255.
[0044] Display client 255 may be associated with a touch panel that can sense interactions between the user and display 131. When the user interacts with display 131, one or more sensors in the touch panel may output signals to display controller 245, indicating which of the one or more sensors is active, the duration of the sensor activity, the pressure applied to the one or more sensors, etc. Display controller 245 can use the sensor outputs to determine how the user interacts with display 131. Display 131 may further be associated with or include other devices, such as a camera, microphone, and / or speaker that operate in conjunction with display client 255.
[0045] Some processing techniques of device 104 can be performed in three stages (e.g., stage 1: rendering stage; stage 2: compositing stage; stage 3: display / transmission stage). However, other processing techniques can combine the compositing stage and the display / transmission stage into a single stage so that the processing techniques can be performed based on the two overall stages (e.g., stage 1: rendering stage; stage 2: compositing / display / transmission stage). During the rendering stage, GPU 210 can process the content buffer based on the execution of the application that generates content pixel by pixel. During the compositing and display stages, pixel elements can be assembled to form a frame, which is then transmitted to the physical display panel / subsystem (e.g., display 131) that displays the frame.
[0046] High refresh rate (HRR) displays (e.g., displays with a refresh rate of 120Hz or higher) may be associated with stringent latency constraints in the following ways: determining the application layers eligible for overlay compositing at the DPU, scheduling GPU compositing for application layers not eligible for overlay compositing at the DPU, and configuring the DPU hardware resources for the final compositing. A combination of a first subset of application layers at the DPU and a second subset of application layers at the GPU can be referred to as a hybrid mode combination. Latency constraints can be particularly stringent in the case of high refresh rate hybrid mode compositing.
[0047] Currently, processing 12 or more layers in hybrid mode compositing at a 120Hz refresh rate can saturate CPU efficiency cores. The display pipeline may be unable to keep up with the compositing deadlines for each rendering cycle. While a 240Hz refresh rate could further reduce the total latency budget to half that associated with a 120Hz refresh rate, more application layers fall back to the GPU due to the segmentation of application layers across multiple overlay planes. As the GPU compositing load increases further, the CPU workload associated with programming the GPU command flow may also increase further. This could further challenge the display pipeline's ability to keep up with compositing deadlines.
[0048] Figure 3 Figure 300 illustrates display composition in a display pipeline. Composer 302 may be an operating system component responsible for compositing or blending layers using DPU and / or GPU hardware. Composer backend 304 may be a vendor-specific software implementation of an interface component that interfaces with the DPU. Composer backend 304 may determine the segmentation of application layers for composition on the DPU and GPU, as further described below. Mixed-mode composition layer segmentation decisions can be made in a deterministic manner. In other words, application layers can be segmented based on a deterministic algorithm that correctly segments the application layer into application layers for DPU processing and program layers for GPU processing on the first pass. The mixed-mode composition layer segmentation operation may be linearly related to the number of application layers (an application layer can refer to an application user interface (UI) element visible on the screen at a given time, such as a status bar, navigation bar, image wallpaper, or launcher, which is a visible application layer on the home screen). In other words, the runtime associated with the mixed-mode composition layer segmentation operation can increase linearly with the number of application layers. As the display refresh rate increases, the CPU workload associated with the layer segmentation operation also increases accordingly. Furthermore, despite optimizations for other parts of the display pipeline (e.g., integration layer selection, commitment to future fencing, etc.), DPU topology selection and GPU composition can still be serialized operations and may lie in the path that determines the longest total execution time. As shown, at compositor 302, programming 320a of the GPU shader for fallback layers (i.e., layers that cannot be composed at DPU 308 and will be composed at GPU 306) and final configuration 322 of the DPU can be cascaded. Composition 324 at GPU 306 and composition 326 at DPU 308 can also be largely cascaded (GPU composition 324 can begin after compositor 302 issues refresh command 320b (with a delay), while DPU composition 326 can begin after compositor 302 begins final configuration 322 of DPU hardware 308 (with a delay)).
[0049] The aspects described in this article may relate to the parallelization of DPU topology selection and GPU shader programming at the CPU, thus enabling faster compositing in the display pipeline to meet tighter compositing deadlines and allowing displays to be scaled to higher refresh rates.
[0050] In one or more aspects, a parallelized model can be implemented, which includes using batch prediction heuristics on the CPU to select consecutive batches of application layers suitable for compositing at the DPU and GPU. The DPU may be energy-efficient, but its functionality may be limited. Some application layers may not support compositing on the DPU. For example, application layers containing invalid rotations (e.g., rotations that are not 90 degrees, 180 degrees, or 270 degrees, etc.) may not support compositing on the DPU. In another example, shadow layers may also not support compositing on the DPU because they may contain dynamically generated content at runtime. On the other hand, the GPU may be more general-purpose. However, the GPU's energy efficiency may be lower than that of the DPU.
[0051] Figure 4 Figure 400 is a composite illustration of one or more techniques according to this disclosure. Figure 4 In the diagram, a solid vertical line can represent vertical synchronization (V). SYNCThe cycle logic is divided into four equal parts (e.g., if the refresh rate is 240 frames per second (fps), each part can be 1 millisecond (ms) long). Horizontal arrows can indicate the time taken for various graphical processes. In one or more aspects, the CPU can perform DPU topology selection 418 (i.e., mapping some application layers to DPU overlay planes) in parallel with some other application layers being composed on the GPU (e.g., via compositor 404) in configuration 406. Specifically, the CPU can program 410 the GPU shaders to select (418) DPU-specific overlay planes for DPU-eligible application layers in parallel with the composition 414 of some application layers on the GPU. A DPU overlay plane can refer to a unit of DPU hardware resources (e.g., DPU hardware with four overlay planes can mix four layers simultaneously). Composition 420 at DPU 424 also overlaps temporally with composition 414 at GPU 422. If, during topology selection 418, it is discovered that the DPU does not support any application layers scheduled for the first pass of synthesis on the DPU (e.g., discovered by the synthesizer backend), the CPU (e.g., synthesizer 404) can configure (408) these application layers as a second batch to be synthesized on the GPU 416, and can schedule (412) the second batch for synthesis on the GPU. Incorrectly scheduling DPU-unqualified application layers for the first pass of synthesis on the DPU is likely rare. In any case, the size of the second batch of application layers used for synthesis on the GPU (which may include DPU-unqualified application layers that were incorrectly scheduled for DPU synthesis, if any) is likely to be much smaller than the initial batch synthesized on the GPU.
[0052] In one or more aspects, the GPU batch predictor 402, which can be executed on the CPU, can select a first application layer for synthesis on the DPU and a second application layer for synthesis on the GPU batch based on a batch prediction heuristic, as described in further detail below. Figure 4 The dashed lines in the diagram can indicate where synthesizer 404 is relative to V. SYNC The logical segmentation of the cycle is used to invoke the expected timeline of the GPU batch predictor 402. Vertical arrows can indicate the input / output call sequence from the synthesizer 404 to the GPU batch predictor 402, and then back.
[0053] Figure 5Figure 500 illustrates layer segmentation according to one or more techniques according to this disclosure. In one aspect, the first application layer selected for compositing at the DPU may be associated with a lower Z-order. A higher Z-order application layer (e.g., facing top) may be positioned above an overlapping application layer with a lower Z-order (e.g., facing bottom). The application layer associated with the lower Z-order (e.g., wallpaper) may be larger in size. The DPU can be optimized for larger application layers. Therefore, the application layer associated with the lower Z-order may be more suitable for compositing at the DPU. In one aspect, the first application layer selected for compositing at the DPU may be associated with a higher priority. The application layer associated with the higher priority may include those associated with camera applications, gaming applications, or high dynamic range (HDR) video applications, etc. In one aspect, the second application layer selected for compositing on the GPU may be associated with a higher Z-order. The application layer associated with the higher Z-order may be smaller in size. Since the DPU may not be optimized for smaller application layers, the application layer associated with the higher Z-order may be more suitable for compositing at the GPU. The pivot can indicate the split between the first application layer selected for composition at the DPU and the second application layer selected for composition at the GPU.
[0054] In one aspect, separate CPU cores (e.g., separate CPU efficiency cores) can be used to parallelize the DPU and GPU execution paths. Since superposition synthesis is associated with linear complexity with respect to the number of application layers, leveraging near-uniformly sized batches of application layers to parallelize the DPU and GPU execution paths on separate CPU cores could potentially result in halving the overall execution time.
[0055] Return to Figure 4 In one aspect, at GPU batch predictor 402, selecting a second application layer for compositing at the GPU may include: application layers already marked for compositing at the GPU. These may include application layers not supported by the DPU, such as shadow layers, layers with empty handles, or layers with invalid rotations (e.g., rotations not 90 degrees, 180 degrees, 270 degrees, etc.). In one aspect, selecting a second application layer for compositing at the GPU may include: application layers marked for compositing at the GPU in one or more previous rounds. In one aspect, selecting a second application layer for compositing at the GPU may include: a new application layer associated with a geometric change based on the application layer marked for compositing at the GPU in one or more previous rounds.
[0056] In one aspect, the batch prediction heuristic used in the GPU batch predictor 402 to select the first and second application layers can be based on a learned model. This learned model can be trained using a hybrid-mode synthesis layer segmentation decision made at the synthesizer backend (e.g., based on artificial intelligence techniques). The learned model can be trained online or offline. For example, if in one round an application layer containing an 8x scaling factor is found to be unsuitable for synthesis at the DPU, the result can be used to update the learned model, and based on the learned model, the GPU batch predictor 402 can classify similar application layers as the second application layer for future synthesis at the GPU.
[0057] In one aspect, the hybrid mode synthesis layer segmentation results can be encoded (e.g., using matrix encoding) before being forwarded from the synthesizer backend to update the learning model. The hybrid mode synthesis layer segmentation results may include reasons why the application layer does not support synthesis at the DPU. These reasons may include, for example, pixel format, downscaling factor, stacking scarcity, blending stage limitations, bandwidth limitations, and so on.
[0058] Figure 6 This document illustrates a call flowchart 600 of example communication between a CPU 602, a DPU 604, and a GPU 606, based on one or more techniques described in this disclosure. At 608, the CPU 602 may receive indications for multiple application layers used for composition at the DPU 604 and GPU 606. At 610, the CPU 602 may identify layer priorities for the multiple application layers used for composition at the DPU 604 and GPU 606. At 612, the CPU 602 may select one or more first application layers from the multiple application layers for attempted composition at the DPU 604, and select one or more second application layers from the multiple application layers for composition at the GPU 606. The selection of the one or more first application layers and the one or more second application layers may be associated with a heuristic process, for example, as described above.
[0059] In one configuration, the selection of the one or more first application layers for synthesis attempted at DPU 604 and the one or more second application layers for synthesis at GPU 606 can be based on layer priority. Specifically, the one or more first application layers for synthesis attempted at DPU 604 can correspond to higher layer priorities. The one or more second application layers for synthesis at GPU 606 can correspond to lower layer priorities. Compared to at least some of the one or more second application layers corresponding to lower layer priorities, at least some of the one or more first application layers corresponding to higher layer priorities can correspond to larger layer sizes or higher priority applications. The one or more first application layers corresponding to higher layer priorities can correspond to at least one of the following: camera applications, gaming applications, or HDR video applications.
[0060] In one configuration, the selection of one or more first application layers and one or more second application layers can be based on the order or sorting (e.g., Z-order) of the multiple application layers. One or more first application layers used for synthesis at DPU 604 may correspond to a lower order or sorting (e.g., lower Z-order). One or more second application layers used for synthesis at GPU 606 may correspond to a higher order or sorting (e.g., higher Z-order).
[0061] In one configuration, the selection of one or more first application layers and one or more second application layers can be based on at least one prior synthesis of multiple application layers at the DPU 604 and GPU 606.
[0062] In one configuration, at least some of the layers in one or more second application layers can be pre-specified (tagged) to be composed on the GPU.
[0063] At 614, CPU 602 can configure each of one or more first application layers to each of the plurality of DPU 604 stacking planes, and configure each of one or more second application layers to the GPU 606 shader. The composition of at least some of the first application layers at the plurality of DPU 604 stacking planes can at least partially overlap in time with the composition of at least some of the second application layers at the GPU 606 shader. The configuration of one or more first application layers to the plurality of DPU stacking planes can overlap in time with the configuration of one or more second application layers to the GPU shader.
[0064] At 616, CPU 602 may send to DPU 604 an instruction to configure each of one or more first application layers to each of the plurality of DPU 604 stacking planes. At 618, CPU 602 may send to GPU 606 an instruction to configure each of one or more second application layers to the GPU 605 shader. At 620, DPU 604 may perform composition on one or more first application layers. At 622, GPU 606 may perform composition on one or more second application layers. The one or more second application layers may correspond to the first batch of application layers used for composition at the GPU. At 624, CPU 602 may recognize that DPU 604 does not support composition of at least one of the one or more first application layers. At 626, if any layer is incorrectly selected, CPU may configure at least one first application layer (the attempted composition incorrectly selected at DPU 604) to the GPU 606 shader. At 628, if any layer is incorrectly selected, GPU 606 may synthesize the at least one first application layer (the attempted synthesis that was incorrectly selected at DPU 604). The at least one first application layer may correspond to a second batch of application layers used for synthesis at GPU 606.
[0065] The multiple DPU 604 stacked planes can correspond to the output synthesized by DPU 604. The GPU 606 shader can correspond to the output synthesized by GPU 605.
[0066] Figure 7 This is a flowchart 700 illustrating a processing method based on one or more techniques according to this disclosure. The method can be combined with... Figure 1 , 2 It is executed by the devices used in aspects 4-6 (e.g., devices for display processing, CPU, wireless communication devices, etc.).
[0067] At 702, the device can receive instructions for multiple application layers used for synthesis at the first and second processors. For example, refer to... Figure 6 At 608, CPU 602 can receive instructions for multiple application layers used for synthesis at the first processor 604 and the second processor 606. Furthermore, Figure 1 The processing unit 120 can execute step 702.
[0068] At 704, the device can select one or more first application layers from the plurality of application layers for attempted synthesis on a first processor, and select one or more second application layers from the plurality of application layers for synthesis on a second processor. For example, refer to Figure 6 At 612, CPU 602 can select one or more first application layers from the plurality of application layers for attempted synthesis on first processor 604, and select one or more second application layers from the plurality of application layers for synthesis on second processor 606. Furthermore, Figure 1 The processing unit 120 can execute step 704.
[0069] At point 706, the device can send each of the one or more first application layers to a first processor for synthesis, and send each of the one or more second application layers to a second processor for synthesis. The first processor and the second processor can synthesize at least some of the one or more first application layers and at least some of the one or more second application layers in parallel. For example, refer to... Figure 6 At point 614, CPU 602 can send each of the one or more first application layers to first processor 604 for compositing, and send each of the one or more second application layers to second processor 606 for compositing. Furthermore, Figure 1 The processing unit 120 can execute step 706.
[0070] Figure 8 This is a flowchart 800 illustrating a processing method based on one or more techniques according to this disclosure. The method can be combined with... Figure 1 , 2 It is executed by the devices used in aspects 4-6 (e.g., devices for display processing, CPU, wireless communication devices, etc.).
[0071] At 802, the device can receive instructions for multiple application layers used for synthesis at the first and second processors. For example, refer to... Figure 6 At 608, CPU 602 can receive instructions for multiple application layers used for synthesis at the first processor 604 and the second processor 606. Furthermore, Figure 1 The processing unit 120 can execute step 802.
[0072] In one configuration, the first processor can be a DPU. The second processor can be a GPU.
[0073] At 804, the device can identify layer priorities for the plurality of application layers used for synthesis at the DPU and GPU. The selection of one or more first application layers for attempted synthesis at the DPU and one or more second application layers for synthesis at the GPU can be based on this layer priority. For example, refer to... Figure 6At 610, CPU 602 can identify layer priorities for the multiple application layers used for composition at the DPU and GPU. Furthermore, Figure 1 The processing unit 120 can execute step 804.
[0074] At 806, the device can select one or more first application layers from the plurality of application layers for attempted synthesis on a first processor, and select one or more second application layers from the plurality of application layers for synthesis on a second processor. For example, refer to Figure 6 At 612, CPU 602 can select one or more first application layers from the plurality of application layers for attempted synthesis on first processor 604, and select one or more second application layers from the plurality of application layers for synthesis on second processor 606. Furthermore, Figure 1 The processing unit 120 can execute step 806.
[0075] In one configuration, the one or more first application layers used for attempted synthesis at the DPU may correspond to a higher layer priority, while the one or more second application layers used for synthesis at the GPU may correspond to a lower layer priority.
[0076] In one configuration, at least some of the one or more first application layers corresponding to higher layer priorities may correspond to at least one of a camera application, a game application, or an HDR video application.
[0077] In one configuration, at least some of the first application layers corresponding to a higher priority layer correspond to a larger layer size or a higher priority application, compared to at least some of the second application layers corresponding to a lower priority layer.
[0078] In one configuration, the selection of the one or more first application layers and the one or more second application layers can be based on the order or sorting of the multiple application layers. The one or more first application layers synthesized at the DPU may correspond to a lower order or sorting, while the one or more second application layers synthesized at the GPU may correspond to a higher order or sorting.
[0079] In one configuration, the selection of one or more first application layers and one or more second application layers can be associated with a heuristic process or a machine learning process.
[0080] In one configuration, the selection of the one or more first application layers and the one or more second application layers may be based on at least one prior synthesis of the plurality of application layers at the DPU and GPU.
[0081] In one configuration, at least some of the one or more second application layers can be pre-specified for compositing at the GPU.
[0082] At point 808, the device can send each of the one or more first application layers to a first processor for synthesis, and send each of the one or more second application layers to a second processor for synthesis. The first processor and the second processor can synthesize at least some of the one or more first application layers and at least some of the one or more second application layers in parallel. For example, refer to... Figure 6 At locations 614, 616, and 618, CPU 602 can send each of the one or more first application layers to first processor 604 for compositing, and send each of the one or more second application layers to second processor 606 for compositing. Furthermore, Figure 1 The processing unit 120 can execute step 808.
[0083] At 810, the device can send each of the one or more first application layers to each of the plurality of DPU overlay planes associated with the DPU, and send each of the one or more second application layers to the GPU shader associated with the GPU. For example, refer to Figure 6 At 616 and 618, CPU 602 can send each of the one or more first application layers to each of the plurality of DPU overlay planes associated with DPU 604, and send each of the one or more second application layers to the GPU shader associated with GPU 606. Furthermore, Figure 1 The processing unit 120 can execute step 810.
[0084] In one configuration, the transport from one or more first application layers to multiple DPU overlay planes may overlap in time with the transport from one or more second application layers to GPU shaders.
[0085] In one configuration, the plurality of DPU overlay planes can correspond to the output synthesized by the DPUs. The GPU shaders can correspond to the output synthesized by the GPUs.
[0086] At 812, the device can identify whether the first processor does not support the synthesis of at least one of the one or more first application layers. For example, refer to Figure 6At point 624, CPU 602 can identify whether the first processor 604 does not support the synthesis of at least one of the one or more first application layers. Furthermore, Figure 1 The processing unit 120 can execute step 812.
[0087] At 814, the device can send the at least one first application layer to the GPU shader for composition on the GPU. The at least one first application layer can be sent to the GPU shader after selecting one or more first application layers for attempted composition at the DPU. The one or more second application layers for composition at the GPU can correspond to the first batch of application layers for composition at the GPU, and the at least one first application layer can correspond to the second batch of application layers for composition at the GPU. For example, refer to... Figure 6 At 626, CPU 602 can send the at least one first application layer to the GPU shader for composition on GPU 606. Furthermore, Figure 1 The processing unit 120 can execute step 814.
[0088] In this configuration, a method or apparatus for display processing is provided. The apparatus may be a CPU or some other processor capable of performing display processing. In various aspects, the apparatus may be a processing unit 120 within device 104, or some other hardware within device 104 or another device. The apparatus may include: a unit for receiving instructions for combining a plurality of application layers at a first processor and a second processor. The apparatus may further include: a unit for selecting one or more first application layers from the plurality of application layers for attempting combination on the first processor, and a unit for selecting one or more second application layers from the plurality of application layers for combination on the second processor. The apparatus may further include: a unit for sending each of the one or more first application layers to the first processor for combination, and sending each of the one or more second application layers to the second processor for combination. The first processor and the second processor may combine at least some of the one or more first application layers and at least some of the one or more second application layers in parallel.
[0089] In one configuration, the first processor may be a DPU. The second processor may be a GPU. In one configuration, the apparatus may further include a unit for identifying layer priorities for the plurality of application layers used for compositing at the DPU and the GPU. The selection of the one or more first application layers for attempted compositing at the DPU and the one or more second application layers for compositing at the GPU may be based on the layer priorities. In one configuration, the one or more first application layers for attempted compositing at the DPU may correspond to higher layer priorities, while the one or more second application layers for compositing at the GPU may correspond to lower layer priorities. In one configuration, at least some of the one or more first application layers corresponding to the higher layer priorities may correspond to at least one of the following: a camera application, a game application, or an HDR video application. In one configuration, at least some of the one or more first application layers corresponding to the higher layer priorities may correspond to larger layer sizes or higher priority applications compared to at least some of the second application layers corresponding to the lower layer priorities. In one configuration, the selection of the one or more first application layers and the one or more second application layers may be based on the order or sorting of the plurality of application layers. The one or more first application layers attempted for synthesis at the DPU may correspond to a lower order or sequence, while the one or more second application layers used for synthesis at the GPU may correspond to a higher order or sequence. In one configuration, the selection of the one or more first application layers and the one or more second application layers may be associated with a heuristic or machine learning process. In one configuration, the selection of the one or more first application layers and the one or more second application layers may be based on at least one previous synthesis of the plurality of application layers at the DPU and the GPU. In one configuration, at least some of the one or more second application layers may be pre-specified for synthesis at the GPU. In one configuration, the apparatus may further include: sending each of the one or more first application layers to each of a plurality of DPU stacking planes associated with the DPU, and sending each of the one or more second application layers to a GPU shader associated with the GPU. In one configuration, the transfer of the one or more first application layers to the plurality of DPU stacking planes may overlap in time with the transfer of the one or more second application layers to the GPU shader. In one configuration, the plurality of DPU overlay planes may correspond to the output synthesized by the DPUs, while the GPU shaders may correspond to the output synthesized by the GPUs.In one configuration, the apparatus may further include: a unit for identifying whether the first processor does not support the composition of at least one of the one or more first application layers. In one configuration, the apparatus may further include: a unit for sending the at least one first application layer to a GPU shader for composition on the GPU. After selecting the one or more first application layers for attempted composition at the DPU, the at least one first application layer may be sent to the GPU shader. The one or more second application layers for composition at the GPU may correspond to a first batch of application layers for composition at the GPU. The at least one first application layer may correspond to a second batch of application layers for composition at the GPU. In one configuration, the apparatus may be a wireless communication device.
[0090] Therefore, the time for each blending mode compositing round can be significantly reduced without using additional hardware. Compositing at extremely high refresh rates (e.g., 240Hz or higher) can be supported without placing excessive demands on the CPU. The saved CPU processing resources can be dedicated to other purposes (e.g., application rendering). CPU efficiency cores are sufficient for compositing, even at very high refresh rates. CPU performance cores can be a scarce resource, contributing to more valuable tasks.
[0091] It should be understood that the specific order or block / step hierarchy in the processing, flowcharts, and / or calling flowcharts disclosed herein is merely one example of an exemplary method. It should be understood that these specific orders or block / step hierarchies in the processing, flowcharts, and / or calling flowcharts may be rearranged based on design preferences. Furthermore, some blocks / steps may be combined and / or omitted. Other blocks / steps may also be added. The appended method claims give the elements of various blocks / steps in an exemplary order, but this does not imply that they are limited to the given specific order or hierarchy.
[0092] To enable any person skilled in the art to implement the various aspects described herein, the foregoing description has been made regarding these aspects. Various modifications to these aspects will be apparent to those skilled in the art, and the general principles defined herein may also be applied to other aspects. Therefore, the invention is not limited to the aspects shown herein, but is consistent with the full scope of the invention disclosure, wherein, unless specifically stated otherwise, the use of the singular to modify a component does not mean "one and only one," but can mean "one or more." The term "exemplary" as used herein means "serving as an example, illustration, or description." Any aspect described herein as "exemplary" should not be construed as preferred or advantageous over other aspects.
[0093] Unless otherwise specified, the term "some" refers to one or more, and the term "or" may be interpreted as "and / or" unless the context otherwise specifies. Combinations such as "at least one of A, B, or C", "one or more of A, B, or C", "at least one of A, B, and C", "one or more of A, B, and C", and "A, B, C, or any combination thereof" include any combination of A, B, and / or C, which may include multiple A, multiple B, or multiple C. Specifically, combinations such as "at least one of A, B, or C", "one or more of A, B, or C", "at least one of A, B, and C", "one or more of A, B, and C", and "A, B, C, or any combination thereof" may be only A, only B, only C, A and B, A and C, B and C, or A and B and C, wherein any such combination may contain one or more members or some members of A, B, or C. All structural and functional equivalents of components throughout the various aspects described in this disclosure are expressly incorporated herein by reference and intended to be covered by the claims, and such structural and functional equivalents are well known or will be known to those skilled in the art. Furthermore, no disclosure herein is intended to be offered to the public, whether or not such disclosure is expressly stated in the claims. Terms such as “module,” “apparatus,” “element,” “device,” etc., are not substitutes for the term “unit.” Therefore, the constituent elements of a claim should not be construed as functional modules unless the constituent element is expressly described using the term “functional module.”
[0094] In one or more examples, the described functionality may be implemented using hardware, software, firmware, or any combination thereof. For example, although the term "processing unit" is used throughout this disclosure, such a processing unit may be implemented using hardware, software, firmware, or any combination thereof. When any functionality, processing unit, technique, or other module described herein is implemented using software, such functionality, processing unit, technique, or other module may be stored on a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium.
[0095] Computer-readable media can include computer data storage media or communication media, including any medium that facilitates the transfer of a computer program from one place to another. In this way, a computer-readable medium can generally correspond to: (1) a non-transitory tangible computer-readable storage medium; or (2) a communication medium such as a signal or carrier waveform. Data storage media can be any available medium accessible to one or more computers or one or more processors to obtain instructions, code, and / or data structures for implementing the techniques described in this disclosure. For example, but not limitingly, such computer-readable media can include RAM, ROM, EEPROM, CD-ROM or other optical disc storage, disk storage, or other magnetic storage devices. As used herein, disks and optical discs include CDs, laser discs, optical discs, DVDs, floppy disks, and Blu-ray discs, wherein disks typically magnetically copy data, while optical discs typically optically copy data using lasers. Combinations of the above should also be included within the scope of protection of computer-readable media. Computer program products can include computer-readable media.
[0096] The techniques disclosed herein can be implemented using a wide variety of devices or apparatuses, including wireless handheld devices, integrated circuits (ICs), or a set of ICs (e.g., chipsets). Various components, modules, or units are described in this disclosure to emphasize functional aspects of a device configured to perform the disclosed techniques, but they do not necessarily need to be implemented by different hardware units. Rather, as described above, individual units can be combined in any hardware unit or provided through a cooperating set of hardware units (including one or more processors as described above) combined with appropriate software and / or firmware. Therefore, as used herein, the term "processor" can refer to any of the foregoing structures or any other structures suitable for implementing the techniques described herein. Furthermore, these techniques can be fully implemented in one or more circuit or logic elements.
[0097] The following aspects are illustrative only and may be combined with, but are not limited to, other aspects or teachings described herein.
[0098] Aspect 1 is an apparatus for display processing, including at least one processor coupled to a memory, configured to: receive instructions for a plurality of application layers for synthesis at a first processor and a second processor; select one or more first application layers from the plurality of application layers for attempted synthesis at the first processor, and one or more second application layers from the plurality of application layers for synthesis at the second processor; and send each of the one or more first application layers to the first processor for synthesis, and send each of the one or more second application layers to the second processor for synthesis, wherein the first processor and the second processor can synthesize at least some of the one or more first application layers and at least some of the one or more second application layers in parallel.
[0099] Aspect 2 can be combined with aspect 1, wherein the first processor is a DPU and the second processor is a GPU.
[0100] Aspect 3 can be combined with aspect 2, wherein the at least one processor is further configured to: identify layer priorities for the plurality of application layers for synthesis at the DPU and the GPU, wherein the selection of the one or more first application layers for attempted synthesis at the DPU and the one or more second application layers for synthesis at the GPU is based on the layer priorities.
[0101] Aspect 4 can be combined with any of Aspects 2 and 3, wherein the one or more first application layers for synthesis at the DPU correspond to a higher layer priority, while the one or more second application layers for synthesis at the GPU correspond to a lower layer priority.
[0102] Aspect 5 can be combined with aspect 4, wherein at least some of the one or more first application layers corresponding to the higher priority layer correspond to at least one of the following: camera application, game application, or HDR video application.
[0103] Aspect 6 can be combined with any of Aspects 4 and 5, wherein at least some of the first application layers corresponding to the higher priority layer correspond to a larger layer size or a higher priority application compared to at least some of the second application layers corresponding to the lower priority layer.
[0104] Aspect 7 can be combined with any of Aspects 2-6, wherein the selection of the one or more first application layers and the one or more second application layers is based on the order or sorting of the plurality of application layers, and the one or more first application layers synthesized at the DPU correspond to a lower order or sorting, while the one or more second application layers synthesized at the GPU correspond to a higher order or sorting.
[0105] Aspect 8 can be combined with any of Aspects 2-7, wherein the selection of the one or more first application layers and the one or more second application layers is associated with a heuristic process or a machine learning process.
[0106] Aspect 9 can be combined with any of Aspects 2-8, wherein the selection of the one or more first application layers and the one or more second application layers is based on at least one prior synthesis of the plurality of application layers at the DPU and the GPU.
[0107] Aspect 10 can be combined with any of aspects 2-9, wherein at least some of the one or more second application layers are pre-specified for synthesis at the GPU.
[0108] Aspect 11 can be combined with any one of aspects 2-10, wherein the at least one processor is further configured to: send each of the one or more first application layers to each of the plurality of DPU overlay planes associated with the DPU, and send each of the one or more second application layers to a GPU shader associated with the GPU.
[0109] Aspect 12 can be combined with aspect 11, wherein the transport from the one or more first application layers to the plurality of DPU overlay planes overlaps in time with the transport from the one or more second application layers to the GPU shaders.
[0110] Aspect 13 can be combined with any of aspects 11 and 12, wherein the plurality of DPU overlay planes correspond to the output of DPU synthesis, and the GPU shaders correspond to the output of GPU synthesis.
[0111] Aspect 14 can be combined with any one of aspects 2-13, wherein the at least one processor is further configured to: identify whether the first processor does not support the synthesis of at least one of the one or more first application layers.
[0112] Aspect 15 can be combined with aspect 14, wherein the at least one processor is further configured to: send the at least one first application layer to a GPU shader for composition on the GPU, wherein, after selecting the one or more first application layers for attempted composition at the DPU, the at least one first application layer is sent to the GPU shader, the one or more second application layers for composition on the GPU correspond to a first batch of application layers for composition on the GPU, and the at least one first application layer corresponds to a second batch of application layers for composition on the GPU.
[0113] Aspect 16 can be combined with any of aspects 1-15, wherein the device is a wireless communication device.
[0114] Aspect 17 is a display processing method for implementing any one of Aspects 1-16.
[0115] Aspect 18 is an apparatus for display processing, comprising units for implementing the method as described in any one of aspects 1-16.
[0116] Aspect 19 is a computer-readable medium storing computer-executable code that, when executed by at least one processor, causes the at least one processor to perform the method as described in any one of aspects 1-16.
[0117] This document describes various aspects, which, along with others, fall within the scope of the appended claims.
Claims
1. An apparatus for display processing, comprising: Memory; as well as At least one processor, coupled to the memory, is configured to: Receive instructions for multiple application layers used for synthesis at the first processor and the second processor; Select one or more first application layers from the plurality of application layers for attempting synthesis at the first processor, and one or more second application layers from the plurality of application layers for synthesis at the second processor; After selecting one or more first application layers from the plurality of application layers for attempting synthesis at the first processor, it is identified that at least one of the one or more first application layers does not support synthesis by the first processor; as well as Each of the one or more first application layers is sent to the first processor for synthesis, and each of the one or more second application layers, as well as each of the at least one first application layer that is not supported for synthesis by the first processor, is sent to the second processor for synthesis. Accordingly, the first processor and the second processor are capable of synthesizing at least some of the one or more first application layers and at least some of the one or more second application layers in parallel, and the at least one unsupported first application layer is configured to be scheduled for synthesis after the one or more second application layers.
2. The apparatus according to claim 1, wherein, The first processor is a display processing unit (DPU), and the second processor is a graphics processing unit (GPU).
3. The apparatus of claim 2, wherein the at least one processor is further configured to: Identify the layer priorities for the plurality of application layers used for synthesis at the DPU and the GPU, wherein, The selection of the one or more first application layers for synthesis at the DPU and the one or more second application layers for synthesis at the GPU is based on the layer priority.
4. The apparatus according to claim 2, wherein, The one or more first application layers used for synthesis at the DPU correspond to higher layer priorities, and the one or more second application layers used for synthesis at the GPU correspond to lower layer priorities.
5. The apparatus according to claim 4, wherein, At least some of the one or more first application layers corresponding to the higher layer priority correspond to at least one of the following: camera application, game application, or high dynamic range (HDR) video application.
6. The apparatus according to claim 4, wherein, Compared to at least some of the second application layers among the one or more second application layers corresponding to the lower layer priority, at least some of the first application layers among the one or more first application layers corresponding to the higher layer priority are associated with larger layer sizes or higher priority applications.
7. The apparatus according to claim 2, wherein, The selection of the one or more first application layers and the one or more second application layers is based on the order or sorting of the plurality of application layers, and the one or more first application layers for synthesis at the DPU correspond to a lower order or sorting, and the one or more second application layers for synthesis at the GPU correspond to a higher order or sorting.
8. The apparatus according to claim 2, wherein, The selection of the one or more first application layers and the one or more second application layers is associated with a heuristic or machine learning process.
9. The apparatus according to claim 2, wherein, The selection of the one or more first application layers and the one or more second application layers is based on at least one prior synthesis of the plurality of application layers at the DPU and the GPU.
10. The apparatus according to claim 2, wherein, At least some of the one or more second application layers are pre-specified for synthesis at the GPU.
11. The apparatus of claim 2, wherein the at least one processor is further configured to: Each of the one or more first application layers is sent to each of the multiple DPU overlay planes associated with the DPU, and each of the one or more second application layers is sent to the GPU shader associated with the GPU.
12. The apparatus according to claim 11, wherein, The transfer from one or more first application layers to the plurality of DPU overlay planes overlaps in time with the transfer from one or more second application layers to the GPU shaders.
13. The apparatus according to claim 11, wherein, The multiple DPU overlay planes correspond to the outputs synthesized by the DPUs, and the GPU shaders correspond to the outputs synthesized by the GPUs.
14. The apparatus according to claim 2, wherein, In order to send each of the at least one first application layer that does not support synthesis by the first processor to the second processor, the at least one processor is further configured to: Each of the at least one first application layer is sent to the GPU shader for compositing at the GPU.
15. The apparatus according to claim 1, wherein, The device is a wireless communication device.
16. A display processing method, comprising: Receive instructions for multiple application layers used for synthesis at the first processor and the second processor; Select one or more first application layers from the plurality of application layers for attempting synthesis at the first processor, and one or more second application layers from the plurality of application layers for synthesis at the second processor; After selecting one or more first application layers from the plurality of application layers for attempting synthesis at the first processor, it is identified that at least one of the one or more first application layers does not support synthesis by the first processor; as well as Each of the one or more first application layers is sent to the first processor for synthesis, and each of the one or more second application layers, as well as each of the at least one first application layer that is not supported for synthesis by the first processor, is sent to the second processor for synthesis. Accordingly, the first processor and the second processor are capable of synthesizing at least some of the one or more first application layers and at least some of the one or more second application layers in parallel, and the at least one unsupported first application layer is configured to be scheduled for synthesis after the one or more second application layers.
17. The method according to claim 16, wherein, The first processor is a display processing unit (DPU), and the second processor is a graphics processing unit (GPU).
18. The method of claim 17, further comprising: Identify layer priorities for the plurality of application layers to be synthesized at the DPU and the GPU, wherein the selection of one or more first application layers for attempted synthesis at the DPU and one or more second application layers for synthesis at the GPU is based on the layer priorities.
19. The method of claim 17, wherein, The one or more first application layers used for synthesis at the DPU correspond to higher layer priorities, and the one or more second application layers used for synthesis at the GPU correspond to lower layer priorities.
20. The method according to claim 19, wherein, At least some of the one or more first application layers corresponding to the higher layer priority correspond to at least one of the following: camera application, game application, or high dynamic range (HDR) video application.
21. The method according to claim 19, wherein, Compared to at least some of the second application layers among the one or more second application layers corresponding to the lower layer priority, at least some of the first application layers among the one or more first application layers corresponding to the higher layer priority are associated with larger layer sizes or higher priority applications.
22. The method according to claim 17, wherein, The selection of the one or more first application layers and the one or more second application layers is based on the order or sorting of the plurality of application layers, and the one or more first application layers for synthesis at the DPU correspond to a lower order or sorting, and the one or more second application layers for synthesis at the GPU correspond to a higher order or sorting.
23. The method according to claim 17, wherein, The selection of the one or more first application layers and the one or more second application layers is associated with a heuristic or machine learning process.
24. The method of claim 17, wherein, The selection of the one or more first application layers and the one or more second application layers is based on at least one prior synthesis of the plurality of application layers at the DPU and the GPU.
25. The method according to claim 17, wherein, At least some of the one or more second application layers are pre-specified for synthesis at the GPU.
26. The method of claim 17, further comprising: Each of the one or more first application layers is sent to each of the multiple DPU overlay planes associated with the DPU, and each of the one or more second application layers is sent to the GPU shader associated with the GPU.
27. The method according to claim 26, wherein, The transfer from one or more first application layers to the plurality of DPU overlay planes overlaps in time with the transfer from one or more second application layers to the GPU shaders.
28. The method according to claim 26, wherein, The multiple DPU overlay planes correspond to the outputs synthesized by the DPUs, and the GPU shaders correspond to the outputs synthesized by the GPUs.
29. The method according to claim 17, wherein, Sending each of the at least one first application layer that is not supported by the first processor for synthesis to the second processor includes: Each of the at least one first application layer is sent to the GPU shader for compositing at the GPU.
30. A non-transitory computer-readable medium storing computer-executable code, said code, when executed by at least one processor, causing said at least one processor to perform the following operations: Receive instructions for multiple application layers used for synthesis at the first processor and the second processor; Select one or more first application layers from the plurality of application layers for attempting synthesis at the first processor, and one or more second application layers from the plurality of application layers for synthesis at the second processor; After selecting one or more first application layers from the plurality of application layers for attempting synthesis at the first processor, it is identified that at least one of the one or more first application layers does not support synthesis by the first processor; as well as Each of the one or more first application layers is sent to the first processor for synthesis, and each of the one or more second application layers, as well as each of the at least one first application layer that is not supported for synthesis by the first processor, is sent to the second processor for synthesis. Accordingly, the first processor and the second processor are capable of synthesizing at least some of the one or more first application layers and at least some of the one or more second application layers in parallel, and the at least one unsupported first application layer is configured to be scheduled for synthesis after the one or more second application layers.