Adjusting memory bandwidth utilization on demand according to display service requirements

The display controller prefetch data and bandwidth adjustment circuit reduces the memory bandwidth of other clients, solving the visual artifacts caused by temporary stopping of memory access, ensuring the quality of service and power consumption management of the display.

CN119968673APending Publication Date: 2025-05-09ADVANCED MICRO DEVICES INC +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380067311.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-29
Filing Date
2023-09-12
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

In multi-workload scenarios, high frequency operation of memory subsystems may lead to increased power consumption, and temporarily stopping memory access may introduce visual artifacts, making it difficult to find a suitable time for blocking periods.

Method used

The display controller prefetches enough data to deal with temporary cessation of memory access, ensuring quality of service by temporarily increasing memory bandwidth. At the same time, the bandwidth adjustment circuit reduces the memory bandwidth of other clients to ensure that the display controller obtains sufficient bandwidth.

Benefits of technology

By prefetching data and adjusting memory bandwidth, the visual artifacts caused by temporary stopping of memory access is solved, and the quality of service of the monitor is maintained, avoiding unnecessary increase in power consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119968673A_ABST
    Figure CN119968673A_ABST
Patent Text Reader

Abstract

Systems, apparatuses, and methods for prefetching data by a display controller are presented. From time to time, a performance state change of the memory is performed. During such a change, a memory clock frequency change of a memory subsystem (220) for driving pixels to a frame buffer (230) of the display device (250) is stored. During the performance state change, memory access may be temporarily blocked. To maintain a desired quality of service of the display, the display controller (150) is configured to prefetch data prior to the performance state change. To ensure that the display controller has sufficient memory bandwidth to complete the prefetch, bandwidth reduction circuitry (112A, 112N) in a client (205) of the system is configured to temporarily reduce the memory bandwidth of the corresponding client.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Related technical description

[0002] Many types of computer systems include display devices for displaying images, video streams, and data. Therefore, these systems often include functionality for generating and / or manipulating image and video information. Typically, in digital imaging, the smallest item of information in an image is referred to as a "picture element," and more generally as a "pixel."

[0003] Some systems include multiple separate displays. In these systems, multi-display technology enables a single graphics processing unit (GPU) (or other device, such as an accelerated processing circuit (APU) or other type of system on chip (SOC) or any application-specific integrated circuit (ASIC) with a display controller) to simultaneously support multiple independent display outputs. In one example, a computing system can independently connect multiple high-resolution displays to a large integrated display surface to provide an expanded visual workspace. Gaming, entertainment, medical, audio and video editing, business and other applications can take advantage of the expanded visual workspace and increase multitasking opportunities.

[0004] For one or more supported displays, the video subsystem maintains a corresponding frame buffer storing data (such as one or more video frames), which may be stored in a dynamic random access memory (DRAM). For each supported display, the video controller reads data via a given one of the one or more DRAM interfaces for accessing the corresponding frame buffer. The memory clock is typically used to control the data rate at which the frame buffer within the DRAM is accessed. In some cases, in order to provide a physical connection for transmitting a pixel bit stream from the frame buffer to a display device, the computer is directly connected to the display device via an interface such as DisplayPort (DP), embedded DisplayPort (eDP), High-Definition Multimedia Interface (HDMI), or other types of interfaces. In a specific implementation, the bandwidth limit of the video stream sent from the computer to the display device will be the maximum bit rate of the DisplayPort, embedded DisplayPort, or HDMI cable.

[0005] In a scenario where multiple workloads (e.g., game rendering, video processing) are accessing the memory subsystem, the memory subsystem may be set to a relatively high frequency (e.g., its maximum possible frequency) to ensure that the operating frequency of the memory subsystem can handle a large number of reads and writes. In some cases, when the memory subsystem is not overstressed, the system may expect to reduce the memory clock frequency in order to reduce power consumption. Changing the memory clock frequency may require performing a training session, a configuration / mode change, or another action on the memory interface that requires access to the temporarily stopped memory. Stopping all memory accesses may be referred to as a "lockout period." Due to this lockout period when the memory interface needs to be retrained or when other types of mode changes need to be performed, it may be difficult or impossible to find a convenient time to stop all memory accesses without introducing visual artifacts on any display in the display. One solution to this problem is for the display controller to prefetch enough data to resolve the temporary stop of memory access. For example, this may require the display controller to temporarily double its memory bandwidth. One way to ensure that the increased memory bandwidth is available to the display controller is to statically allocate the amount of memory bandwidth to the display controller. However, this method reduces the bandwidth available to other clients, even if the display controller does not always need the increased amount of bandwidth. Therefore, it is desirable to have an improved system and method for managing memory bandwidth. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] Advantages of the methods and mechanisms described herein may be better understood by referring to the following description in conjunction with the accompanying drawings, in which:

[0007] Figure 1 is a block diagram of a specific implementation of a computing system.

[0008] Figure 2 is a block diagram of a specific implementation of a computing system.

[0009] Figure 3 A timing diagram of one specific implementation of the timing of memory clock frequency updates and pre-fetching of display data in a computing system.

[0010] Figure 4 A method for pre-fetching display data before a performance state of a memory storing the display data is changed. DETAILED DESCRIPTION

[0011] In the following description, many specific details are set forth to provide a thorough understanding of the methods and mechanisms presented herein. However, it will be appreciated by those of ordinary skill in the art that each specific implementation may be practiced without these specific details. In some cases, well-known structures, components, signals, computer program instructions, and techniques are not shown in detail to avoid blurring the methods described herein. It should be understood that, for simplicity and clear illustration, the elements shown in the drawings are not necessarily drawn to scale. For example, the size of some of these elements may be enlarged relative to other elements.

[0012] Disclosed are systems, apparatus, and methods for prefetching data by a display controller in a computing system. From time to time, a performance state change of a memory is performed. During such a change, a memory clock frequency of a memory subsystem storing a frame buffer used to drive pixels to a display device is changed. During the performance state change, memory access may be temporarily blocked. In order to maintain the desired quality of service of the display, the display controller is configured to prefetch data before the performance state change. In order to ensure that the display controller has sufficient memory bandwidth to complete the prefetch, a bandwidth reduction circuit in a client of the system is configured to temporarily reduce the memory bandwidth available or otherwise consumed by the corresponding client. By reducing memory accesses generated by other clients, other clients are prevented from competing with the display controller for memory bandwidth, which may cause the display controller to fail to meet the desired quality of service requirements.

[0013] Reference now Figure 1 , a block diagram of one implementation of a computing system 100 is shown. In one implementation, computing system 100 includes at least processors 105A-N, input / output (I / O) interface 120, bus 125, memory controller 130, network interface 135, memory device 140, display controller 150, display 155, and control circuit 160. In other implementations, computing system 100 includes other components and / or computing system 100 is arranged in a different manner.

[0014] Display controller 150 represents any number of display controllers included in system 100, where the number varies depending on the specific implementation. Display controller 150 is configured to drive corresponding display 155, where display 155 represents any number of displays. In some specific implementations, a single display controller drives multiple displays. As shown in the example, display controller 150 includes a buffer 152 for storing frame data to be displayed.

[0015] In one implementation, control circuit 160 determines whether a condition for performing a power state change (also referred to as a "Pstate" change) has been detected. In various implementations, a change in Pstate results in a change in the operating frequency and / or power consumption of a given device. For example, an increase in Pstate may require an increase in the operating frequency and voltage supplied to the device. Conversely, a decrease in Pstate may require a decrease in the operating frequency and / or voltage supplied to the device.

[0016] When a condition for executing a power state change of the memory device 140 is detected, the control circuit 160 determines when to implement the power state change. Before implementing the power state change, the control circuit 160 is configured to transmit a signal 116 to the display controller 150. In response to the signal 116, the display controller 150 is configured to prefetch additional data into the buffer 152 in anticipation of an upcoming memory lockout period (i.e., a period during which memory access is not allowed). This will prevent interruptions in the display data that may cause visual artifacts, etc. Therefore, the memory bandwidth requirements of the display controller may be temporarily increased. When the display controller 150 is caused to prefetch additional data from the memory 140, there may not be enough bandwidth available due to many other clients (e.g., processor 105, I / O 120, etc.) that generate memory accesses. In other words, in order to complete the prefetch, the display controller 150 may need a bandwidth of X. However, other clients in the system may be allocated various amounts of memory bandwidth, making X bandwidth unavailable to the display controller 150. Therefore, to ensure that sufficient bandwidth is available to the display controller 150, a bandwidth adjustment circuit 112 (eg, Figure 1 112 in FIG. 11). The control circuit 160 is configured to transmit a signal / indication 114 to each of the bandwidth adjustment circuits. In response to the indication, the bandwidth adjustment circuit causes the corresponding client to temporarily reduce its memory bandwidth during the period when the display controller 150 increases its bandwidth.

[0017] The bandwidth adjustment circuit 112 includes circuitry configured to cause a corresponding client to reduce memory accesses transmitted to the memory 140. In some implementations, the bandwidth adjustment circuitry is configured to cause one corresponding client to reduce memory bandwidth. In other implementations, the bandwidth adjustment circuitry 112 is configured to cause more than one client to reduce memory bandwidth. In some implementations, the bandwidth adjustment circuitry is part of the client circuitry, while in other implementations, the bandwidth adjustment circuitry is implemented separately from a given client. These and other implementations are possible and contemplated. In this way, sufficient bandwidth is provided to the display controller to pre-fetch additional data.

[0018] In one specific implementation, the power state change involves adjusting the memory clock frequency of one or more memory devices 140. The control circuit 160 may be implemented using any suitable combination of circuits, memory elements, and program instructions. It should be noted that the control circuit 160 may also be referred to by other names, such as system management controller, system management circuit, system controller, controller, etc. Although in Figure 1 105A-N, but it should be understood that this is merely representative of one implementation. In other implementations, the system 100 may include multiple control circuits 160 located in any suitable location. Moreover, in another implementation, the control circuit 160 is implemented by one of the processors 105A-N.

[0019] Processors 105A-N represent any number of processors included in system 100. In one implementation, processor 105A is a general purpose processor, such as a central processing unit (CPU). In this implementation, processor 105A executes a driver 110 (e.g., a graphics driver) for communicating with and / or controlling the operation of one or more other processors in system 100. It should be noted that driver 110 may be implemented using any suitable combination of hardware, software, and / or firmware, depending on the implementation.

[0020] In one implementation, processor 105N is a data parallel processor with a highly parallel architecture. Data parallel processors include graphics processing circuits (GPUs), digital signal processors (DSPs), field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), etc. In some implementations, processors 105A-N include multiple data parallel processors. In one implementation, processor 105N is a GPU that renders pixel data into a frame buffer 142 representing an image. The pixel data is then provided to display controller 150 to be driven to display 155.

[0021] The memory controller 130 represents any number and type of memory controllers that can be accessed by the processors 105A-N. Although the memory controller 130 is shown as being separate from the processors 105A-N, it should be understood that this merely represents one possible implementation. In other implementations, the memory controller 130 may be embedded within one or more of the processors 105A-N, and / or the memory controller 130 may be located on the same semiconductor chip as one or more of the processors 105A-N. The memory controller 130 is coupled to any number and type of memory devices 140. The memory devices 140 represent any number and type of memory devices. For example, the types of memory in the memory devices 140 include dynamic random access memory (DRAM), static random access memory (SRAM), graphic double data rate (GDDR) synchronous DRAM (SDRAM), NAND flash memory, NOR flash memory, ferroelectric random access memory (FeRAM), etc.

[0022] The I / O interface 120 represents any number and type of I / O interfaces (e.g., a peripheral component interconnect (PCI) bus, a PCI-Extended (PCI-X), a PCIE (PCI Express) bus, a Gigabit Ethernet (GBE) bus, a Universal Serial Bus (USB)). Various types of peripheral devices (not shown) are coupled to the I / O interface 120. Such peripheral devices include (but are not limited to) a display, a keyboard, a mouse, a printer, a scanner, a joystick or other type of game controller, a media recording device, an external storage device, a network interface card, etc. The network interface 135 is capable of receiving and sending network messages over a network.

[0023] In various implementations, computing system 100 is any of a computer, a laptop, a mobile device, a game console, a server, a streaming device, a wearable device, or various other types of computing systems or devices. It should be noted that the number of components of computing system 100 varies from implementation to implementation. For example, in other implementations, there are more than 100 components. Figure 1 More or less of each component shown. It should also be noted that in other specific implementations, the computing system 100 includes Figure 1 Other components (e.g., phase-locked loops, voltage regulators) are not shown in the figure to avoid cluttering the figure. In addition, in other specific implementations, the computing system 100 is configured to be different from Figure 1 Constructed in the manner shown.

[0024] Now turn to Figure 2, a block diagram of one implementation of the system 200 is shown. In one implementation, the system 200 includes a processing element 205, a control circuit 210, a structure 215, a memory subsystem 220, a display controller 150, a pre-fetch controller 240, and a display device 250. Although the pre-fetch controller 240 is shown as being included in the display controller 150, this does not exclude the pre-fetch controller 240 from being integrated into the display device 250. In other words, depending on the implementation, the pre-fetch controller 240 may be located inside or outside the display device 250. Similarly, although the buffer 245 is shown as being located inside the pre-fetch controller 240, in other implementations, this does not exclude the buffer 245 from being located outside the pre-fetch controller 240. In general, the display controller 150 receives video images and frame data from various sources, processes the data, and then sends the data out in a format compatible with its target display.

[0025] Processing element 205 represents any number, type, and arrangement of processing resources (e.g., CPU, GPU, FPGA, ASIC). Figure 1 As illustrated, bandwidth adjustment circuits 112A-112N are associated with processing elements 205 that generate memory accesses. In this example, each bandwidth adjustment circuit 112A-112N is associated with a corresponding processing element (PE) 262A-262N. In addition, a queue 262A-262N is associated with each processing element 262, which is configured to store pending memory accesses generated by the processing element 262. In various specific implementations, the bandwidth adjustment circuit 112 is configured to control the generation of memory accesses by the processing element 262 and to serve any one or both of the already generated memory accesses stored in the pending queues 264A-264N. It should be noted that various possible arrangements of the processing elements 262 and the queues 264 are possible and contemplated. The control circuit 210 includes any suitable combination of execution circuits, circuits, memories, and program instructions. Although the control circuit 210 is shown as a separate component from the processing element 205, this represents a specific specific implementation. In another implementation, the functionality of control circuit 210 is performed at least in part by processing element 205. Fabric 215 represents any number and types of buses, communication devices / interfaces, interconnects, and other interface modules used to connect the various components of system 200 together.

[0026] In one implementation, the processing element 205 generates pixel data for display on the display device 250. In one implementation, the pixel data is written by the processing element 205 to the frame buffer 230 in the memory 220 and then driven from the frame buffer 230 to the display device 250. In one implementation, the pixel data stored in the frame buffer 230 represents a frame of a video sequence. In another implementation, the pixel data stored in the frame buffer 230 represents the screen content of a laptop or desktop personal computer (PC). In another implementation, the pixel data stored in the frame buffer 230 represents the screen content of a mobile device (e.g., a smart phone, a tablet).

[0027] The memory subsystem 220 includes any number and type of memory controllers and memory devices. In one specific implementation, the memory subsystem 220 is capable of operating at various different clock frequencies that can be adjusted according to various operating conditions. However, when implementing a memory clock frequency change, memory training is typically performed to modify various parameters, adjust the characteristics of signals generated for data transfer, etc. For example, the phase, delay and / or voltage level of various memory interface signals are tested and adjusted during memory training. Various signal transmissions can be performed between the memory controller and the memory in order to train these memory interface signals. During this training, memory access is typically stopped. When modifying the memory clock frequency, it may be challenging to find an appropriate time to perform this memory training.

[0028] In one implementation, the control circuit 210 is configured to cause a performance state change of the memory subsystem 220. When the performance state change is to be performed, the control circuit 210 causes the display controller 150 to initiate prefetching of display data from the memory 220 prior to the performance state change. This causes a memory training to be performed that temporarily blocks access to the memory 220 while the performance state of the memory 220 is changing. By causing the display controller 150 to prefetch the display data (via the prefetch controller 240), the display controller is not deprived of video data during the training. The prefetched data (e.g., pixel data) is stored in a buffer 245 of the prefetch controller 240 and driven to the display device 250.

[0029] In one implementation, the control circuit 210 includes a memory bandwidth monitor 212, a tracking circuit 213, and a frequency adjustment circuit 214. The memory bandwidth monitor 212, the tracking circuit 213, and the frequency adjustment circuit 214 may be implemented using any combination of circuits, execution circuits, and program instructions. Moreover, in another implementation, the memory bandwidth monitor 212, the tracking circuit 213, and the frequency adjustment circuit 214 are separate circuits from the control circuit 210, rather than being part of the control circuit 210. In other implementations, the control circuit 210 may include other arrangements of components that perform similar functionality as the memory bandwidth monitor 212, the tracking circuit 213, and the frequency adjustment circuit 214.

[0030] In one specific implementation, the memory bandwidth monitor 212 compares the real-time memory bandwidth requirement of the memory subsystem 220 with the memory bandwidth available at the current memory clock frequency. If the available memory bandwidth at the current memory clock frequency differs from the real-time memory bandwidth requirement by more than a threshold, the control circuit 210 changes the frequency of one or more clocks of the memory subsystem 220.

[0031] In one implementation, the control circuit 210 sends a signal to the pre-fetch controller 240 via a sideband interface 247. It should be noted that the sideband interface 247 is separate from the main interface 242 used to pass pixels to the pre-fetch controller 240. In one implementation, the main interface 242 is an embedded display port (eDP) interface. In other implementations, the main interface 242 is compatible with any of a variety of other protocols. Sending signals through the sideband interface 247 allows the timing and scheduling of pre-fetches to be performed in a relatively short time. This is in contrast to the traditional method of sending requests through the main interface 242, which may result in a delay of several frames. Figure 2 Also shown is a signal 249 transmitted by the control circuit 210 to the memory subsystem 220 that is configured to cause the memory subsystem to change its current Pstate.

[0032] Once the prefetch controller 240 has completed prefetching data from the memory 220, the frequency adjustment circuit 214 generates a command to program the clock signal generator 225 to generate a memory clock at a different frequency. In other specific implementations, the control circuit 210 includes other logic and / or circuit arrangements to adjust the memory clock frequency. As used herein, the terms "logic" and "unit" refer to circuits or circuits configured to perform the described functions. For example, in another specific implementation, the tracking circuit 213 and the frequency adjustment circuit 214 are combined together into a single circuit. Other arrangements of circuits, processing elements, execution circuits, interface circuits, program instructions and other components for implementing the functionality of the control circuit 210 are possible and contemplated.

[0033] System 200 can be any of various types of computing systems. For example, in one implementation, system 200 includes a laptop computer connected to an external display. In this implementation, display device 250 is an internal display of the laptop computer, and display device 270 is an external display. In another implementation, system 200 includes a mobile device connected to an external display. In this implementation, display device 250 is an internal display of the mobile device, and display device 270 is an external display. Other scenarios that employ components of system 200 to implement the techniques described herein are possible and contemplated.

[0034] Reference now Figure 3 , shows a timing diagram 300, which shows waveforms of one specific implementation of the timing of memory clock frequency updates for a multi-display system. In the example shown, a signal is generated that enables the display controller to temporarily increase the memory bandwidth before the Pstate of the memory device changes. As shown, Figure 3 The Pstate before change signal generated when it is determined that a memory Pstate change will occur is illustrated.Such a determination may be made by a control circuit (eg, 160 or 210) including a power management circuit.

[0035] At time 312, the Pstate before change signal 302 is indicated. It is noted that although the discussion describes various signals and indications as "asserted" and / or "transmitted", such assertion / transmission takes various forms depending on the specific implementation. For example, in some implementations, the assertion of the signal is achieved by causing the signal to reach a specific value or voltage level. In other implementations, the assertion of the signal or indication is performed by writing a specific value to a register or memory location. All such implementations are possible and contemplated. In various implementations, this can be a signal asserted by a controller. In response to detecting the signal 302, one or more bandwidth throttling signals 304 are generated at time 314. In another implementation, bandwidth throttling may also be asserted by the control circuit directly prior to initiating the Pstate before change. The amount of time that elapses between the assertion of the signal 302 and the signal 304 varies depending on the specific implementation. The bandwidth throttling signal (e.g., Figure 1 The signal 114 in is transmitted to one or more circuits configured to generate memory access. In various specific implementations, corresponding to such as Figure 1The bandwidth reduction circuit of the bandwidth adjustment 112 or other circuits in the memory detects the bandwidth throttling signal and causes the corresponding memory access generating device to temporarily reduce the rate at which the memory access is generated. In some embodiments, all memory accesses generated by the corresponding circuit are temporarily stopped (i.e., the rate becomes zero). In other embodiments, the rate is reduced or otherwise limited, but does not become zero. In such embodiments, the device is allowed to generate memory accesses, but the rate is limited or otherwise reduced. The duration of the reduction (or "throttling") varies depending on the specific implementation. In some embodiments, the duration is a fixed amount of time (which may be programmable), after which memory access generation is no longer limited. In other embodiments, the duration lasts for a period of time that is determined based on another signal indicating that prefetching has been completed. A variety of such embodiments are possible and contemplated.

[0036] After asserting the bandwidth throttling signal 304, at time 316, the control circuit (e.g., control circuit 160, control circuit 210) transmits a prefetch signal 306 to the display controller. In some implementations, the prefetch signal 306 may be transmitted simultaneously with the assertion of the bandwidth throttling signal 304. In other implementations, there is a delay between the assertion of the signal 304 and the signal 306. In response to the assertion of the prefetch signal 306, the display controller (e.g., 150, 240) initiates prefetching data from the memory. As described above, in the process of prefetching data from the memory, other memory access generating clients temporarily reduce their memory bandwidth to ensure that the display controller has the desired bandwidth increase. In this way, the desired quality of service (QoS) of the data being displayed can be maintained. After the display controller completes its access to the memory, the bandwidth throttling 304 is de-asserted, and then the control circuit causes the Pstate of the memory to change. In the example shown, the controller asserts the Pstate change signal 308 at time 318. In various implementations, the control circuit (160, 210) also transmits or stores an indication of the new Pstate and clock frequency to which the memory is to be transitioned. In response to the Pstate change signal 308 at time 318, the memory subsystem enters the training period described above. As indicated, many memory devices (e.g., Graphics Double Data Rate (GDDR) Synchronous Dynamic Random Access Memory (SDRAM) devices) require memory training when the memory clock frequency is changed. For these memory devices, memory training is performed as part of the memory clock frequency change. After a period of time, the memory training is completed at time 320, and the memory (subsystem) reaches a stable state at the new Pstate. At this point, access to the memory is no longer blocked (i.e., the memory lockout period ends).

[0037] Reference now Figure 4 , a specific implementation of a method 500 for performing a display controller prefetch before a memory clock frequency change is shown. For the purpose of discussion, the steps in this specific implementation are shown in a sequential order. However, it should be noted that in each specific implementation of the described method, one or more of the described elements are performed simultaneously, in a different order than shown, or omitted entirely. Other additional elements are also performed as needed. Any of the various systems or devices described herein is configured to implement the method 500.

[0038] exist Figure 4 In a specific implementation, a control circuit (such as Figure 1 The control circuit 160 or Figure 2 The control circuitry 210 in the memory subsystem determines that one or more conditions are satisfied that cause a change in the memory clock frequency of the memory subsystem (block 405). As an example, an increase or decrease in required memory bandwidth may be detected based on tasks being executed (or tasks queued for execution), thermal conditions, etc. For example, if an increase in memory accesses is detected, the memory clock frequency may be increased to increase the rate at which memory accesses can be completed. Conversely, if a decrease in the number of memory accesses is detected, the memory clock frequency may be decreased to reduce power consumption. Numerous such examples are possible and are contemplated. In response to detecting the condition, a signal (e.g., Figure 3A signal 302 is provided that causes one or more bandwidth reduction circuits 112 to temporarily reduce the memory bandwidth of the corresponding client. In various specific implementations, this reduction can be achieved by preventing one or more pending memory accesses from being selected for service. For example, in some specific implementations, the client is configured to store the generated memory accesses in a queue or other location (e.g., an output or pending queue), where the memory accesses are then selected and transmitted to the memory subsystem for service. In some specific implementations, the reduction in bandwidth is achieved by causing the corresponding client to temporarily stop or slow down the generation of the client's memory accesses. In one specific implementation, the change in the memory clock frequency of the memory subsystem is performed as part of the power state change. One or more conditions that trigger the change in the memory clock frequency may vary with the specific implementation. For example, the condition may be triggered in response to detecting an increased memory bandwidth requirement. For example, a task corresponding to a particular type of application may have a higher demand for bandwidth. In response, the Pstate of the memory is indicated to increase. As another example, one or more processing circuits in the computing system are detected to be in an idle condition or otherwise have a reduced demand for memory bandwidth. In response, the Pstate of the memory begins to be reduced to reduce the power consumption of the system. In other specific implementations, other conditions may cause the memory clock frequency to change. For example, in one specific implementation, connecting or disconnecting alternating current (AC) power or direct current (DC) power may cause the memory clock frequency to change. Depending on the power source, there may be different allowable clock ranges. In another specific implementation, a change in the temperature of a host system or device may trigger a desired change in the memory clock frequency. For example, if the temperature of the host system / device exceeds a first threshold, the control circuit will attempt to reduce power consumption in order to reduce the temperature. One way to reduce power consumption is by reducing the memory clock frequency. In another specific implementation, if the temperature drops below a second threshold, the control circuit may increase the memory clock frequency because doing so will not overheat the system / device. In still another specific implementation, if there is a requested performance increase, or the performance increase is otherwise considered desirable (e.g., increasing computing speed, frame rate of video display, etc.), the control circuit will attempt to increase performance by increasing the memory clock frequency. Other conditions for changing the memory clock frequency are possible and can be expected.

[0039] In some specific implementations, the conditions for triggering the change of the memory clock frequency can be event-driven. For example, in various specific implementations, when the throughput exceeds or is lower than a certain threshold, the memory controller issues an event related to the throughput. Such events can be monitored during a programmable time window, or filtered in time in some way. There can also be a mechanism based on software, firmware or hardware, which knows or predicts that the workload needs resources before scheduling or executing the workload when the workload is submitted. Similarly, when the workload is completed, the mechanism knows what resources are no longer needed (that is, the workload in question has been completed and no longer needs resources). Moreover, similar mechanisms can solve periodic workloads. In another specific implementation, a real-time operating system (RTOS) can know the deadline, and the RTOS can select a more preferred clock according to the approaching deadline.

[0040] In response to detecting the condition that the memory clock frequency is changed (405), the control circuit generates a bandwidth throttling signal, which is then detected by one or more bandwidth reduction circuits in the computing system. As described above, the detection of the bandwidth reduction signal causes one or more devices in the computing system to reduce the memory access rate it transmits to the memory system. The control circuit then generates 415 or otherwise transmits a prefetch signal (e.g., such as Figure 2 In response to detecting the prefetch signal, the display controller initiates prefetching display data from the memory subsystem. After the display controller completes (420) the prefetching of data, the control circuit (160, 210) initiates or otherwise causes a change in the Pstate of the memory. In various implementations, the completion of the prefetch (420) is determined based on the elapse of a given time period (which may be programmable). In other implementations, the display controller may transmit an indication of the completion of the prefetch. In such an implementation, the display controller may transmit the indication in response to receiving the prefetch data or otherwise determining that the prefetching of data from the memory is complete and transferred to the display controller. In other words, even if all the prefetched data has not yet arrived at the display controller, no further access to the memory is deemed necessary. These and other implementations are possible and contemplated.

[0041] In response to the completion of the display controller prefetching data, bandwidth throttling is released 422 (i.e., bandwidth throttling stops), and the control circuit initiates a Pstate change for the memory. In various implementations, the Pstate change includes changing the memory clock frequency (block 425). In one implementation, memory training is performed as part of the memory clock frequency update. After the memory clock frequency update and training are completed (430), memory access can be performed again. In some implementations, the control circuit (e.g., 160, 210) transmits a signal to the bandwidth reduction circuit (112) that causes the bandwidth reduction circuit to stop throttling the memory bandwidth of the corresponding device. In other implementations, as described above, bandwidth throttling continues for a given period of time. It should be noted that method 400 can be repeated each time a condition for changing the memory clock frequency is detected.

[0042] In various specific implementations, the program instructions of the software application are used to implement the methods and / or mechanisms described herein. For example, it is envisioned that the program instructions can be executed by a general-purpose processor or a special-purpose processor. In various specific implementations, such program instructions are represented by a high-level programming language. In other specific implementations, the program instructions are compiled from the high-level programming language into binary, intermediate or other forms. Alternatively, program instructions describing the behavior or design of the hardware are written. Such program instructions are represented by a high-level programming language such as C. Alternatively, a hardware design language (HDL) such as Verilog is used. In various specific implementations, the program instructions are stored on any non-transient computer-readable storage medium in a variety of non-transient computer-readable storage media. The storage medium can be accessed by a computing system during use to provide the computing system with program instructions for program execution. Typically, such a computing system includes at least one or more memories and one or more processors configured to execute program instructions.

[0043] It should be emphasized that the above specific implementation is only a non-limiting example of a specific implementation. Once the above disclosure is fully understood, many variations and modifications will become apparent to those skilled in the art. It is intended that the following claims be interpreted as covering all such variations and modifications.

Claims

1. A device, comprising: control circuitry, wherein in response to determining that a condition for changing a performance state of a memory subsystem is satisfied, the control circuitry is configured to: reducing memory bandwidth of clients configured to generate memory accesses to the memory subsystem; and A display controller is caused to pre-fetch display data from the memory subsystem. 2 . The apparatus of claim 1 , wherein in order to reduce the memory bandwidth, the control circuit is configured to transmit an instruction to a bandwidth adjustment circuit corresponding to the client. 3 . The apparatus of claim 1 , wherein after the pre-fetching, the control circuit is configured to cause the memory subsystem to enter a training period. The apparatus of claim 3 , wherein access to memory is blocked during the training period. The apparatus of claim 4 , wherein the reduction of memory bandwidth ceases after the training period is completed.

6. The apparatus of claim 5, wherein determining that the condition for changing the performance state of the memory subsystem is satisfied comprises one or more of: detecting an idle condition of a client, detecting an increased memory bandwidth requirement of a client, detecting a temperature change, determining that a memory bandwidth requirement differs from available memory bandwidth at a current memory clock frequency by more than a threshold, or detecting a requested performance increase. 7 . The apparatus of claim 1 , wherein the control circuit is configured to cause the display controller to prefetch the display data by transmitting a signal to the display controller.

8. The apparatus of claim 7, wherein the control circuit is configured to reduce the memory bandwidth within a given time period.

9. A method comprising: In response to satisfying a condition for changing a performance state of the memory subsystem: reducing memory bandwidth of clients configured to generate memory accesses to the memory subsystem; and A display controller is caused to pre-fetch display data from the memory subsystem.

10. The method according to claim 9, wherein in order to reduce the memory bandwidth, the method comprises: An indication is transmitted to a bandwidth adjustment circuit corresponding to the client.

11. The method according to claim 9, wherein after the pre-fetching, the method comprises: The memory subsystem is caused to enter a training period.

12. The method of claim 11, wherein during the training period, access to memory is blocked. The method of claim 12 , wherein the reduction of memory bandwidth ceases after the training period is completed.

14. The method of claim 13, wherein determining that the condition for changing the performance state of the memory subsystem is satisfied comprises one or more of: detecting an idle condition of a client, detecting an increased memory bandwidth requirement of a client, detecting a temperature change, determining that a memory bandwidth requirement differs from available memory bandwidth at a current memory clock frequency by more than a threshold, or detecting a requested performance increase.

15. The method according to claim 9, further comprising: The display controller is caused to pre-fetch the display data by transmitting a signal to the display controller.

16. The method according to claim 15, wherein the method comprises: Reduces memory bandwidth over a given period of time.

17. A system, comprising: Memory subsystem; one or more clients configured to generate memory accesses to the memory subsystem; and Display controller; control circuitry, wherein in response to determining that a condition for changing a performance state of the memory subsystem is satisfied, the control circuitry is configured to: reducing memory bandwidth of one or more of the clients; and The display controller is caused to pre-fetch display data from the memory subsystem.

18. The system of claim 17, wherein the system further comprises bandwidth reduction circuits corresponding to the one or more clients, and the control circuit is configured to reduce the memory bandwidth of the one or more of the clients by transmitting an indication to the bandwidth reduction circuit.

19. The system of claim 18, wherein after the display controller pre-fetches display data, the control circuit is configured to cause the memory subsystem to enter a training period.

20. The system of claim 19, wherein access to memory is blocked during the training period.