Cache allocation method and electronic device
By identifying thread groups and optimizing the cache limit, the problem of high power consumption caused by unreasonable cache resource allocation was solved, achieving more efficient cache resource utilization and performance improvement.
Patent Information
- Application Number
- CN202510243057.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2045-03-03
AI Technical Summary
Traditional cache allocation methods lead to unreasonable allocation of cache resources, resulting in high power consumption.
By identifying thread groups and considering the memory stagnation that occurs when the CPU executes memory access instructions in each cycle, the cache limits for critical and non-critical thread groups are periodically optimized and adjusted to ensure a reasonable allocation of cache resources between critical and non-critical thread groups.
It reduces the overall power consumption of electronic devices and improves the utilization and performance of cache resources.
Smart Images

Figure CN119739536B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of cache management technology, and in particular to a cache allocation method and electronic device. Background Technology
[0002] A cache is a type of memory that enables high-speed data exchange. Its working principle is as follows: when the CPU requests to read data, it first attempts to search the cache (which typically includes a multi-level cache structure). If the data exists at a certain level of the cache, it is read directly and sent to the CPU; this process is called a cache hit. If the data is not found in the cache (i.e., a cache miss), it is read from the relatively slower main memory.
[0003] Cache allocation refers to the distribution of cache resources to different processor cores, threads, or processes according to a certain mechanism or proportion. However, traditional cache allocation methods suffer from the technical problem of high power consumption due to unreasonable allocation of cache resources. Summary of the Invention
[0004] This application provides a cache allocation method and an electronic device to solve the technical problem of high power consumption caused by unreasonable cache resource allocation.
[0005] To achieve the above objectives, the embodiments of this application adopt the following technical solutions:
[0006] Firstly, a cache allocation method is provided, which includes:
[0007] A first thread group and a second thread group are determined for the current scenario. The importance of the first thread in the first thread group is higher than that of the second thread in the second thread group. Based on the first information corresponding to each preset period, the cache limits corresponding to the first and second thread groups are periodically optimized and adjusted to obtain the first target cache limit for the first thread group and the second target cache limit for the second thread group. The first information corresponding to each period is used to characterize the memory stagnation that occurs when the CPU executes memory access instructions within that period, and the first target cache limit is higher than the second target cache limit.
[0008] It should be understood that the first thread group is the critical thread group, and the second thread group is the non-critical thread group. Within the critical thread group, the first thread is more important than the second thread in the non-critical thread group. The first piece of information is the memory stall situation generated when the CPU executes memory access instructions each cycle. Based on this information, the cache limits for each of the critical and non-critical thread groups are periodically optimized and adjusted to obtain the target cache limits for each group. The target cache limit for the critical thread group is higher than that for the non-critical thread group.
[0009] In the above cache allocation method, the memory stalls caused by the CPU executing memory access instructions based on the target cache limit within a preset cycle are lower than those caused by using other cache limits during the adjustment process. This ensures a reasonable allocation of cache resources between critical and non-critical thread groups, thereby reducing the overall power consumption of electronic devices.
[0010] In one possible implementation of the first aspect, the first information corresponding to each cycle is the weighted sum of the second information corresponding to the first thread group and the second thread group in the cycle; the weight corresponding to the first thread group is greater than the weight corresponding to the second thread group; wherein, the second information corresponding to each thread group in the first thread group and the second thread group represents the memory stagnation situation when the CPU executes the memory access instructions of the threads in the thread group within the cycle.
[0011] It should be understood that the second information corresponding to each thread group is the memory stagnation status when the CPU executes memory access instructions of the threads in the corresponding thread group within a cycle. The first information is the weighted sum of the second information corresponding to critical thread groups and non-critical thread groups in the cycle, with the weight corresponding to critical thread groups being greater than the weight corresponding to non-critical thread groups.
[0012] In the aforementioned cache allocation method, different weights are assigned to groups of varying importance to accurately represent the importance of each group. Based on this, the optimization objective can be the first piece of information, aiming to minimize memory stagnation during CPU execution of memory access instructions, thereby ensuring the rational allocation of cache resources and reducing the power consumption of electronic devices.
[0013] In another possible implementation of the first aspect, the second information corresponding to each thread group is the ratio of the first duration to the first total duration; the first duration represents the duration of memory stagnation when the CPU executes the memory access instructions of each thread in the thread group within the cycle, and the first total duration represents the total duration occupied by the CPU when executing the instructions of each thread in the thread group within the cycle.
[0014] It should be understood that the first duration is the time during which the CPU pauses memory access instructions from each thread in the thread group within a given period, and the first total duration is the total time occupied by the CPU executing instructions from each thread in the thread group within a given period. The second piece of information is the ratio of the first duration to the first total duration. In the above cache allocation method, by collecting the first and second durations and calculating the second piece of information, an accurate assessment of the CPU's memory access performance for executing thread groups can be achieved, providing an accurate data basis for periodically optimizing and adjusting the cache limits of each group.
[0015] In another possible implementation of the first aspect, based on the first information corresponding to each preset period, the cache upper limit corresponding to the first thread group and the second thread group is periodically optimized and adjusted, including: for each new period, under the condition of satisfying a first condition, optimizing and adjusting the cache upper limit corresponding to the first thread group or the second thread group in the first period. The first condition includes that the first information corresponding to the first period is less than the first information corresponding to the second period, and / or that the difference between the first information corresponding to the first period and the first information corresponding to the second period is less than or equal to a first threshold; the first threshold is a positive number; the first period is the period preceding the new period, and the second period is the period preceding the first period.
[0016] It should be understood that the first condition includes a decrease in the first information and / or an increase in the first information, but the increase does not exceed the first threshold. Under the first condition, the cache limit corresponding to the first cycle is optimized and adjusted.
[0017] In the above cache allocation method, minimizing the first information is the optimization objective. A decrease in the first information indicates that the overall memory stagnation phenomenon has been improved, and the current cache upper limit for each group is more optimal. Further adjustments are made based on the first cycle. If the first information increases but the increase does not exceed the first threshold, it indicates that the overall memory stagnation phenomenon has not been improved. However, considering that a small increase in the first information may stem from other practical factors, further adjustments can be made based on the first cycle. Under the first condition, further adjustments and optimizations based on the first cycle are beneficial for more accurately finding the optimal cache allocation setting that minimizes memory stagnation.
[0018] In another possible implementation of the first aspect, the cache limit corresponding to the first thread group or the second thread group in the first period is optimized and adjusted, including: lowering the cache limit corresponding to the first target thread group in the first period, and / or raising the cache limit corresponding to the second target thread group in the first period. Wherein, the third information corresponding to the first target thread group is less than the third information corresponding to thread groups other than the first target thread group; the third information corresponding to the second target thread group is greater than the third information corresponding to thread groups other than the second target thread group; the third information corresponding to each thread group in the first and second thread groups is the product of the second information corresponding to the thread group and the weight corresponding to the thread group.
[0019] It should be understood that, under the first condition, when optimizing the cache limit for each thread group, the cache limit for the thread group with the smallest third information can be lowered, and / or the cache limit for the thread group with the largest third information can be raised. Here, the third information is the product of the second information corresponding to the thread group and the weight corresponding to the thread group.
[0020] In the above cache allocation method, the third information comprehensively considers the importance of each group and the memory stagnation situation when the CPU executes memory access instructions of threads in the corresponding thread group within a cycle. By lowering the cache share limit of the lowest group in the third information and / or raising the cache share limit of the highest group in the third information, dynamic reallocation of cache resources is achieved, optimizing cache resource utilization and enabling cache resources to more accurately match the actual needs of each group.
[0021] In another possible implementation of the first aspect, based on the first information corresponding to each preset period, the cache upper limits corresponding to the first thread group and the second thread group are periodically optimized and adjusted. This further includes: for each new period, under the condition that a second condition is met, optimizing and adjusting the cache upper limit corresponding to the first thread group or the second thread group in the second period. The second condition includes that the difference between the first information corresponding to the first period and the first information corresponding to the second period is greater than a first threshold; the first threshold is a positive number.
[0022] It should be understood that the second condition includes an increase in the first information level, and the increase being greater than the first threshold. Under the first condition, the cache limit corresponding to the second cycle is optimized and adjusted. The second cycle is the cycle preceding the first cycle.
[0023] In the above cache allocation method, if the first information increases and the increase exceeds the first threshold, it indicates that the overall memory stagnation phenomenon is intensifying. The cache upper limit settings for each group in the first cycle are unreasonable, so the process is directly backtracked to the second cycle. Based on the second cycle, a new round of adjustments is made. Under the second condition, continuing to adjust and optimize based on the backtracking of the second cycle helps to more accurately find the optimal cache allocation setting scheme that minimizes memory stagnation.
[0024] In another possible implementation of the first aspect, the cache limit corresponding to the first thread group or the second thread group in the second period is optimized and adjusted, including: lowering the cache limit corresponding to the third target thread group in the second period, and / or raising the cache limit corresponding to the fourth target thread group in the second period. Wherein, the third information corresponding to the third target thread group is less than the third information corresponding to thread groups other than the third target thread group; the third information corresponding to the fourth target thread group is greater than the third information corresponding to thread groups other than the fourth target thread group; the third information corresponding to each thread group in the first and second thread groups is the product of the second information corresponding to the thread group and the weight corresponding to the thread group.
[0025] It should be understood that, under the second condition, when optimizing and adjusting the cache limit for each thread group, the cache limit for the thread group with the smallest third information can be lowered, and / or the cache limit for the thread group with the largest third information can be raised. This optimizes cache resource utilization, allowing cache resources to more accurately match the actual needs of each group.
[0026] In another possible implementation of the first aspect, based on the first information corresponding to each preset period, the cache upper limits corresponding to the first thread group and the second thread group are periodically optimized and adjusted to obtain the first target cache upper limit corresponding to the first thread group and the second target cache upper limit corresponding to the second thread group. This includes: periodically optimizing and adjusting the cache upper limits corresponding to the first thread group and the second thread group based on the first information corresponding to each preset period; determining the target period during the periodic optimization and adjustment process; determining the cache upper limit corresponding to the first thread group in the target period as the first target cache upper limit; and determining the cache upper limit corresponding to the second thread group in the target period as the second target cache upper limit. Wherein, the first information corresponding to the target period is less than the first information corresponding to each of the N consecutive periods following the target period, and N is a positive integer.
[0027] It should be understood that the target period is determined during the periodic optimization and adjustment process. The first information corresponding to the target period is less than the first information corresponding to each of the N consecutive periods following the target period, where N is a positive integer. The cache limit corresponding to each group of the target period is taken as the target cache limit for each group.
[0028] In the above cache allocation methods, the first information in the target period is the smallest, and after N consecutive rounds of adjustments, the memory stagnation situation has not been substantially improved. Therefore, the cache allocation setting scheme for the target period can be regarded as the optimal cache allocation setting scheme under the current conditions. By setting the cache upper limit for the target period, the optimal allocation of cache resources in critical thread groups and non-critical thread groups is ensured, significantly reducing the overall power consumption of electronic devices.
[0029] In another possible implementation of the first aspect, after allocating a first target cache limit to the first thread group and a second target cache limit to the second thread group, if the second information of any thread group changes beyond the second threshold for M consecutive periods, the cache limit of each thread group is reset; M is an integer greater than or equal to 2.
[0030] It should be understood that after allocating target cache limits to each group, the second information for each group is monitored. If the changes in the second information of any group exceed the second threshold for M consecutive periods, the cache limit for each thread group is reset. In the above cache allocation method, changes in sub-scenes directly affect the memory access behavior of threads, thereby causing fluctuations in the second information. Using the second information as the basis for judging whether a sub-scene has changed has high accuracy. After a sub-scene changes, the electronic device resets the cache limits for each group and uses the method of this application to dynamically readjust the cache limits for each group to ensure rapid adaptation to the new sub-scene's demand for cache resources.
[0031] In another possible implementation of the first aspect, the adjustment granularity of the optimization is a preset unit granularity.
[0032] In the above-described cache allocation method, adjusting at the unit granularity makes the adjustment process faster and more efficient. Since the adjustment granularity is preset, electronic devices do not need to perform complex calculations or judgments during adjustment; they only need to perform simple addition or subtraction operations based on the current needs and the preset unit granularity.
[0033] In another possible implementation of the first aspect, there are multiple first thread groups and multiple second thread groups. The multiple first thread groups include logic thread groups and / or rendering thread groups in the game scene, and the multiple second thread groups include auxiliary thread groups and non-game thread groups in the game scene. The auxiliary thread groups include auxiliary threads related to the game scene other than logic thread groups and rendering thread groups, and the non-game thread groups include threads unrelated to the game scene.
[0034] In the aforementioned cache allocation method, the game logic thread and game rendering thread frequently access and call large amounts of game data during processing, resulting in a very high memory access ratio. If the game logic thread and game rendering thread can improve the cache hit rate of memory access instructions, memory stagnation caused by waiting for memory data can be effectively reduced. In game scenarios, optimizing the game logic thread and game rendering thread can significantly improve game rendering performance and effectively reduce power consumption.
[0035] In a second aspect, this application provides an electronic device comprising: one or more processors and a memory; the memory and the processor being coupled; the memory being used to store computer program code, the computer program code including computer instructions, which, when executed by one or more processors, cause the electronic device to perform the method described in the first aspect and any of its possible design embodiments.
[0036] Thirdly, this application provides a computer-readable storage medium storing instructions that, when executed on an electronic device, cause the electronic device to perform the method described in the first aspect and any of its possible design embodiments.
[0037] Fourthly, this application provides a computer program product that, when run on a computer, causes the computer to perform the method described in the first aspect and any possible design thereof.
[0038] It is understood that the beneficial effects achieved by the electronic device described in the second aspect, the computer-readable storage medium described in the third aspect, and the computer program product described in the fourth aspect can be referred to in the beneficial effects of the first aspect and any of its possible design embodiments, and will not be repeated here. Attached Figure Description
[0039] Figure 1 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application;
[0040] Figure 2 A schematic diagram of the software architecture of an electronic device provided in an embodiment of this application;
[0041] Figure 3 A schematic diagram illustrating the comparison of runtime of a game thread within a single frame, provided as an embodiment of this application;
[0042] Figure 4 This is a schematic diagram of a CPU data reading hierarchy provided in an embodiment of this application;
[0043] Figure 5 This application provides a schematic diagram of the software workflow of an electronic device in a game scenario.
[0044] Figure 6 This is a schematic diagram of a cache allocation method provided in an embodiment of this application;
[0045] Figure 7 This application provides a schematic diagram of a weighted sum calculation of memory stall rate.
[0046] Figure 8 An illustration of an adjustment to a cache allocation setting scheme provided in this application embodiment. Figure 1 ;
[0047] Figure 9 An illustration of an adjustment to a cache allocation setting scheme provided in this application embodiment. Figure 2 ;
[0048] Figure 10 An illustration of an adjustment to a cache allocation setting scheme provided in this application embodiment. Figure 3 ;
[0049] Figure 11 An illustration of an adjustment to a cache allocation setting scheme provided in this application embodiment. Figure 4 ;
[0050] Figure 12 An illustration of an adjustment to a cache allocation setting scheme provided in this application embodiment. Figure 5 ;
[0051] Figure 13 An illustration of an adjustment to a cache allocation setting scheme provided in this application embodiment. Figure 6 ;
[0052] Figure 14 This is a schematic diagram of a sub-scene change provided in an embodiment of this application;
[0053] Figure 15 This is a schematic diagram of a chip system provided in an embodiment of this application. Detailed Implementation
[0054] The technical solutions of the embodiments of this application are described below with reference to the accompanying drawings. In the description of the embodiments of this application, the terminology used in the following embodiments is for the purpose of describing specific embodiments only and is not intended to limit the application. As used in the specification and appended claims of this application, the singular expressions "a," "the," "the," "the," and "this" are intended to also include expressions such as "one or more," unless the context clearly indicates otherwise. It should also be understood that in the following embodiments of this application, "at least one" and "one or more" refer to one or more (including two). The term "and / or" is used to describe the relationship between related objects, indicating that three relationships can exist; for example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship.
[0055] References to "one embodiment" or "some embodiments" in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized. The term "connection" includes direct connections and indirect connections, unless otherwise stated. "First" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated.
[0056] In the embodiments of this application, the words "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplarily" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of the words "exemplarily" or "for example" is intended to present the relevant concepts in a specific manner.
[0057] In some schemes, a support vector machine model is used to classify programs based on the degree of interference the program causes to other programs when it occupies the cache and the program's sensitivity to cache allocation size. Then, a Bayesian optimization algorithm is applied to schedule last-level cache (LLC) resources among program categories, searching for a cache allocation scheme that maximizes system throughput.
[0058] In some solutions, cache behavior prediction models are constructed to predict cache demand. Model input parameters include a cache service state transition matrix and initial cache state parameters, while model output parameters include cache utilization, cache miss rate, and process conflict probability. Based on these model output parameters, a dynamic allocation scheme for shared cache in embedded real-time systems is developed.
[0059] Neither of the above two solutions takes into account the difference in the importance of threads in the current scenario, which may lead to unreasonable allocation of cache resources, thereby increasing the overall power consumption of electronic devices.
[0060] In some solutions, cache space is allocated evenly across threads based on a simple principle of equal distribution. This cache allocation method may result in insufficient cache resources for critical threads, while non-critical threads consume excessive cache space. When critical threads experience performance limitations due to insufficient cache resources, the system has to increase the CPU frequency to maintain the target performance level, leading to a significant increase in power consumption.
[0061] In other approaches, excessive favoritism towards critical threads—for example, consistently allocating too much cache space to critical threads according to a fixed ratio while compressing the cache space of non-critical threads—may lead to a significant increase in the runtime of these non-critical threads. Even with the CPU frequency remaining constant, this can result in increased power consumption due to interference and competition between threads.
[0062] To address this, this application provides a cache allocation method applied to electronic devices. In the current scenario, the electronic device identifies threads and groups them into critical thread groups and non-critical thread groups. Based on the memory stagnation caused by the CPU executing memory access instructions within each cycle, the cache upper limits corresponding to the critical thread groups and non-critical thread groups are periodically optimized and adjusted to obtain the target cache upper limit for each group (i.e., each thread group). This ensures that the memory stagnation caused by the CPU executing memory access instructions based on the target cache upper limit within a preset cycle (i.e., the memory stagnation during CPU execution throughout the entire cycle) is lower than the memory stagnation caused by using other cache upper limits adjusted during the adjustment process. Furthermore, the target cache upper limit corresponding to the critical thread group is greater than the target cache upper limit for the non-critical thread group. This ensures a reasonable allocation of cache resources between the critical thread groups and non-critical thread groups, thereby reducing the overall power consumption of the electronic device.
[0063] Cache limits can also be called cache share limits. Cache share limits refer to the maximum cache resource quota that each group can use in the cache space. For example, within a preset period, if the cache share limit for critical thread group A is 100%, and the cache share limit for non-critical thread group B is 60%, then 40% of the cache is guaranteed to be available only to critical thread group A. Thus, in most cases, critical thread group A is guaranteed to receive more cache resources than non-critical thread group B.
[0064] For example, the aforementioned electronic devices may specifically include mobile phones, tablets, televisions (also known as smart TVs, smart screens, or large-screen devices), laptops, ultra-mobile personal computers (UMPCs), handheld computers, cameras (such as digital cameras), camcorders (such as digital camcorders), netbooks, personal digital assistants (PDAs), wearable electronic devices (e.g., smartwatches, smart bracelets, smart glasses), in-vehicle devices, virtual reality devices, and other electronic devices with cache management functions. This application embodiment does not impose any limitations on this.
[0065] For ease of understanding, a mobile phone is used as an example to illustrate the hardware structure and software system of the electronic device to which the method in the embodiments of this application is applied.
[0066] Figure 1 A schematic diagram of the mobile phone structure is shown.
[0067] The mobile phone may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc.
[0068] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the mobile phone. In other embodiments of this application, the mobile phone may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0069] Processor 110 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.
[0070] Processor 110 can be used to read and execute computer-readable instructions. Optionally, processor 110 may also include a controller, an arithmetic logic unit (ALU), and registers. The controller is primarily responsible for instruction decoding and issuing control signals for the operations corresponding to the instructions. The ALU is primarily responsible for storing register operands and intermediate operation results temporarily stored during instruction execution. Registers are high-speed storage components with limited storage capacity, used to temporarily store instructions, data, and addresses.
[0071] In a specific implementation, the hardware architecture of the processor 110 can be an application-specific integrated circuit (ASIC) architecture, a microprocessor without interlocked piped stages (MIPS) architecture, an ARM (advanced RISC machines) architecture, or a net processor (NP) architecture, etc.
[0072] In some embodiments, the mobile phone can execute a cache allocation method through the processor 110 to ensure the reasonable allocation of cache resources between critical thread groups and non-critical thread groups.
[0073] Internal memory 121 can be used to store computer executable program code, which includes instructions. Processor 110 executes various functional applications and data processing of electronic device 100 by running the instructions stored in internal memory 121. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of electronic device 100 (such as audio data, phonebook, etc.). Furthermore, internal memory 121 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.
[0074] Internal memory 121 includes a cache for storing instructions and data. In some embodiments, the cache is a high-speed buffer memory. The cache can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the cache. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves system efficiency. The method in this embodiment aims to allocate a reasonable cache limit to threads, thereby enabling the rational use of the cache and improving performance.
[0075] The wireless communication function of a mobile phone can be implemented through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor, and baseband processor.
[0076] The mobile communication module 150 can provide wireless communication solutions, including 2G / 3G / 4G / 5G, for use in mobile phones. In some embodiments, at least some functional modules of the mobile communication module 150 can be housed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 can be housed in the same device.
[0077] The wireless communication module 160 can provide solutions for wireless communication applications in mobile phones, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR) technology, and other wireless communication technologies.
[0078] Mobile phones can perform audio functions, such as music playback and recording, through components like the audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.
[0079] Button 190 includes a power button, volume buttons, etc. Button 190 can be a mechanical button or a touch button. The mobile phone can receive button input and generate key signal inputs related to the phone's user settings and function control.
[0080] Motor 191 can generate vibration alerts. Motor 191 can be used for incoming call vibration alerts or for touch vibration feedback.
[0081] Indicator 192 can be an indicator light, used to indicate charging status, power changes, or to indicate messages, missed calls, notifications, etc.
[0082] Camera 193 is an actual physical camera. In some embodiments of this application, the mobile phone may have multiple cameras 193; for example, multiple cameras 193 may include a front-facing camera and a rear-facing camera. There may be one or more front-facing cameras, and there may also be one or more rear-facing cameras.
[0083] The mobile phone can perform shooting functions through the ISP, camera 193, video codec, GPU, display 194, and application processor. For example, an application in the mobile phone can access the camera 193 and control the camera 193 to capture images.
[0084] The mobile phone implements its display function through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0085] Display screen 194 is used to display images, videos, etc.
[0086] The SIM card interface 195 is used to connect the SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to achieve contact and separation with the mobile phone.
[0087] Mobile phone software systems can adopt layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This embodiment of the invention uses a layered architecture software system as an example to illustrate the software structure of a mobile phone.
[0088] Figure 2 This is a software structure block diagram of a mobile phone according to an embodiment of the present invention.
[0089] A layered architecture for electronic devices divides software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, an electronic device may include, for example, an application layer, an application framework layer, a hardware abstraction layer (HAL layer), and a kernel layer. For example, the application layer may include a series of applications.
[0090] The application layer can include a series of application packages.
[0091] like Figure 2 As shown, the application package can include applications such as camera, call, music, video, and games.
[0092] In other embodiments, the application package may also include Figure 2 Modules not shown in the image include applications such as gallery, calendar, map, navigation, WLAN, and Bluetooth.
[0093] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions.
[0094] like Figure 2 As shown, the application framework layer can include game services, content providers, view systems, phone managers, resource managers, etc.
[0095] The game service manages the lifecycle of game applications, including operations such as starting, pausing, resuming, and stopping the game. It coordinates various resources and logic within the game to ensure smooth gameplay.
[0096] Content providers store and retrieve data, making that data accessible to applications. This data may include videos, images, audio, made and received phone calls, browsing history and bookmarks, phone books, etc.
[0097] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.
[0098] The phone manager is used to provide mobile phone communication functions, such as managing call status (including connection and hang-up).
[0099] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.
[0100] In other embodiments, the application framework layer described above may further include Figure 2 Modules not shown in the diagram, such as the window manager and notification manager.
[0101] The window manager is used to manage windowed applications. It can retrieve screen size, determine the presence of a status bar, lock the screen, and capture screenshots, among other things.
[0102] The notification manager allows applications to display notifications in the status bar. These notifications can be used to deliver informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of completed downloads or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting sounds, vibrating electronic devices, and flashing indicator lights.
[0103] For example, the hardware abstraction layer (HAL) can encapsulate drivers in the kernel layer and provide an interface for calling the application framework layer, shielding the implementation details of the lower-level hardware.
[0104] like Figure 2 As shown, the hardware abstraction layer of the electronic device includes at least an identification and grouping module and a partitioning and calculation module. The identification and grouping module is used to identify each thread and group them according to the importance of each thread to the target scene. The partitioning and calculation module is used to calculate the upper limit of the cache share of each group in the cache for the next cycle.
[0105] The kernel layer is the layer between hardware and software. It contains at least a data collector and a cache share setting module. The data collector gathers target information for each group. This target information represents the memory stalls that occur when the CPU executes memory access instructions during the execution of threads in the corresponding group within a given period. The cache share setting module sets the upper limit of the cache share for each group based on the calculation results from the partitioning calculation module.
[0106] It should be understood that the cache allocation method in this application embodiment can be applied to various stable and continuous scenarios, including but not limited to game scenarios, video playback scenarios, music playback scenarios, and recording scenarios. The following description uses a game scenario as an example to illustrate the cache allocation method. Cache allocation methods for other scenarios are similar to those for the game scenario and will not be elaborated upon further.
[0107] In a game scenario, game frame rendering refers to the process by which the game engine translates the virtual game world into a visible image on the electronic device screen in real time while the device is running a game. For each frame, the game logic thread first calculates the content that needs to be displayed. Based on the calculation results from the game logic thread, the game rendering thread executes the graphics rendering task to generate the final image. In addition, other auxiliary threads may exist to coordinate the rendering of game frames.
[0108] Figure 3 This shows a comparison of the runtime of the game thread within a single frame. The first function refers to the function called during the execution of the game rendering thread; there may be multiple first functions. Figure 3 The text uses only different patterns to illustrate two different first functions for simple explanation. Figure 3 It is evident that the runtime of the game logic thread and the game rendering thread is significantly longer than that of the game auxiliary thread. In other words, the total rendering time of a single game frame largely depends on the runtime of the game logic thread and the game rendering thread.
[0109] Through in-depth research and analysis, the inventors of this application discovered that because the game logic thread and the game rendering thread frequently access and call large amounts of game data during processing, they have an extremely high memory access ratio; that is, these two threads account for a very high percentage of memory access during runtime. The high memory access demands of the game logic thread and the game rendering thread make memory access latency a significant factor affecting the runtime of these two threads.
[0110] For ease of understanding, please refer to Figure 4The illustration shows the CPU's hierarchical structure when requesting to read data. When the CPU initiates a data read request, it first searches the L1 cache. If the required data is not found in the L1 cache, it continues searching the L2 cache. If the L2 cache also fails to find the target data, it further searches the L3 cache. If the data is still not found in the L3 cache (i.e., a cache miss occurs), the required data is finally read from main memory.
[0111] Because the game logic thread and game rendering thread need to access a large amount of game data, they frequently initiate data read requests. If these requests frequently read data from main memory instead of hitting the faster cache, it leads to a phenomenon known as "cache stall." Cache stall increases data access latency, thus affecting the runtime of these two threads. Therefore, if the game logic thread and game rendering thread can improve the cache hit rate of memory access instructions, they can effectively reduce memory stall caused by waiting for memory data, thereby increasing the number of instructions per cycle (ipc) and ultimately shortening the thread execution time.
[0112] Based on this, optimizing the game logic thread and the game rendering thread in game scenarios can significantly improve game rendering performance and effectively reduce power consumption.
[0113] For example, in a game scenario, electronic devices can divide threads into four groups based on their importance to the game environment: game logic threads, game rendering threads, game support threads, and non-game threads. The game logic threads and game rendering threads are considered critical threads, while the game support threads and non-game threads are considered non-critical. The grouping of game threads is not limited to four groups; the specific number of groups can be flexibly adjusted according to actual needs and is not limited here.
[0114] In some embodiments, after the game starts, the electronic device identifies and groups the threads. Specifically, the electronic device can identify game threads and non-game threads based on their process IDs. Game threads include game logic threads, game rendering threads, and game auxiliary threads. Within the game threads, the electronic device further identifies game logic threads and marks them as game logic thread groups; identifies game rendering threads and marks them as game rendering thread groups; and identifies game auxiliary threads and marks them as game auxiliary thread groups. The electronic device marks non-game threads as non-game thread groups.
[0115] Based on the memory stalls that occur when the CPU executes memory access instructions within each cycle, electronic devices periodically optimize and adjust the cache limits for the game logic thread group, game rendering thread group, game auxiliary thread group, and non-game thread group to obtain the target cache limits for these four groups. This ensures that the memory stalls caused by the CPU executing memory access instructions based on the target cache limits within a preset cycle are lower than the memory stalls caused by using other cache limits during the adjustment process. Please refer to [link / reference]. Figure 5 Using a game scenario as an example, the workflow of the mobile software is illustrated, specifically including the following steps:
[0116] (1) Launch the game application at the application layer.
[0117] (2) The game application transmits state information to the game service in the application framework layer.
[0118] This status information is used to indicate that the game application has started.
[0119] (3) The game service transmits state information to the game scheduler of the hardware abstraction layer.
[0120] In this embodiment, the group identification module and the partitioning calculation module are set in the game scheduler. In other embodiments, they may also be set in other modules, which is not limited here.
[0121] (4) After the game scheduler’s identification and grouping module groups each thread, it sends the grouping results to the kernel layer’s data collector.
[0122] For example, the identification and grouping module can begin identifying each thread after the game application starts, and group each thread according to its importance to the game scene.
[0123] (5) The data collector collects the memory stagnation situation (i.e. target information) generated when the CPU executes memory access instructions during the running of each group thread in this cycle.
[0124] (6) The data collector sends the target information of each group to the game scheduler.
[0125] (7) The game scheduler's partitioning calculation module sends the cache allocation setting scheme for the next cycle to the kernel layer's cache share setting module.
[0126] Specifically, the partitioning calculation module can calculate the cache allocation setting scheme for the next cycle (that is, the upper limit of the cache share in the cache corresponding to each group) based on the target information of each group, and send the cache allocation setting scheme for the next cycle to the cache share setting module in the kernel layer. For example, the cache allocation setting scheme can be represented as configuration information such as <group number, upper limit of cache share>.
[0127] Furthermore, the cache share setting module can set the upper limit of the corresponding cache share in the cache for each group based on the cache allocation setting scheme.
[0128] In some embodiments, the cache share setting module can set the upper limit of the cache share for each group in the last level cache (LLC) according to the above cache allocation setting scheme.
[0129] It should be understood that, in cases such as Figure 4 In the illustrated embodiment, the Level 3 cache (L3 Cache) is an LLC. In other embodiments, the cache consists of only a Level 1 cache (L1 Cache) and a Level 2 cache (L2 Cache). In this case, the Level 2 cache (L2 Cache) is the last-level cache (LLC), and the LLC in other embodiments can be understood similarly.
[0130] For ease of understanding, the cache allocation method of this application is described in detail in the following embodiments.
[0131] In some embodiments, such as Figure 6 As shown, the electronic device determines a first thread group and a second thread group for the current scene, with the first thread group being more important than the second thread group in the current scene. The first thread group is the critical thread group, and the second thread group is the non-critical thread group.
[0132] Based on the first information corresponding to each preset cycle, the cache limits corresponding to the first thread group and the second thread group are periodically optimized and adjusted. This first information is used to characterize the memory stagnation that occurs when the CPU executes memory access instructions within the cycle. Figure 6 In this context, the cache limit of the first thread group within period n is denoted as a. n The cache limit of the second thread group within period n is denoted as b. n For example, the cache limit for the first thread group within period 1 is denoted as a1, and the cache limit for the second thread group within period 1 is denoted as b1. After optimization and adjustment of the cache limits for the first and second thread groups over multiple periods, the first target cache limit a corresponding to the first thread group is obtained. n+1 The second target cache limit b corresponding to the second thread group n+1 Thus, the memory stalls that occur when the CPU executes memory access instructions based on the target cache limit within a preset cycle are lower than the memory stalls that occur when using other cache limits during the adjustment process.
[0133] In some embodiments, the optimization objective can be the first information, which aims to minimize the first information, that is, to minimize the overall memory stagnation when the CPU executes memory access instructions of all threads, thereby ensuring the reasonable allocation of cache resources and reducing the power consumption of electronic devices.
[0134] Specifically, thread grouping depends on the importance of each thread to the target scenario. The grouping result includes at least one critical thread group and one non-critical thread group.
[0135] For example, in a game scene, the game logic thread and the game rendering thread are critical threads, while the game auxiliary threads and non-game threads are non-critical threads. Thus, in a game scene, threads can be divided into four groups: the game logic thread group, the game rendering thread group, the game auxiliary thread group, and the non-game thread group.
[0136] For example, in video playback scenarios, threads differ in their importance. The video playback rendering thread is considered a critical thread, while video playback auxiliary threads and non-video playback threads are considered non-critical threads. Thus, threads can be divided into three groups: the video playback rendering thread group, the video playback auxiliary thread group, and the non-video playback thread group. Alternatively, threads can be divided into two groups: the video playback rendering thread group and the non-critical thread group, where the non-critical thread group includes all threads except the video playback rendering thread.
[0137] The electronic device can calculate the second information corresponding to each group based on the target information of each group. For example, the second information can be the ratio of the first duration to the first total duration. Furthermore, for each cycle, the electronic device can perform a weighted summation of the second information corresponding to each group in that cycle to obtain the first information corresponding to that cycle.
[0138] Specifically, within a preset period, the electronic device collects target information for each group. The target information may include a first duration and a first total duration. The first duration represents the duration of memory stagnation caused by the CPU executing memory access instructions of each thread in the group within the period, and the first total duration represents the total time occupied by the CPU executing instructions of each thread in the group within the period.
[0139] For example, the first duration can be the memory stall cycle for each group, and the first total duration can be the total number of instruction cycles.
[0140] A memory stall cycle refers to the number of clock cycles during which the CPU remains idle while executing instructions, waiting for memory operations (such as data reads). Total instruction cycles refer to the total number of clock cycles occupied by all instructions executed by the CPU within a preset cycle. Here, "clock cycle" refers to the unit of time required for the CPU to execute a basic operation. For example, for each preset cycle, target information for each group can be collected through a Performance Monitoring Unit (PMU). The PMU is essentially a counter that can be used to collect memory stall cycles and total cycles for each group.
[0141] In some embodiments of this application, within each cycle, memory stall cycle and total cycles are collected by group to obtain the memory stall cycle and total cycles for each group. Specifically, when a thread switches, it is first checked whether the thread before and after the switch (e.g., thread A switches to thread B) belongs to the same group. If they belong to the same group, the PMU continues to count. If thread B and thread A belong to different groups, the electronic device reads the PMU data, that is, records the memory stall cycle and total cycles of the group to which thread A belongs in the previous time slice. After that, the PMU is reset to zero, and counting begins for the group to which thread B belongs.
[0142] Within a preset period, for any group, the memory stall cycles from different time slices are summed to obtain the memory stall cycle for that group, and the total cycles from different time slices are summed to obtain the total cycles for that group. This yields the memory stall cycle and total cycles for each group. At the end of a preset period, the data collected in that period is read and cleared, and counting restarts for the next preset period.
[0143] In some embodiments, the electronic device can assign different weights to different groups to accurately characterize the importance of each group. Furthermore, for each cycle, the electronic device can perform a weighted summation of the second information corresponding to each group in that cycle to obtain the first information corresponding to that cycle. The second information corresponding to each group characterizes the memory stagnation status of the CPU when executing memory access instructions of threads in that thread group during that cycle. For example, the second information can be the ratio of a first duration to a first total duration.
[0144] Specifically, the electronic device can multiply the second information corresponding to each thread group in that period with the weight of the thread group to obtain the third information of each thread group. Then, the electronic device can add the third information of each thread group to obtain the first information corresponding to that period.
[0145] In some embodiments, for any given cycle and for any given thread group, the first duration is the memory stall cycle, the first total duration is the total instruction cycles, the second information is the ratio of the memory stall cycle to the total cycles (i.e., the memory stall rate), and the third information is the product of the memory stall rate and the weight of the corresponding thread group (i.e., the weighted memory stall rate). The electronic device sums the third information for each group to obtain the first information (i.e., the weighted sum of memory stall rates).
[0146] Based on this, the optimization objective can be a weighted sum of memory stagnation rates, aiming to minimize the total weighted memory stagnation rate of each group, thereby ensuring the reasonable allocation of cache resources and reducing the power consumption of electronic devices.
[0147] For the same preset period, the weighted memory stall rates of each group are summed to obtain the weighted sum of memory stall rates for the corresponding preset period. For any group, the weighted memory stall rate can be the product of the memory stall rate and the corresponding group weight coefficient; the memory stall rate can be calculated based on the memory stall cycle and total cycles. For example, the memory stall rate can be memory stall cycle / total cycles.
[0148] For example, such as Figure 7 As shown, for the target scenario, threads are divided into four groups, group 1 to group 4. Weight coefficients are assigned based on the importance of threads within each group, with weight coefficients a, b, c, and d for groups 1 to 4, respectively. Within a preset period, for any group, the electronic device calculates the memory stall rate based on the collected memory stall cycle and total cycles. This memory stall rate is multiplied by the corresponding weight coefficient to obtain the weighted memory stall rate for that group. The weighted memory stall rates of all groups are then summed to obtain the weighted sum of memory stall rates.
[0149] Electronic devices can dynamically adjust cache allocation settings based on a preset period to ensure reasonable and timely optimization of cache resources.
[0150] For the current preset period, the electronic device calculates the memory stall rate, weighted memory stall rate, and weighted sum of memory stall rates for each group. For the next preset period, the electronic device can adjust the cache allocation settings based on the weighted memory stall rate and weighted sum of memory stall rates for each group. For example, the cache share cap for the group with the lowest weighted stall rate can be lowered, and / or the cache share cap for the group with the highest weighted stall rate can be raised.
[0151] By lowering the cache share cap for the group with the lowest weighted memory stagnation rate and / or raising the cache share cap for the group with the highest weighted memory stagnation rate, dynamic reallocation of cache resources is achieved, optimizing cache resource utilization and enabling cache resources to more accurately match the actual needs of each group.
[0152] Electronic devices can adjust the cache share limit by lowering or raising it each time, based on a pre-defined unit granularity. The unit granularity can be set according to needs. For example, if more granular cache resource allocation is desired, the unit granularity can be set to a smaller value, such as 2% or 5%; if simplicity of adjustment and reduced adjustment frequency are desired, the unit granularity can be set to a larger value, such as 10% or 15%.
[0153] Adjusting using unit granularity makes the adjustment process faster and more efficient. Since the adjustment granularity is preset, electronic devices do not need to perform complex calculations or judgments during adjustment; they only need to perform simple additions or subtractions based on the current needs and the preset unit granularity.
[0154] Figure 8 and Figure 9 The diagram illustrates how to adjust two cache allocation settings.
[0155] Figure 8 and Figure 9 In this embodiment, within period 1, the maximum cache share for groups 1 to 4 are A, B, C, and D, respectively. For period 1, the electronic device calculates the weighted memory stagnation rate for groups 1 to 4. Assuming that group 1 has the highest weighted memory stagnation rate and group 4 has the lowest weighted memory stagnation rate, this indicates that group 1 has fewer cache resources compared to other groups, or group 4 has more cache resources compared to other groups.
[0156] In some embodiments, such as Figure 8 As shown, the electronic device can reduce the cache share cap for Group 4 in the next preset cycle (i.e., cycle 2). For example, the adjusted cache share cap for Group 4 is D-10%.
[0157] In other embodiments, such as Figure 9 As shown, electronic devices can increase the cache share cap for Group 1 in cycle 2. For example, the adjusted cache share cap for Group 1 is A+5%.
[0158] The electronic device calculates the difference in the weighted sum of memory stall rates between two preset periods, and adjusts the upper limit of cache share for each group in the next preset period based on this difference. Using the weighted sum of memory stall rates as the optimization target, and comprehensively considering the memory access efficiency of each group, the allocation of cache resources is dynamically optimized through monitoring over multiple consecutive periods, significantly improving the overall memory stall phenomenon.
[0159] If the weighted sum of memory stall rates decreases, it indicates that the overall memory stalling situation has improved, the current cache allocation settings are better, and the electronic device can continue to adjust based on the current cache allocation settings. That is, lower the cache share cap for the group with the lowest current weighted memory stall rate, and / or raise the cache share cap for the group with the highest current weighted memory stall rate. If the weighted sum of memory stall rates increases, but the increase does not exceed a preset first threshold, it indicates that the overall memory stalling situation has not improved. In some embodiments, considering that a small increase in memory stall rate may originate from instantaneous load changes, to avoid over-adjustment causing system instability, the electronic device continues to adjust based on the current cache allocation settings. That is, based on the current cache allocation settings, lower the cache share cap for the group with the lowest current weighted memory stall rate, and / or raise the cache share cap for the group with the highest current weighted memory stall rate.
[0160] In other embodiments, if the weighted sum of the memory stall rate increases, but the increase does not exceed a preset first threshold, the electronic device may also choose to revert to the previous cache allocation setting scheme and make a new round of adjustments based on it.
[0161] If the weighted sum of memory stall rates increases and the increase is greater than the preset first threshold, it indicates that the overall memory stall phenomenon has intensified, the current cache allocation settings are significantly unreasonable, and the electronic device will directly revert to the previous cache allocation settings and make a new round of adjustments based on that.
[0162] In the event that an electronic device reverts to the previous cache allocation setting and makes a new round of adjustments to the previous cache allocation setting, in order to avoid reverting to the current unreasonable cache allocation setting during the adjustment process, the electronic device can change the adjustment method in the new round of adjustments, such as changing the adjustment granularity or the adjustment object.
[0163] In some embodiments, a target period is determined during the periodic optimization and adjustment process. The target period is the period with the smallest weighted sum of historical memory stall rates. After the target period, the weighted sum of memory stall rates for each of the N consecutive periods (N is a preset positive integer) is not less than the weighted sum of memory stall rates corresponding to the target period. This indicates that despite the gradual adjustment over N periods, the memory stall situation has not been substantially improved.
[0164] In this scenario, given that the weighted sum of memory stall rates corresponding to the cache allocation settings for the target period is at its lowest level, this scheme is considered the optimal configuration under the current operating conditions. Therefore, the electronic device adopts the cache allocation settings corresponding to the target period, that is, the target cache limit for each group is set to the cache limit for each group in the target period, in order to reduce the overall power consumption of the electronic device.
[0165] After allocating target cache limits to each group, if the memory stall rate of any group changes by more than a preset second threshold for M consecutive periods (M being a preset integer greater than or equal to 2), then the sub-scenario is considered to have changed. It should be understood that the periods cannot be too short; it is necessary to ensure that the memory stall rate of each group remains basically stable while the sub-scenario remains unchanged.
[0166] In this context, a sub-scene refers to a smaller scene within the target scene, which can be further subdivided based on specific activity content, environmental characteristics, or behavioral patterns. For example, a game scene might include sub-scenes such as free movement, resource collection, and combat. When a sub-scene changes, the electronic device can reset the cache limit for each group and then dynamically readjust the cache limit for each group using the method described in the above embodiment. The memory stall rate for each group is an important indicator reflecting thread memory access efficiency. When a sub-scene changes, it directly affects the thread's memory access behavior, thus causing fluctuations in the memory stall rate. Using the memory stall rate as the criterion for determining whether a sub-scene has changed has high accuracy. After a sub-scene changes, resetting the cache limit for each group and dynamically readjusting the cache limit for each group using the method of this application ensures rapid adaptation to the new sub-scene's demand for cache resources.
[0167] For ease of understanding, the following embodiments describe how the upper limit of the cache share in this application is adjusted.
[0168] In some embodiments, such as Figures 10 to 12 As shown, within period 1, the maximum cache share of groups 1 to 4 is A, B, C and D respectively. In the next preset period (i.e. period 2), the maximum cache share of groups 1 to 4 is A+5%, B, C and D respectively.
[0169] In some embodiments, such as Figure 10 As shown, the difference in the weighted sum of memory stall rates between two consecutive periods is statistically analyzed. If the weighted sum of memory stall rates in period 1 (1) is greater than the weighted sum of memory stall rates in period 2 (2), meaning the weighted sum of memory stall rates decreases, then the cache allocation settings in period 2 are better. For the next period (i.e., period 3), the electronic device continues to adjust the cache share limits for each group based on the cache allocation settings in period 2. For example, in period 3, the cache share limits for groups 1 to 4 are A+5%+5% (i.e., A+10%), B, C, and D, respectively.
[0170] In some embodiments, such as Figure 11 As shown, the difference in the weighted sum of memory stagnation rates between two consecutive cycles is statistically analyzed. If the weighted sum of memory stagnation rates in cycle 1 (1) is less than the weighted sum of memory stagnation rates in cycle 2 (2), meaning the weighted sum of memory stagnation rates increases and the increase exceeds the first threshold, it indicates that the current cache allocation setting scheme is unreasonable. Therefore, for the next preset cycle (i.e., preset cycle 3), the electronic device uses the preset cache allocation setting scheme of cycle 1. That is, in cycle 3, the cache limits for groups 1 to 4 are reset to A, B, C, and D respectively.
[0171] like Figure 11 As shown, for Period 4, to avoid reverting to the unreasonable cache allocation settings of Period 2 during the adjustment process, the electronic device can change the adjustment method in the new round of adjustments. For example, the electronic device can change the adjustment method of "increasing the cache share limit of the group with the highest weighted stagnation rate" to "decreasing the cache share limit of the group with the lowest weighted stagnation rate". Thus, in Period 4, the cache limits of Groups 1 to 4 are A, B, C, and D-5%, respectively. Another example is that the electronic device can reduce the adjustment granularity; for instance, the adjustment method of "increasing the cache share limit of the group with the highest weighted stagnation rate by 5%" can be changed to "increasing the cache share limit of the group with the highest weighted stagnation rate by 3%".
[0172] In some embodiments, such as Figure 12As shown, the electronic device can statistically analyze the difference in the weighted sum of memory stagnation rates between two consecutive preset periods. If the weighted sum of memory stagnation rates 1 in preset period 1 is less than or equal to the weighted sum of memory stagnation rates 2 in preset period 2, and the increase is less than or equal to a first threshold (i.e., the weighted sum of memory stagnation rates increases, and the increase is less than or equal to the first threshold), then for the next preset period (i.e., preset period 3), the electronic device can continue to adjust the cache allocation settings based on the settings in period 2. For example, in period 3, the maximum cache share for groups 1 to 4 are A+5%+5% (i.e., A+10%), B, C, and D, respectively. In some embodiments, a target period is determined during the periodic optimization and adjustment process, and the maximum cache share corresponding to each group in the target period is determined as the target maximum cache share for each group.
[0173] like Figure 13 As shown, in period Y, the maximum cache share for groups 1 to 4 are A, B, C, and D, respectively; in period Y+1, the maximum cache share for groups 1 to 4 are A+5%, B, C, and D, respectively. Compared to all periods before period Y, period Y has the smallest weighted sum of memory stall rates. Assuming the weighted sum of memory stall rates for period Y is y, and continuously recording the weighted sum of memory stall rates for each subsequent period, the weighted sum of memory stall rates for period Y+1 is y+1, and so on, until the weighted sum of memory stall rates for period Y+N (where N is a preset positive integer) is y+n. If y+1 to y+n are all not less than y, then period Y is the target period. In periods Y+N+1 and beyond, the electronic device sets the target cache limit for each group to the corresponding cache limit for each group in period Y. That is, the maximum cache share for groups 1 to 4 are A, B, C, and D, respectively.
[0174] After allocating target cache limits to each group, electronic devices can record the memory stall rate for each group. For example... Figure 14 As shown, assuming that in period 8, the electronic devices allocate the target cache limit to groups 1 to 4, then for subsequent periods, the memory stall rate for groups 1 to 4 will be continuously observed. Within period a, the memory stall rate for group b is denoted as memory stall rate ab. For example, within period 8, the memory stall rate for group 1 is denoted as memory stall rate 8-1.
[0175] If the memory stall rate of any group exceeds the second threshold for multiple consecutive M periods, for example, such as Figure 14As shown, if the difference in memory stall rate for group 4 exceeds the second threshold for three consecutive cycles, the sub-scenario is considered to have changed. The cache share limit for each group during the electronic device reset cycle 12 is dynamically adjusted according to the method of this application.
[0176] This application also provides an electronic device, which may include: one or more processors and a memory; the memory and the processor are coupled; the memory is used to store computer program code, the computer program code including computer instructions, which, when executed by one or more processors, enable the electronic device to perform the various functions or steps performed in the above method embodiments.
[0177] This application also provides a chip system, such as... Figure 15 As shown, the chip system 1500 includes at least one processor 1501 and at least one interface circuit 1502. The processor 1501 and the interface circuit 1502 are interconnected via lines. For example, the interface circuit 1502 can be used to receive signals from other devices (e.g., the memory of an electronic device). As another example, the interface circuit 1502 can be used to send signals to other devices (e.g., the processor 1501). Exemplarily, the interface circuit 1502 can read instructions stored in memory and send those instructions to the processor 1501. When the instructions are executed by the processor 1501, the electronic device can perform the steps in the above embodiments. Of course, the chip system may also include other discrete devices, which are not specifically limited in this application embodiment.
[0178] This application also provides a computer storage medium that includes computer instructions. When the computer instructions are executed on the electronic device, the electronic device performs the steps described in the method embodiments.
[0179] This application also provides a computer program product that, when run on a computer, causes the computer to perform the steps described in the above method embodiments.
[0180] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0181] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0182] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0183] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0184] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially or in other words, the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0185] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A cache allocation method, characterized in that, Applied to electronic devices, the method includes: Determine a first thread group and a second thread group for the current scenario; the importance of the first thread in the first thread group in the current scenario is higher than the importance of the second thread in the second thread group in the current scenario. Based on the first information corresponding to each preset period, a backtracking mechanism is used to periodically optimize and adjust the cache limits corresponding to the first thread group and the second thread group, respectively, to obtain the first target cache limit corresponding to the first thread group and the second target cache limit corresponding to the second thread group. The backtracking mechanism includes, for a new period, optimizing and adjusting based on the cache limit of the first period if a first condition is met, and optimizing and adjusting based on the cache limit of the second period if a second condition is met. The first period is the previous period of the new period, and the second period is the previous period of the first period. The first information corresponding to each cycle is used to characterize the memory stagnation situation generated when the CPU executes a memory access instruction within the cycle, and the upper limit of the first target cache is higher than the upper limit of the second target cache. The first condition includes that the first information corresponding to the first period is less than the first information corresponding to the second period, and / or that the difference between the first information corresponding to the first period and the first information corresponding to the second period is less than or equal to a first threshold; the first threshold is a positive number. The second condition includes that the difference between the first information corresponding to the first period and the first information corresponding to the second period is greater than the first threshold.
2. The cache allocation method according to claim 1, characterized in that, The first information corresponding to each cycle is a weighted sum of the second information corresponding to the first thread group and the second thread group respectively in the cycle; the weight corresponding to the first thread group is greater than the weight corresponding to the second thread group. The second information corresponding to each thread group in the first thread group and the second thread group represents the memory stagnation situation when the CPU executes the memory access instructions of the threads in the thread group within the cycle.
3. The cache allocation method according to claim 2, characterized in that, The second information corresponding to each thread group is the ratio of the first duration to the first total duration; the first duration represents the duration of memory stagnation when the CPU executes the memory access instructions of each thread in the thread group within the cycle, and the first total duration represents the total duration occupied by the CPU when executing the instructions of each thread in the thread group within the cycle.
4. The cache allocation method according to claim 2 or 3, characterized in that, Based on the first information corresponding to each preset period, a backtracking mechanism is used to periodically optimize and adjust the cache limits corresponding to the first thread group and the second thread group, including: For each new cycle, if the first condition is met, the cache limit of the first thread group or the second thread group in the first cycle is optimized and adjusted.
5. The cache allocation method according to claim 4, characterized in that, The optimization adjustment of the cache limit corresponding to the first thread group or the second thread group in the first cycle includes: Lower the cache limit for the first target thread group in the first period, and / or raise the cache limit for the second target thread group in the first period; Wherein, the third information corresponding to the first target thread group is less than the third information corresponding to the thread groups other than the first target thread group; the third information corresponding to the second target thread group is greater than the third information corresponding to the thread groups other than the second target thread group; the third information corresponding to each thread group in the first thread group and the second thread group is the product of the second information corresponding to the thread group and the weight corresponding to the thread group.
6. The cache allocation method according to claim 2 or 3, characterized in that, The method of periodically optimizing and adjusting the cache limits corresponding to the first thread group and the second thread group based on the first information corresponding to each preset period, using a backtracking mechanism, further includes: For each new cycle, if the second condition is met, the cache limit of the first thread group or the second thread group in the second cycle is optimized and adjusted.
7. The cache allocation method according to claim 6, characterized in that, The optimization adjustment of the cache limit for the first thread group or the second thread group in the second cycle includes: Lower the cache limit for the third target thread group in the second cycle, and / or raise the cache limit for the fourth target thread group in the second cycle; Wherein, the third information corresponding to the third target thread group is less than the third information corresponding to the thread groups other than the third target thread group; the third information corresponding to the fourth target thread group is greater than the third information corresponding to the thread groups other than the fourth target thread group; the third information corresponding to each thread group in the first thread group and the second thread group is the product of the second information corresponding to the thread group and the weight corresponding to the thread group.
8. The cache allocation method according to claim 1, characterized in that, The step of periodically optimizing and adjusting the cache limits corresponding to the first thread group and the second thread group based on the first information corresponding to each preset period, to obtain the first target cache limit corresponding to the first thread group and the second target cache limit corresponding to the second thread group, includes: Based on the first information corresponding to each preset period, the cache limit corresponding to the first thread group and the second thread group is periodically optimized and adjusted. During the periodic optimization and adjustment process, the target period is determined, and the cache limit corresponding to the first thread group in the target period is determined as the first target cache limit, and the cache limit corresponding to the second thread group in the target period is determined as the second target cache limit. Wherein, the first information corresponding to the target period is less than the first information corresponding to each of the N consecutive periods following the target period, and N is a positive integer.
9. The cache allocation method according to claim 1, characterized in that, The method further includes: After allocating the first target cache limit to the first thread group and the second target cache limit to the second thread group, if the second information of any thread group changes beyond the second threshold for M consecutive periods, the cache limit of each thread group is reset; M is an integer greater than or equal to 2.
10. The cache allocation method according to claim 1, characterized in that, The adjustment granularity for optimization is the preset unit granularity.
11. The cache allocation method according to claim 1, characterized in that, There are multiple first thread groups and multiple second thread groups. The multiple first thread groups include logic thread groups and / or rendering thread groups in the game scene. The multiple second thread groups include auxiliary thread groups and non-game thread groups in the game scene. The auxiliary thread groups include auxiliary threads related to the game scene other than the logic thread groups and the rendering thread groups. The non-game thread groups include threads unrelated to the game scene.
12. An electronic device, characterized in that, The electronic device includes: one or more processors and a memory; the memory and the processor are coupled; the memory is used to store computer program code, the computer program code including computer instructions, which, when executed by the one or more processors, cause the electronic device to perform the method as described in any one of claims 1-11.
13. A computer-readable storage medium storing instructions, characterized in that, When the instructions are executed on an electronic device, the electronic device causes the electronic device to perform the method as described in any one of claims 1-11.
14. A computer program product, characterized in that, When the computer program product is run on a computer, it causes the computer to perform the method as described in any one of claims 1-11.
Citation Information
Patent Citations
Thread scheduling method, equipment and related device
CN116225632A