Load balancing method and device

By calculating the cache utilization of multi-core processors and selecting low-utilization tasks for migration, the problem of high resource overhead during load balancing among multi-core processors is solved, achieving more efficient load balancing.

CN120973523APending Publication Date: 2025-11-18VIVO MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511072120.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing technologies incur significant resource overhead when performing load balancing across multi-core processors, resulting in high costs for load balancing.

Method used

By acquiring historical processing data of candidate tasks in the source processor core, including the number of executed instructions and the number of cache misses, the cache utilization rate is calculated, and tasks with utilization rates below a threshold are selected for migration to reduce resource overhead.

Benefits of technology

This reduces the amount of cached data reloaded during task migration, lowers the resource overhead of load balancing, and improves task processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120973523A_ABST
    Figure CN120973523A_ABST
Patent Text Reader

Abstract

The invention discloses a load balancing method and device, and belongs to the technical field of communication. The load balancing method is applied to the electronic equipment, the electronic equipment comprises a plurality of processor cores, and the load balancing method comprises the steps that in response to a received load balancing instruction, historical processing data of all candidate tasks in a source processor core are obtained, the load balancing instruction comprises the source processor core and a target processor core, and the target processor core comprises a target task; the historical processing data comprises the number of execution instructions of each candidate task and the number of cache misses; according to the number of execution instructions and the number of cache miss times of each candidate task, determining the utilization rate of each candidate task to the cache of the source processor core; determining a target task according to the utilization rate of each candidate task to the cache, the utilization rate of the target task being less than or equal to a utilization rate threshold; and migrating the target task from the source processor core to the target processor core.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of communication, and particularly relates to a load balancing method and device. BACKGROUND

[0002] With the development of processor technology, currently, the processor of an electronic device is mainly a multi-core processor, and an operating system maintains a task queue for each processor core. In actual application, the electronic device has a problem of unbalanced load, which causes some processor cores to be busy and some processor cores to be idle. In order to improve the utilization rate of resources and the processing efficiency of tasks, the operating system can migrate part of tasks from the busy processor cores to the relatively idle processor cores to achieve load balancing of the processor cores.

[0003] The currently used load balancing method is to select a task from the end of the task queue of the busy processor core and migrate the task to the idle processor core. However, because the resource overhead caused by the migration of each task between different processor cores is different, the cost of load balancing is large. SUMMARY

[0004] The embodiments of the present application aim to provide a load balancing method and device, which can effectively solve the problem of large resource overhead caused by related technologies in load balancing of different processor cores.

[0005] In a first aspect, the embodiments of the present application provide a load balancing method applied to an electronic device, the electronic device comprising a plurality of processor cores, and the load balancing method comprising:

[0006] In response to a received load balancing instruction, obtaining historical processing data of each candidate task in a source processor core, the load balancing instruction comprising the source processor core and a target processor core, the source processor core and the target processor core being different processor cores in the plurality of processor cores, and the historical processing data comprising an execution instruction number and a cache miss number of each candidate task;

[0007] According to the execution instruction number and the cache miss number of each candidate task, determining a utilization rate of a cache of the source processor core by each candidate task;

[0008] According to the utilization rate of the cache by each candidate task, determining a target task, the utilization rate of the target task being less than or equal to a utilization rate threshold;

[0009] Migrating the target task from the source processor core to the target processor core.

[0010] In a second aspect, the embodiments of the present application provide a load balancing device applied to an electronic device, the electronic device comprising a plurality of processor cores, and the load balancing device comprising:

[0011] The acquisition module is configured to acquire historical processing data of each candidate task in the source processor core in response to the received load balancing instruction, the load balancing instruction comprising the source processor core and the target processor core, the source processor core and the target processor core being different processor cores in the plurality of processor cores, the historical processing data comprising the number of execution instructions and the number of cache misses of each candidate task;

[0012] The determination module is configured to determine, according to the number of execution instructions and the number of cache misses of each candidate task, a utilization rate of a cache of the source processor core by each candidate task, and determine, according to the utilization rate of the cache by each candidate task, a target task, the utilization rate of the target task being less than or equal to a utilization rate threshold.

[0013] The migration module is configured to migrate the target task from the source processor core to the target processor core.

[0014] In a third aspect, an electronic device is provided, which includes a processor and a memory. The memory stores programs or instructions executable on the processor. When the programs or instructions are executed by the processor, the steps of the method according to the first aspect are implemented.

[0015] In a fourth aspect, a readable storage medium is provided. The readable storage medium stores programs or instructions. When the programs or instructions are executed by a processor, the steps of the method according to the first aspect are implemented.

[0016] In a fifth aspect, a chip is provided. The chip includes a processor and a communication interface. The communication interface and the processor are coupled. The processor is configured to execute programs or instructions to implement the steps of the method according to the first aspect.

[0017] In a sixth aspect, a computer program product is provided. The program product is stored in a storage medium. The program product is executed by at least one processor to implement the steps of the method according to the first aspect.

[0018] In the embodiment of the present application, in response to the received load balancing instruction, the historical processing data of each candidate task in the source processor core is obtained, the load balancing instruction includes the source processor core and the target processor core, the source processor core and the target processor core are different processor cores in the plurality of processor cores, and the historical processing data includes the number of execution instructions and the number of cache misses of each candidate task; the utilization rate of each candidate task to the cache of the source processor core is determined according to the number of execution instructions and the number of cache misses of each candidate task; the target task is determined according to the utilization rate of each candidate task to the cache, and the utilization rate of the target task is less than or equal to the utilization rate threshold; and the target task is migrated from the source processor core to the target processor core. Since the utilization rate of the task to the cache can represent the dependency of the task to the cache, that is, the lower the utilization rate, the lower the dependency of the task to the cache, the embodiment selects the task with the utilization rate less than or equal to the utilization rate threshold as the target task to be migrated, that is, selects the task with lower dependency to the cache for migration, so that the target task needs to reload less cache data after migration, thereby reducing the resource overhead brought by load balancing. BRIEF DESCRIPTION OF DRAWINGS

[0019] Figure 1 An application scenario of a load balancing method provided by the embodiment of the present application is shown in the figure.

[0020] Figure 2 A flowchart of a load balancing method provided by the embodiment of the present application is shown in the figure.

[0021] Figure 3 A flowchart of another load balancing method provided by the embodiment of the present application is shown in the figure.

[0022] Figure 4 A flowchart of another load balancing method provided by the embodiment of the present application is shown in the figure.

[0023] Figure 5 A structural diagram of a load balancing device provided by the embodiment of the present application is shown in the figure.

[0024] Figure 6 A structural diagram of an electronic device provided by the embodiment of the present application is shown in the figure.

[0025] Figure 7 A hardware structural diagram of an electronic device provided by the embodiment of the present application is shown in the figure. DETAILED DESCRIPTION

[0026] The technical solutions in the embodiments of the present application will be clearly described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.

[0027] The terms "first," "second," etc., used in this application's specification are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class, without limiting the number of objects; for example, a first object can be one or more. Furthermore, in the specification, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects have an "or" relationship.

[0028] For multi-core processors, to achieve load balancing between different processor cores, related technologies typically select tasks directly from the end of the task queue of a busy processor core and migrate them to an idle processor core. However, in reality, the resource overhead of migrating tasks between different processor cores varies. By directly selecting tasks from the end of the task queue, these technologies ignore the resource overhead of migrating tasks between different processor cores, resulting in a high cost for load balancing.

[0029] Therefore, this application provides a load balancing method and apparatus that can effectively solve the problem of high resource overhead caused by related technologies when performing load balancing on different processor cores, and reduce the cost of resource overhead.

[0030] The load balancing method and apparatus provided in this application will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0031] Figure 1 This is a schematic diagram illustrating an application scenario of a load balancing method provided in an embodiment of this application, such as... Figure 1 As shown, this application scenario can include multiple processor cores, which can be integrated into the same electronic device. This electronic device can be a device with an operating system, such as a mobile phone, tablet, or laptop. The operating system can be, for example, Android or Linux. In some embodiments, the multiple processor cores can be integrated into the central processing unit (CPU) of the electronic device; that is, the CPU is a multi-core CPU.

[0032] Figure 1 Taking a device with three processor cores as an example, in practical applications, electronic devices may also include two or more processor cores.

[0033] Suppose that the processor core 101 currently has 10 tasks to be processed, the processor core 102 currently has 2 tasks to be processed, and the processor core 103 currently has 6 tasks to be processed, in order to improve the processing efficiency of the tasks and improve the performance of the electronic device, the operating system can migrate a part of the tasks on the processor core 101 to the processor core 102, so as to realize load balancing of the processor core 101 and the processor core 102.

[0034] Figure 2 A flowchart of a load balancing method provided by an embodiment of the present application is shown in the figure. The load balancing method can be applied to the electronic device with multiple processor cores as described above, and in particular can be applied to a scheduler in the electronic device, which is in communication connection with an operating system. As shown in the figure, the load balancing method can include the following steps: Figure 2

[0035] S210, in response to the received load balancing instruction, obtaining historical processing data of each candidate task in the source processor core.

[0036] The load balancing instruction includes a source processor core and a target processor core, and the source processor core and the target processor core are different processor cores in the multiple processor cores. The historical processing data includes the number of execution instructions and the number of cache misses of each candidate task.

[0037] S220, determining the utilization rate of the cache of the source processor core by each candidate task according to the number of execution instructions and the number of cache misses of each candidate task.

[0038] S230, determining a target task according to the utilization rate of the cache of each candidate task.

[0039] The utilization rate of the target task is less than or equal to a utilization rate threshold.

[0040] S240, migrating the target task from the source processor core to the target processor core.

[0041] ​In the embodiment of the present application, in response to the received load balancing instruction, the historical processing data of each candidate task in the source processor core is obtained, the load balancing instruction comprises the source processor core and the target processor core, the source processor core and the target processor core are different processor cores in the plurality of processor cores, and the historical processing data comprises the number of execution instructions and the number of cache misses of each candidate task; the utilization rate of each candidate task to the cache of the source processor core is determined according to the number of execution instructions and the number of cache misses of each candidate task; the target task is determined according to the utilization rate of each candidate task to the cache, and the utilization rate of the target task is less than or equal to the utilization rate threshold; and the target task is migrated from the source processor core to the target processor core. Since the utilization rate of the task to the cache can represent the dependency of the task to the cache, that is, the lower the utilization rate, the lower the dependency of the task to the cache, the embodiment selects the task with a utilization rate less than or equal to the utilization rate threshold as the target task to be migrated, that is, selects the task with a lower dependency to the cache for migration, so that the target task needs to reload less cache data after migration, thereby reducing the resource overhead brought by load balancing.

[0042] The above steps are described in detail as follows:

[0043] In S210, the load balancing instruction is used to instruct the operating system to perform load balancing on the source processor core and the target processor core. The source processor core and the target processor core are different processor cores in the plurality of processor cores. For example, the source processor core can be a processor core with a high load, and the target processor core can be a processor core with a low load. Here, the load level can be determined based on the number, complexity, etc. of the tasks to be processed on each processor core. For example, as shown in FIG. 1, the source processor core can be processor core 101, and the target processor core can be processor core 102. Figure 1

[0044] For example, the operating system can periodically receive the load balancing instruction, that is, the operating system can receive the load balancing instruction every certain period of time. For example, a timer can be set in the electronic device, and after the timer ends, the operating system sends the load balancing instruction to the scheduler. For example, the load balancing process can be triggered every 5 seconds, so that the tasks on different processor cores can be adjusted in time, and the processing efficiency of the tasks can be improved. At this time, the operating system can determine the source processor core and the target processor core according to the number of tasks to be processed on each processor core and the complexity of each task. After determining the source processor core and the target processor core, the operating system traverses all the processor cores.

[0045] ​Exemplarily, the operating system can also send a load balancing instruction to the scheduler upon receiving an input of a user, that is, the user can also adjust the tasks on different processor cores as needed. At this time, the source processor core and the target processor core can be determined by the user according to the needs. The input here can include but is not limited to touch input, voice input, gesture input, etc.

[0046] The candidate task here is a task currently in need of processing in the source processor core. In actual application, there can be multiple candidate tasks. The historical processing data is the processing data of the candidate task in a historical time period, which can include, for example, the number of executed instructions and the number of cache misses.

[0047] The number of executed instructions is the total number of all machine instructions executed by the processor core to complete a task, which can be used to measure the computational complexity of the task. For example, the more the number of executed instructions, the more complex the task represents.

[0048] The number of cache misses is the total number of times that the requested data is not in the cache. For example, during the execution of a task, if the requested data is not in the cache for 5 times, it is considered that the number of cache misses is 5. The cache here is the cache corresponding to the CPU.

[0049] The number of executed instructions and the number of cache misses can be recorded by the performance monitoring unit (PMU) of the electronic device. That is, during the execution of the task, the PMU can record the number of executed instructions and the number of cache misses of the task in real time.

[0050] Exemplarily, after receiving the load balancing instruction, the scheduler can read the number of executed instructions and the number of cache misses of each candidate task in the source processor core recorded by the PMU during the historical execution process to obtain the historical processing data, which provides a basis for subsequent load balancing. Exemplarily, the number of executed instructions and the number of cache misses of each candidate task recorded by the PMU during the most recent execution process can be read. Exemplarily, if a candidate task is executed multiple times, the average of the number of executed instructions corresponding to multiple executions and the average of the number of cache misses corresponding to multiple executions can also be calculated, and the average of the number of executed instructions and the average of the number of cache misses are taken as the number of executed instructions and the number of cache misses during the most recent execution process.

[0051] In S220, the utilization rate of the candidate task to the cache can represent the degree of dependence of the candidate task on the cache. The lower the utilization rate, the lower the degree of dependence of the candidate task on the cache. The lower the degree of dependence on the cache, the less data that needs to be reloaded when migrating the task, the smaller the resource overhead generated, and the smaller the cost of load balancing.

[0052] In actual application, the resource overheads generated by different tasks during migration are different. By determining the cache utilization of each candidate task, a more accurate basis can be provided for task migration.

[0053] The cache utilization of each candidate task can be determined based on the number of execution instructions and the number of cache misses of each candidate task.

[0054] For example, the cache utilization of each candidate task can be determined based on a machine learning model or a deep learning model in combination with the number of execution instructions and the number of cache misses of each candidate task.

[0055] For example, the number of instructions executed by each candidate task at each cache miss can also be determined based on the number of execution instructions and the number of cache misses of each candidate task. The fewer the number of instructions executed by each candidate task at each cache miss, the lower the dependency of the candidate task on the cache. Therefore, the number of instructions executed by each candidate task at each cache miss can also be determined as the cache utilization of each candidate task.

[0056] In S230, the size of the utilization threshold can be set according to actual needs, or can be dynamically determined according to the cache utilization of each candidate task.

[0057] Since the utilization can represent the resource overhead generated by the candidate task during migration, according to the cache utilization of each candidate task, the target task that needs to be migrated can be more accurately determined to ensure that the resource overhead generated by task migration is small.

[0058] In this embodiment, the target task can be a task whose utilization is less than or equal to the utilization threshold, that is, the target task can generate smaller resource overhead relative to other tasks. The target task can be located at any position in the task queue.

[0059] In S240, after the target task is determined, the target task can be migrated from the source processor core to the target processor core to achieve load balancing of the source processor core and the target processor core and improve the processing efficiency of the target task.

[0060] For example, before the target task is migrated, the scheduler can save the current state of the target task, such as registers, program counters, stack pointers, floating-point registers, and the like, in a shared memory, such as a task control block, and remove the target task from the task queue of the source processor core. Subsequently, the scheduler can read the saved information from the shared memory and copy it to the target processor core, and add the target task to the task queue of the target processor core. In this way, the scheduler can execute the target task on the target processor core, thereby achieving task migration.

[0061] Exemplarily, to ensure that the target task can be correctly resumed on the target processor core, the source processor core can also directly send the current state of the target task to the target processor core. The updating operation of the task queue of the source processor core and the target processor core is still performed by the scheduler.

[0062] The embodiment can select a task with small migration overhead from the task queue of the source processor core based on the historical processing data of each candidate task, and can ensure that the cost of load balancing is small.

[0063] Figure 3 The flowchart of another load balancing method provided by the embodiment of the application, Figure 3 Different from Figure 2 , the difference is that Figure 2 S210 in the embodiment can be refined as Figure 3 S310-S320 in the embodiment.

[0064] S310, determining cache sharing information of the source processor core and the target processor core.

[0065] The cache sharing information is used to represent which caches are shared by the source processor core and the target processor core, and which caches are not shared. The source processor core and the target processor core share different caches, and the resource overhead generated when the task is migrated can also be different.

[0066] In actual application, the processor core can include a two-level cache, a three-level cache, or a cache with more levels. The cache sharing information of each processor core is determined and saved in the design stage of the processor, and can be directly used subsequently.

[0067] Taking a processor core including three levels of caches as an example, the three levels of caches are respectively a first-level cache, a second-level cache and a third-level cache, in some embodiments, the source processor core and the target processor core can not share all the caches. In some embodiments, the source processor core and the target processor core can also share only the third-level cache, and do not share the first-level cache and the second-level cache. In some embodiments, the source processor core and the target processor core can also share the third-level cache and the second-level cache respectively, and do not share the first-level cache.

[0068] S320, obtaining the historical processing data of each candidate task from the target variable of the thread corresponding to each candidate task according to the cache sharing information.

[0069] Exemplarily, a variable can be added in the structure of the thread corresponding to each candidate task, and is used to store the historical processing data of each candidate task. Different cache sharing information corresponds to different target variables.

[0070] For example, when neither the source processor core nor the target processor core shares all caches, the historical processing data can be stored in variable 1, when the source processor core and the target processor core share only a three-level cache, the historical processing data can be stored in variable 2, and when the source processor core and the target processor core share a three-level cache and a two-level cache respectively, the historical processing data can be stored in variable 3.

[0071] Therefore, according to the cache sharing information of the source processor core and the target processor core, the storage location of the historical processing data can be determined, and the historical processing data corresponding to the candidate task can be obtained.

[0072] According to the cache sharing information of the source processor core and the target processor core, the historical processing data of the candidate task can be read from the corresponding target variable in the embodiment, which can ensure the accuracy of the historical processing data, and further can more accurately determine the cache utilization rate of the candidate task, and further can ensure that the cost of load balancing is small.

[0073] Taking an example that the cache includes at least two levels of caches, the target variable includes an instruction variable and a cache miss variable, and the cache miss variable includes at least two, the number of cache miss variables corresponds to the number of levels of the cache, and exemplarily, S320 can include the following steps:

[0074] Obtaining the number of execution instructions of each candidate task from the instruction variable corresponding to each candidate task;

[0075] In a case where the cache sharing information indicates that neither the source processor core nor the target processor core shares all caches, obtaining the cache miss number of the candidate task from the first cache miss variable, the first cache miss variable being used to store the total number of cache misses of each level of cache of the candidate task;

[0076] In a case where the cache sharing information indicates that the source processor core and the target processor core share part of the cache, obtaining the cache miss number of the candidate task from the second cache miss variable, the second cache miss variable being used to store the number of cache misses of the target cache of the candidate task, the target cache being a cache other than the shared cache among the caches.

[0077] Exemplarily, the instruction variable and the cache miss variable can be added in the structure corresponding to the thread of each candidate task, wherein the instruction variable is used to record the number of execution instructions of the candidate task, and the cache miss variable is used to record the cache miss number of the candidate task. According to the cache sharing information of the two processor cores and the number of caches, a plurality of cache miss variables can be set to store the cache miss numbers corresponding to different cache sharing information through different cache miss variables.

[0078] Exemplarily, the execution instruction number of each candidate task can be obtained from the instruction variable of the structure corresponding to each candidate task.

[0079] For the cache miss number, since different cache sharing information corresponds to different cache miss variables, the cache miss number of the candidate task can be obtained from the corresponding cache miss variable according to the cache sharing information of the source processor core and the target processor core. In actual application, the data in the instruction variable and the cache miss variable are updated synchronously.

[0080] Exemplarily, when neither the source processor core nor the target processor core shares all caches, the total number of cache misses of each level of cache of the candidate task can be obtained from the first cache miss variable. Exemplarily, when the source processor core and the target processor core share a part of the cache, the number of target cache misses of the candidate task can be obtained from the second cache miss variable, and the target cache is the cache that is not shared by the source processor core and the target processor core.

[0081] Taking a cache including three levels as an example, the first level cache, the second level cache and the third level cache, exemplarily, in the case that the first level cache, the second level cache and the third level cache of the source processor core and the target processor core are not shared, the total number of cache misses of each level of cache of the candidate task can be obtained from variable 1, in the case that the third level cache of the source processor core and the target processor core is shared, but the second level cache and the first level cache are not shared, the total number of second level cache misses and first level cache misses of the candidate task can be obtained from variable 2, and in the case that the third level cache and the second level cache of the source processor core and the target processor core are shared respectively, the number of first level cache misses of the candidate task can be obtained from variable 3. Here, variable 1 can be the first cache miss variable in the above embodiment, and variable 2 and variable 3 are collectively referred to as the second cache miss variable.

[0082] The embodiment can obtain the cache miss number of the candidate task from the corresponding cache miss variable according to the cache sharing information of the source processor core and the target processor core, which ensures the accuracy of the cache miss number, and further can more accurately determine the target task that needs to be migrated.

[0083] In order to determine the cache utilization of each candidate task and the target task, Figure 4 Exemplarily, a flowchart of a load balancing method is provided, Figure 4 which is different from Figure 2 in that Figure 2 S220 in Figure 4 may be refined as S410-S420 in Figure 2 S230 in Figure 4 may be refined as S430-S440 in.

[0084] S410, determine the ratio of the number of execution instructions and the number of cache misses of each candidate task.

[0085] S420, determine the utilization of the cache of the source processor core according to the ratio.

[0086] S430, determine the target candidate task with utilization less than or equal to the utilization threshold from the candidate tasks.

[0087] S440, in the case where the target candidate task includes multiple, determine the target candidate task with the minimum utilization as the target task.

[0088] Exemplarily, the ratio of the number of execution instructions and the number of cache misses of each candidate task can be calculated, and in actual application, an additional variable can be added in the structure corresponding to each candidate task to store the ratio, which provides a basis for selecting a core in subsequent load balancing. The ratio changes with the number of execution instructions and the number of cache misses to maintain the consistency of data changes.

[0089] Exemplarily, the ratio can be directly determined as the utilization of the cache of the candidate task, or the result after correction of the ratio can be used as the utilization of the cache of the candidate task.

[0090] After the utilization of each candidate task is determined, the target task to be migrated can be determined based on the utilization of each candidate task.

[0091] Exemplarily, the target candidate task with utilization less than or equal to the utilization threshold can be selected from the candidate tasks based on the utilization of each candidate task. In actual application, there can be one or more target candidate tasks.

[0092] In the case where there is only one target candidate task, the target candidate task can be directly determined as the target task. In the case where there are multiple target candidate tasks, the target candidate task with the minimum utilization can be determined as the target task.

[0093] According to the ratio of the number of execution instructions and the number of cache misses of each candidate task, the utilization of the cache of each candidate task is determined in the embodiment, which can more accurately evaluate the dependency of each candidate task on the cache, so that the target task to be migrated can be more accurately determined. Moreover, when there are multiple target candidate tasks, the target task with the minimum utilization is preferentially selected as the target task, which can achieve the purpose of load balancing with the least resource consumption.

[0094] It should be noted that the load balancing method provided in the embodiments of the present application can be executed by a load balancing device or a processing module in the load balancing device for executing the load balancing method. The load balancing device provided in the embodiments of the present application is taken as an example to illustrate the load balancing device provided in the embodiments of the present application.

[0095] Figure 5 A structural schematic diagram of a load balancing device provided in the embodiments of the present application is shown in FIG. 5. The load balancing device 500 is applied to an electronic device, and the electronic device includes a plurality of processor cores.

[0096] As shown in FIG. 5, the load balancing device 500 can include: Figure 5

[0097] The obtaining module 501 is configured to obtain historical processing data of each candidate task in the source processor core in response to a received load balancing instruction, the load balancing instruction including the source processor core and the target processor core, the source processor core and the target processor core being different processor cores in the plurality of processor cores, and the historical processing data including the number of execution instructions and the number of cache misses of each candidate task.

[0098] The determining module 502 is configured to determine the utilization rate of the cache of the source processor core by each candidate task according to the number of execution instructions and the number of cache misses of each candidate task, and determine a target task according to the utilization rate of the cache of each candidate task, the utilization rate of the target task being less than or equal to a utilization rate threshold.

[0099] The migrating module 503 is configured to migrate the target task from the source processor core to the target processor core.

[0100] In the embodiments of the present application, in response to a received load balancing instruction, the historical processing data of each candidate task in the source processor core is obtained, the load balancing instruction including the source processor core and the target processor core, the source processor core and the target processor core being different processor cores in the plurality of processor cores, and the historical processing data including the number of execution instructions and the number of cache misses of each candidate task. The utilization rate of the cache of the source processor core by each candidate task is determined according to the number of execution instructions and the number of cache misses of each candidate task. The target task is determined according to the utilization rate of the cache of each candidate task, the utilization rate of the target task being less than or equal to a utilization rate threshold. The target task is migrated from the source processor core to the target processor core. Since the utilization rate of the cache by a task can represent the dependency of the task on the cache, that is, the lower the utilization rate, the lower the dependency of the task on the cache, the embodiments select a task with a utilization rate less than or equal to a utilization rate threshold as a target task to be migrated, that is, select a task with a lower dependency on the cache for migration, so that the target task needs to reload less cache data after migration, thereby reducing the resource overhead caused by load balancing. ​

[0101] In some possible implementation of the embodiments of the present application, the determining module 502 is further configured to determine cache sharing information of the source processor core and the target processor core.

[0102] The obtaining module 501 is specifically configured to:

[0103] According to the cache sharing information, the historical processing data of each candidate task is obtained from the target variable of the thread corresponding to each candidate task, and different cache sharing information corresponds to different target variables.

[0104] In some possible implementation of the embodiments of the present application, the cache includes at least two levels of caches, the target variable includes an instruction variable and cache miss variables, the number of cache miss variables corresponds to the number of levels of the cache, and the cache miss variables include at least two.

[0105] The obtaining module 501 is specifically configured to:

[0106] The execution instruction number of each candidate task is obtained from the instruction variable corresponding to each candidate task.

[0107] In a case where the cache sharing information indicates that neither the source processor core nor the target processor core shares all the caches, the cache miss number of the candidate task is obtained from a first cache miss variable, and the first cache miss variable is used to store the total number of cache misses of each level of cache of the candidate task.

[0108] In a case where the cache sharing information indicates that the source processor core and the target processor core share part of the cache, the cache miss number of the candidate task is obtained from a second cache miss variable, and the second cache miss variable is used to store the number of cache misses of the target cache of the candidate task, and the target cache is a cache other than the shared cache in each cache.

[0109] In some possible implementation of the embodiments of the present application, the determining module 502 is specifically configured to:

[0110] The target candidate task with the utilization rate less than or equal to the utilization rate threshold is determined from the candidate tasks.

[0111] In a case where the target candidate task includes multiple target candidate tasks, the target candidate task with the minimum utilization rate is determined as the target task.

[0112] The load balancing apparatus in the embodiments of the present applicationapplicationbe an apparatus or a component in an electronic device, such as an integrated circuit or a chip. For example, the electronic deviceapplicationbe a mobile phone, a tablet computer, a notebook computer, a palm computer, a vehicle-mounted electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), andapplicationalso be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a cash register, or a self-service machine, and the like, and the embodiments of the present application do not make a specific limitation.

[0113] The electronic device in the embodiments of the present applicationapplicationbe a terminal with an operating system. The operating systemapplicationbe an Android operating system, an iOS operating system, or other possible operating systems, and the embodiments of the present application do not make a specific limitation.

[0114] The load balancing apparatus provided in the embodiments of the present applicationapplicationimplement each process in the load balancing method embodiments and achieve the same technical effects. To avoid repetition, details are not described herein. Figure 2 to Figure 4 The load balancing apparatus provided in the embodiments of the present applicationapplicationimplement each process in the load balancing method embodiments and achieve the same technical effects. To avoid repetition, details are not described herein.

[0115] As shown in Figure 6 The embodiments of the present application further provide an electronic device 600, which includes a processor 601 and a memory 602. The memory 602 stores programs or instructions that can be run on the processor 601. When the programs or instructions are executed by the processor 601, each step of the load balancing method embodiments described above is implemented, and the same technical effects are achieved. To avoid repetition, details are not described herein.

[0116] It should be noted that the electronic device in the embodiments of the present applicationapplicationinclude the mobile terminal and the non-mobile terminal described above.

[0117] Figure 7 A hardware structure schematic diagram of an electronic device is provided in the embodiments of the present application. The electronic device 700applicationinclude a plurality of processor cores.

[0118] The electronic device 700 includes, but is not limited to, a radio frequency unit 701, a network module 702, an audio output unit 703, an input unit 704, a sensor 705, a display unit 706, a user input unit 707, an interface unit 708, a memory 709, and a processor 710, etc.

[0119] Those skilled in the art can understand that the electronic device 700 can further include a power supply (such as a battery) for supplying power to each component, and the power supply can be logically connected to the processor 710 through a power management system, so as to realize the functions of managing charging, discharging, and power consumption management through the power management system. Figure 7 The structure of the electronic device 700 shown in the figure does not constitute a limitation on the electronic device 700, and the electronic device 700 can include more or fewer components than those shown, or combine some components, or different component arrangements, which will not be described here.

[0120] The processor 710 is configured to, in response to the received load balancing instruction, obtain historical processing data of each candidate task in the source processor core, the load balancing instruction including the source processor core and the target processor core, the source processor core and the target processor core being different processor cores in a plurality of processor cores, and the historical processing data including the number of execution instructions and the number of cache misses of each candidate task.

[0121] According to the number of execution instructions and the number of cache misses of each candidate task, the utilization rate of the cache of the source processor core by each candidate task is determined.

[0122] According to the utilization rate of the cache of each candidate task, a target task is determined, and the utilization rate of the target task is less than or equal to a utilization rate threshold.

[0123] The target task is migrated from the source processor core to the target processor core.

[0124] In the embodiment of the present application, in response to the received load balancing instruction, the historical processing data of each candidate task in the source processor core is obtained, the load balancing instruction comprises the source processor core and the target processor core, the source processor core and the target processor core are different processor cores in the plurality of processor cores, and the historical processing data comprises the number of execution instructions and the number of cache misses of each candidate task; the utilization rate of each candidate task to the cache of the source processor core is determined according to the number of execution instructions and the number of cache misses of each candidate task; the target task is determined according to the utilization rate of each candidate task to the cache, and the utilization rate of the target task is less than or equal to the utilization rate threshold; and the target task is migrated from the source processor core to the target processor core. Since the utilization rate of the task to the cache can represent the dependency of the task to the cache, that is, the lower the utilization rate, the lower the dependency of the task to the cache, the embodiment selects the task with the utilization rate less than or equal to the utilization rate threshold as the target task to be migrated, that is, selects the task with lower dependency to the cache for migration, so that the target task needs to reload less cache data after migration, thereby reducing the resource overhead brought by load balancing.

[0125] In some possible implementations of the embodiment of the present application, the processor 710 is specifically configured to:

[0126] determine the cache sharing information of the source processor core and the target processor core;

[0127] according to the cache sharing information, obtain the historical processing data of each candidate task from the target variable corresponding to the thread of each candidate task, and different cache sharing information corresponds to different target variables.

[0128] In some possible implementations of the embodiment of the present application, the cache comprises at least two levels of caches, the target variable comprises an instruction variable and a cache miss variable, the number of cache miss variables is at least two, and the number of cache miss variables corresponds to the number of levels of the cache;

[0129] The processor 710 is specifically configured to:

[0130] obtain the number of execution instructions of each candidate task from the instruction variable corresponding to each candidate task;

[0131] in a case where the cache sharing information indicates that the source processor core and the target processor core do not share all caches, obtain the number of cache misses of the candidate task from the first cache miss variable, and the first cache miss variable is used to store the total number of cache misses of each level of cache of the candidate task;

[0132] In a case where the cache sharing information indicates that the source processor core shares part of the cache with the target processor core, the cache miss number of the candidate task is obtained from a second cache miss variable, the second cache miss variable being used to store the number of target caches in which the candidate task misses, and the target cache being a cache other than the shared cache among the caches.

[0133] In some possible implementations of the embodiments of the present application, the processor 710 is specifically configured to:

[0134] determine the ratio of the number of execution instructions to the cache miss number of each candidate task;

[0135] determine the cache utilization of the source processor core by each candidate task according to the ratio.

[0136] In some possible implementations of the embodiments of the present application, the processor 710 is specifically configured to:

[0137] determine a target candidate task from the candidate tasks, the target candidate task having a cache utilization less than or equal to a utilization threshold;

[0138] In a case where the target candidate task includes multiple target candidate tasks, the target task is determined as the target candidate task having the minimum cache utilization.

[0139] It should be understood that, in the embodiments of the present application, the input unit 704 can include a graphics processing unit (GPU) 7041 and a microphone 7042. The graphics processing unit 7041 processes image data of a still picture or a video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 706 can include a display panel 7061, which can be configured in the form of a liquid crystal display, an organic light-emitting diode, and the like. The user input unit 707 includes at least one of a touch panel 7071 and other input devices 7072. The touch panel 7071 is also referred to as a touch screen. The touch panel 7071 can include a touch detection device and a touch controller. The other input devices 7072 can include, but are not limited to, a physical keyboard, function keys (such as volume control keys, on-off keys, and the like), a trackball, a mouse, an operation lever, and the like, which are not described herein.

[0140] The memory 709 can be used to store software programs and various data. The memory 709 can mainly include a first storage area storing programs or instructions and a second storage area storing data, wherein the first storage area can store an operating system, application programs or instructions required by at least one function (such as a sound playing function, an image playing function, etc.), and the like. In addition, the memory 709 can include a volatile memory or a non-volatile memory, or the memory 709 can include both a volatile memory and a non-volatile memory. The non-volatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically EPROM (EEPROM), or a flash memory. The volatile memory can be a Random Access Memory (RAM), a Static RAM (SRAM), a Dynamic RAM (DRAM), a Synchronous DRAM (SDRAM), a Double Data Rate SDRAM (DDR SDRAM), an Enhanced SDRAM (ESDRAM), a Synch link DRAM (SLDRAM), and a Direct Rambus RAM (DRRAM). The memory 709 in the embodiments of the present application includes but is not limited to these and any other suitable types of memory.

[0141] The processor 710 can include one or more processing units; optionally, the processor 710 integrates an application processor and a modem processor, wherein the application processor mainly processes operations related to an operating system, a user interface, and an application program, and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above-mentioned modem processor can also not be integrated into the processor 710.

[0142] The embodiments of the present application also provide a readable storage medium, and the readable storage medium stores programs or instructions, the programs or instructions are executed by a processor to realize various processes of the above-mentioned load balancing method embodiments and achieve the same technical effects. To avoid repetition, details are not described here.

[0143] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes a computer readable storage medium, such as a computer readable only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.

[0144] The chip provided in the embodiments of the present application includes a processor and a communication interface, the communication interface is coupled with the processor, the processor is used to run programs or instructions to realize the processes of the load balancing method embodiments and achieve the same technical effects. To avoid repetition, details are not described herein.

[0145] It should be understood that the chip mentioned in the embodiments of the present application can also be referred to as a system-level chip, a system chip, a chip system, or a system-on-chip, etc.

[0146] The embodiments of the present application provide a computer program product stored in a storage medium, the program product is executed by at least one processor to realize the processes of the load balancing method embodiments and achieve the same technical effects. To avoid repetition, details are not described herein.

[0147] It should be noted that in this document, the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusion, so that processes, methods, articles, or devices including a series of elements not only include those elements, but also include other elements not explicitly listed, or include elements inherent to such processes, methods, articles, or devices. Without more limitations, the element defined by the statement "including a" does not exclude the presence of additional identical elements in the process, method, article, or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to the order of performing the functions as shown or discussed, but can also include performing the functions in a substantially simultaneous manner or in a reverse order, for example, the described method can be performed in an order different from the described order, and various steps can be added, omitted, or combined. In addition, the features described with reference to certain examples can be combined in other examples.

[0148] From the above description of the embodiments, those skilled in the art can clearly understand that the above-mentioned method embodiments can be realized by means of software and necessary general hardware platforms, of course, they can also be realized by hardware, but in many cases the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk, etc.), and includes a plurality of instructions for making a terminal (which can be a mobile phone, computer, server, or network device, etc.) execute the methods described in the embodiments of the present application.

[0149] The embodiments of the present application are described above with reference to the accompanying drawings, but the present application is not limited to the above-described specific embodiments, and the above-described specific embodiments are merely illustrative, but not restrictive, and a person of ordinary skill in the art can make many forms under the inspiration of the present application without departing from the purpose of the present application and the scope protected by the claims.

Claims

1. A load balancing method applied to electronic devices, characterized in that, The electronic device includes multiple processor cores, and the method includes: In response to a received load balancing instruction, historical processing data of each candidate task in the source processor core is obtained. The load balancing instruction includes the source processor core and the target processor core, which are different processor cores among a plurality of processor cores. The historical processing data includes the number of execution instructions and the number of cache misses for each candidate task. The cache utilization rate of the source processor core for each candidate task is determined based on the number of execution instructions and cache misses for each candidate task. Based on the utilization rate of the cache by each of the candidate tasks, a target task is determined, wherein the utilization rate of the target task is less than or equal to a utilization threshold. The target task is migrated from the source processor core to the target processor core.

2. The method according to claim 1, characterized in that, The acquisition of historical processing data for each candidate task in the source processor core includes: Determine the cache sharing information of the source processor core and the target processor core; Based on the cache sharing information, the historical processing data of each candidate task is obtained from the target variable of the thread corresponding to each candidate task. Different cache sharing information corresponds to different target variables.

3. The method according to claim 2, characterized in that, The cache includes at least two levels of cache, the target variable includes instruction variables and cache miss variables, the cache miss variables include at least two, and the number of cache miss variables corresponds to the number of levels of the cache; The step of obtaining historical processing data for each candidate task from the target variable of the thread corresponding to each candidate task based on the cache sharing information includes: Obtain the number of execution instructions for each candidate task from the instruction variable corresponding to each candidate task; When the cache sharing information indicates that neither the source processor core nor the target processor core shares all caches, the number of cache misses of the candidate task is obtained from the first cache miss variable, where the first cache miss variable is used to store the total number of times the candidate task misses all levels of cache. When the cache sharing information indicates that the source processor core and the target processor core share a portion of the cache, the number of cache misses of the candidate task is obtained from the second cache miss variable. The second cache miss variable is used to store the number of times the candidate task misses the target cache, and the target cache is the cache other than the shared cache among the caches.

4. The method according to any one of claims 1-3, characterized in that, The step of determining the cache utilization of the source processor core for each candidate task based on the number of executed instructions and cache misses for each candidate task includes: Determine the ratio of the number of execution instructions to the number of cache misses for each candidate task; The utilization rate of the source processor core's cache for each candidate task is determined based on the ratio.

5. The method according to any one of claims 1-3, characterized in that, The step of determining the target task based on the utilization efficiency of the cache by each of the candidate tasks includes: From each of the candidate tasks, identify target candidate tasks whose utilization rate is less than or equal to the utilization rate threshold; When there are multiple target candidate tasks, the target candidate task with the lowest utilization rate is determined as the target task.

6. A load balancing device, applied to electronic equipment, characterized in that, The electronic device includes multiple processor cores, and the device includes: The acquisition module is used to acquire historical processing data of each candidate task in the source processor core in response to the received load balancing instruction. The load balancing instruction includes the source processor core and the target processor core. The source processor core and the target processor core are different processor cores among multiple processor cores. The historical processing data includes the number of execution instructions and the number of cache misses for each candidate task. The determining module is configured to determine the cache utilization rate of the source processor core for each candidate task based on the number of execution instructions and cache misses of each candidate task; and to determine the target task based on the cache utilization rate of each candidate task, wherein the utilization rate of the target task is less than or equal to a utilization threshold. A migration module is used to migrate the target task from the source processor core to the target processor core.

7. The apparatus according to claim 6, characterized in that, The determining module is further configured to determine cache sharing information between the source processor core and the target processor core; The acquisition module is specifically used for: Based on the cache sharing information, the historical processing data of each candidate task is obtained from the target variable of the thread corresponding to each candidate task. Different cache sharing information corresponds to different target variables.

8. The apparatus according to claim 7, characterized in that, The cache includes at least two levels of cache, the target variable includes instruction variables and cache miss variables, the cache miss variables include at least two, and the number of cache miss variables corresponds to the number of levels of the cache; The acquisition module is specifically used for: Obtain the number of execution instructions for each candidate task from the instruction variable corresponding to each candidate task; When the cache sharing information indicates that neither the source processor core nor the target processor core shares all caches, the number of cache misses of the candidate task is obtained from the first cache miss variable, where the first cache miss variable is used to store the total number of times the candidate task misses all levels of cache. When the cache sharing information indicates that the source processor core and the target processor core share a portion of the cache, the number of cache misses of the candidate task is obtained from the second cache miss variable. The second cache miss variable is used to store the number of times the candidate task misses the target cache, and the target cache is the cache other than the shared cache among the caches.

9. The apparatus according to any one of claims 6-8, characterized in that, The determining module is specifically used for: Determine the ratio of the number of execution instructions to the number of cache misses for each candidate task; The utilization rate of the source processor core's cache for each candidate task is determined based on the ratio.

10. The apparatus according to any one of claims 6-8, characterized in that, The determining module is specifically used for: From each of the candidate tasks, identify target candidate tasks whose utilization rate is less than or equal to the utilization rate threshold; When there are multiple target candidate tasks, the target candidate task with the lowest utilization rate is determined as the target task.