Task processing method and apparatus, and device and storage medium

By gradually adjusting the calculation density level of the computing unit before and after task processing, the problem of excessive chip current change rate is solved, ensuring the stability of chip performance and functional integrity.

WO2025118503A1PCT designated stage expired Publication Date: 2025-06-12HYGON INFORMATION TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/096229
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-04
Filing Date
2024-05-30
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

When processing tasks, the rate of change of chip current over time (DIDT) is too large, resulting in a decrease in the power supply voltage of the internal devices of the chip and an increase in the device delay, which may cause the entire chip function to fail.

Method used

By gradually adjusting the calculation density level of the calculation unit before and after task processing, ensure that the current change rate is within a reasonable range. The specific method includes gradually increasing the calculation density to the first calculation density level before processing the task, and gradually reducing it to the second calculation density level after processing the task is completed.

Benefits of technology

It effectively solves the chip performance problem caused by excessive DIDT, ensures stable power supply of internal components of the chip, and avoids functional failure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024096229_12062025_PF_FP_ABST
    Figure CN2024096229_12062025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present disclosure are a task processing method and apparatus, and a device and a storage medium. The method comprises: determining a first computational density level used for processing one or more tasks; within a first time period before the one or more tasks are processed, gradually increasing the computational density of a computing unit to the first computational density level, which computing unit is used for processing the one or more tasks; processing the one or more tasks by means of the computing unit; and within a second time period after the processing of the one or more tasks is completed, gradually reducing the computational density of the computing unit to a second computational density level.
Need to check novelty before this filing date? Find Prior Art

Description

Method, device, apparatus and storage medium for task processing

[0001] This application claims priority to Chinese Patent Application No. 202311644046.2 filed on December 4, 2023, the entire text of which is incorporated herein by reference as a part of this application. Technical Field

[0002] Embodiments of the present disclosure relate to a method, device, apparatus, and storage medium for task processing. Background Art

[0003] A typical chip power supply scenario is shown in Figure 1. The power output by power supply 102 enters the chip substrate 103 through the internal trace AB of circuit board 101, and then enters the chip 104 through the power connection BC between the substrate and the chip pin. Inside the chip, it is connected to the power network 105 through the power trace CD. The internal chip device 106 is connected to the power network 105 through the power pin of the device through the connection E to obtain power. When the chip load is small or no load, the power supply can provide a stable voltage output and all the voltage acts on the internal chip devices. However, when the chip load suddenly increases, the voltage on the internal chip devices may drop sharply due to the following two reasons:

[0004] 1) A sudden increase in chip load may increase the rate of change of current over time (DIDT) on the chip, increasing the energy required by the chip. However, the power supply needs a certain amount of time to respond to the sudden increase in energy. If the DIDT is too large (i.e., the current changes too quickly), the power supply may not be able to provide enough energy to meet the chip's needs in a short period of time.

[0005] 2) As the chip load increases, the current increases sharply, the internal resistance of the chip decreases, and the voltage distributed to the internal devices decreases.

[0006] These two issues often occur simultaneously, causing the supply voltage of internal chip components to drop, increasing device latency and delays. This can lead to register setup times on critical paths not being met, potentially causing the entire chip to fail. Further analysis reveals that the cause of excessive DIDT is primarily due to the simultaneous activation of multiple cores during multi-core computing, resulting in a rapid increase in current within a short period of time.

[0007] Therefore, a method is needed to effectively solve the problem of excessive change rate of chip current over time (DIDT) when processing tasks.

[0008] Summary of the Invention

[0009] This disclosure section is provided to briefly introduce concepts that will be described in detail in the detailed description section below. This disclosure section is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0010] An embodiment of the present disclosure provides a method for task processing, comprising: determining a first computing density level for processing one or more tasks; gradually increasing the computing density of a computing unit for processing the one or more tasks to the first computing density level within a first time period before processing the one or more tasks; processing the one or more tasks by the computing unit; and gradually reducing the computing density of the computing unit to a second computing density level within a second time period after processing of the one or more tasks is completed.

[0011] According to an embodiment of the present disclosure, determining the first computing density level for processing one or more tasks includes: determining the classified statistical number of instructions to be processed included in the one or more tasks according to a predetermined category; and determining the first computing density level for processing the one or more tasks based on the classified statistical number.

[0012] According to an embodiment of the present disclosure, the predetermined categories include one or more of the following: low-density operation instructions, medium-low-density operation instructions, medium-density operation instructions, medium-high-density operation instructions, and high-density operation instructions.

[0013] According to an embodiment of the present disclosure, the method further includes: gradually increasing the computing density of the computing unit to a preset third computing density level within a third time period before the first time period.

[0014] According to an embodiment of the present disclosure, the first time period is a time period for acquiring instructions to be processed included in the one or more tasks.

[0015] According to an embodiment of the present disclosure, within a first time period before processing the one or more tasks, gradually increasing the computing density of the computing unit used to process the one or more tasks to the first computing density level includes: within the first time period, sequentially providing the computing unit with one or more first preset data with gradually increasing computing density for processing by the computing unit, wherein the first first preset data among the one or more first preset data corresponds to the current computing density level of the computing unit, and the last first preset data among the one or more first preset data corresponds to the first computing density level.

[0016] According to an embodiment of the present disclosure, the one or more first preset data are pre-stored in a database.

[0017] According to an embodiment of the present disclosure, within a second time period after the processing of the one or more tasks is completed, gradually reducing the computing density of the computing unit to a second computing density level includes: within the second time period, sequentially providing one or more second preset data with gradually decreasing computing density to the computing unit for processing by the computing unit, wherein the first second preset data among the one or more second preset data corresponds to the current computing density level of the computing unit, and the last second preset data among the one or more second preset data corresponds to the second computing density level.

[0018] According to an embodiment of the present disclosure, the one or more second preset data are pre-stored in a database.

[0019] According to an embodiment of the present disclosure, the method also includes: when one or more new tasks arrive within the second time period, stopping the gradual reduction of the computing density; and gradually increasing the computing density of the computing unit to a fourth computing density level within a fourth time period from the time of stopping the gradual reduction of the computing density, wherein the fourth computing density level is a computing density level for processing the new one or more tasks.

[0020] An embodiment of the present disclosure provides an apparatus for task processing, comprising: a determination unit configured to determine a first computing density level for processing one or more tasks; a first preheating unit configured to gradually increase the computing density of a computing unit for processing the one or more tasks to the first computing density level within a first time period before processing the one or more tasks; a computing unit configured to process the one or more tasks through the computing unit; and a stretching unit configured to gradually reduce the computing density of the computing unit to a second computing density level within a second time period after processing of the one or more tasks is completed.

[0021] According to an embodiment of the present disclosure, the determination unit is configured to determine a first computing density level for processing one or more tasks, including: determining a classified statistical number of instructions to be processed included in the one or more tasks according to a predetermined category; and determining the first computing density level for processing the one or more tasks based on the classified statistical number.

[0022] According to an embodiment of the present disclosure, the predetermined categories include one or more of the following: low-density operation instructions, medium-low-density operation instructions, medium-density operation instructions, medium-high-density operation instructions, and high-density operation instructions.

[0023] According to an embodiment of the present disclosure, the apparatus further includes: a second preheating unit configured to gradually increase the computing density of the computing unit to a preset third computing density level within a third time period before the first time period.

[0024] According to an embodiment of the present disclosure, the first time period is a time period for acquiring instructions to be processed included in the one or more tasks.

[0025] According to an embodiment of the present disclosure, the first preheating unit is configured to gradually increase the computing density of the computing unit used to process the one or more tasks to the first computing density level within a first time period before processing the one or more tasks, including: within the first time period, sequentially providing the computing unit with one or more first preset data with gradually increasing computing density for processing by the computing unit, wherein the first first preset data among the one or more first preset data corresponds to the current computing density level of the computing unit, and the last first preset data among the one or more first preset data corresponds to the first computing density level.

[0026] According to an embodiment of the present disclosure, the one or more first preset data are pre-stored in a database.

[0027] According to an embodiment of the present disclosure, the stretching unit is configured to gradually reduce the computing density of the computing unit to a second computing density level within a second time period after the processing of the one or more tasks is completed, including: within the second time period, sequentially providing one or more second preset data with gradually decreasing computing density to the computing unit for processing by the computing unit, wherein the first second preset data among the one or more second preset data corresponds to the current computing density level of the computing unit, and the last second preset data among the one or more second preset data corresponds to the second computing density level.

[0028] According to an embodiment of the present disclosure, the one or more second preset data are pre-stored in a database.

[0029] According to an embodiment of the present disclosure, the stretching unit is further configured to: stop the gradual reduction of the computing density when one or more new tasks arrive within the second time period; and the first preheating unit is further configured to: gradually increase the computing density of the computing unit to a fourth computing density level within a fourth time period from the time when the gradual reduction of the computing density is stopped, wherein the fourth computing density level is a computing density level for processing the new one or more tasks.

[0030] An embodiment of the present disclosure provides a device for task processing, comprising: at least one processor; and a storage device having at least one program stored thereon, which enables the at least one processor to implement any method according to an embodiment of the present disclosure when the at least one program is executed by the at least one processor.

[0031] An embodiment of the present disclosure provides an electronic device including any device for task processing according to an embodiment of the present disclosure.

[0032] An embodiment of the present disclosure provides a non-transitory storage medium containing computer-executable instructions, wherein the computer-executable instructions, when executed by a computer processor, are used to perform any method according to an embodiment of the present disclosure.

[0033] The method, device and apparatus provided in the present disclosure can effectively solve the problem that the chip current change rate over time (DIDT) is too large when processing tasks, thereby affecting system performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.

[0035] Figure 1 shows a chip power supply scenario;

[0036] FIG2 shows a schematic diagram of a general-purpose computing multi-core processor according to an embodiment of the present disclosure;

[0037] Figure 3 shows the current situation of a general computing multi-core processor during operation;

[0038] FIG4 shows a schematic diagram of an improved general-purpose computing multi-core processor according to an embodiment of the present disclosure;

[0039] FIG5 shows a schematic flow chart of counting unfinished tasks according to an embodiment of the present disclosure;

[0040] FIG6 shows an example classification of instructions according to their power consumption according to an embodiment of the present disclosure;

[0041] FIG7 is a schematic flow chart showing the counting of pending instructions of a specific category in an unfinished task according to an embodiment of the present disclosure;

[0042] FIG8 is a schematic diagram showing data sources acquired by a computing unit according to an embodiment of the present disclosure;

[0043] FIG9 illustrates an example decision-making mechanism of a DIDT control unit according to an embodiment of the present disclosure;

[0044] FIG10 shows an example control flow of a control core according to an embodiment of the present disclosure;

[0045] FIG11 shows an example of an immediate number database according to an embodiment of the present disclosure;

[0046] FIG12 shows a graph showing changes in current over time of a computing unit during a preheating phase according to an embodiment of the present disclosure;

[0047] FIG13 shows the current of the computing unit during operation after being processed by the DIDT control unit according to an embodiment of the present disclosure;

[0048] FIG14 shows a method for task processing according to an embodiment of the present disclosure;

[0049] FIG15 shows a schematic diagram of an apparatus for task processing according to an embodiment of the present disclosure; and

[0050] FIG16 shows a schematic diagram of another apparatus for task processing according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0051] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0052] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0053] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.

[0054] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0055] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0056] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0057] Before beginning the description of the embodiments of the present disclosure, some technical terms that may be involved in the present disclosure are first introduced.

[0058] PVT: chip process, voltage and temperature.

[0059] DIDT: The rate of change of current over time. The greater the current change amplitude and / or the shorter the change time, the greater the DIDT.

[0060] IR Drop: Voltage drop, which specifically refers to the instantaneous excessive DIDT, causing the chip power supply voltage output to be lower than the specified requirement and / or the overall voltage drop generated from the power supply output to the chip transistor power supply.

[0061] In this disclosure, multi-core may refer to multiple computing units. In some embodiments, a core may be a computing unit. In some embodiments, a core may also be a combination of multiple computing units.

[0062] The following technologies can be used to solve the chip failure problem caused by DIDT:

[0063] 1. Estimate the voltage drop during severe DIDT and increase the output voltage of the power supply on the chip board to ensure that the power supply to the internal components of the chip still meets the minimum power supply requirements in severe DIDT conditions.

[0064] 2. Voltage droop detection and clock stretch are introduced inside the chip. The clock stretch amplitude is set according to the voltage droop detection amplitude, thereby reducing the clock frequency and chip load, thereby alleviating the voltage droop situation.

[0065] 3. A module that automatically estimates the current load is added to the chip. When the load increases sharply, the chip load is reduced to alleviate DIDT by reducing task dispatching through the internal data processing unit of the gating chip.

[0066] 4. For multi-core processing systems, when the chip is turned on, control the time when multiple cores are turned on, and gradually turn on multiple cores to avoid the generation of large currents.

[0067] The above technology can alleviate the voltage drop problem caused by DIDT to a certain extent, thereby preventing chip function failure, but it also has the following disadvantages:

[0068] Regarding #1: Since the chip power supply is increased, the chip power consumption will increase. When the load decreases sharply, the voltage of the internal components of the chip will rise, causing voltage overshoot and the risk of burning out the internal components of the chip.

[0069] Regarding #2: The relationship between the amplitude of the internal chip voltage drop detection and the amplitude of the clock stretching is difficult to determine. If the relationship is too loose, the clock stretching amplitude will be too large, resulting in performance loss. It also does not take into account the individual differences between chips. In some specific situations, such as when spread spectrum clocking (SSC) is enabled, the clock stretching may fail.

[0070] Regarding #3: Reducing task allocation based on load changes may cause system performance to degrade when DIDT occurs.

[0071] For #4: The DIDT problem can be better controlled, but it may lead to performance loss because it takes some time for all cores to start running.

[0072] This disclosure addresses the issue of chip performance degradation caused by excessive DIDT. Based on further hardware analysis, this disclosure proposes an adaptive DIDT solution for general-purpose multi-core processors. According to the disclosed method, the computing unit can be "preheated" (i.e., gradually achieving a high current consumption state or computing density level before officially starting the actual computing task, which may also be referred to as pre-starting the computing unit) during the interval between instruction and / or data fetches by each core, depending on the task allocation. After the data is retrieved, the computing unit is directly switched to the computing task for calculation. Throughout the computing process, the computing unit's state is adaptively adjusted based on the task allocation, and the current consumption state or computing density level of the computing unit is gradually and slowly reduced after the computing task is completed. This achieves "preheating" of the computing unit's state before officially starting the actual computing task, and "stretching" of the computing unit's state after the computing task is completed. Since both the "preheating" and "stretching" time are idle time for the computing unit to fetch instructions and / or data, the problem of excessive DIDT can be solved without reducing chip performance. In this article, the (operating) state of a computing unit may refer to the computing unit's current, power consumption, computing density, and their equivalents. Furthermore, in this document, computing density may refer to instantaneous current or instantaneous power consumption of a computing unit when executing a calculation (eg, an instruction, a task, etc.).

[0073] FIG2 shows a schematic diagram of a general-purpose computing multi-core processor 200 according to an embodiment of the present disclosure.

[0074] As shown in FIG2 , the general computing multi-core processor 200 mainly includes a task fetch & first-level task allocation unit 201, a second-level task allocation unit 202, a third-level task allocation unit 203, an instruction fetch unit 204, a decoding & operand fetch unit 205, a computing unit 206, a cache 207, a data routing 208, an off-chip storage 209, and a high-speed interface 210. The structure shown in FIG2 is merely an example, and the general computing multi-core processor 200 may also include more or fewer units than FIG2 . For example, in other embodiments, it may include less than three levels of task allocation units or more than three levels of task allocation units, etc., which is not limited herein. An example workflow of the general computing multi-core processor 200 is as follows:

[0075] 1) After the chip is started, the processor, such as the CPU, writes the task to be executed (including instructions and / or related data) to the off-chip storage 209 through the high-speed interface 210 (e.g., PCIE), and notifies the task fetching & primary task allocation unit 201 through the internal connection 221.

[0076] 2) After receiving the notification, the task retrieval & first-level task allocation unit 201 retrieves the task according to the notification address (in this article, the task retrieved by the task retrieval & first-level task allocation unit 201 can be called a total task or a first-level task or a first task, or a task set including one or more tasks), and disassembles and allocates the first-level task according to the resource usage information fed back by the second-level task allocation unit, and sends the allocated second-level tasks to multiple second-level task allocation units 202.

[0077] 3) The secondary task allocation unit 202 further decomposes and allocates the secondary tasks according to the resource usage feedback from the tertiary task allocation unit 203 , and distributes the decomposed tertiary tasks to multiple tertiary task allocation units 203 .

[0078] 4) After the third-level task allocation unit 203 accepts the task, it allocates the third-level task to multiple different instruction fetch units 204 based on the resource usage feedback from the instruction fetch unit 204. An instruction fetch unit 204 may be allocated one or more tasks (in this article, it can be described using wave as an example) (in this article, it can also be called the second task).

[0079] 5) The instruction fetch unit 204 fetches the corresponding instruction through the instruction fetch bus 222 according to the address of the task assigned by the three-level task allocation unit. Instructions of multiple tasks can be fetched in a round-robin manner and sent to the decoding & operand fetching unit 205.

[0080] 6) Decode & Operand Fetch Unit 205 decodes the instruction and reads data from cache 207 via internal bus 223. If the required data is stored in cache 207 (or a cache hit), the data is directly fed back. If the required data is not stored in cache 207, cache 207 reads data from off-chip storage 209 via data routing 208, and feeds the data back to Decode & Operand Fetch Unit 205.

[0081] 7) After all data are ready, the decoded instruction will be sent to the calculation unit 206 for calculation, and the calculation result can be written into the cache 207.

[0082] 8) The cache 207 can further write the calculation results to the off-chip storage 209.

[0083] 9) When the task fetching & first-level task assignment unit 201 detects that all assigned tasks are completed, it can generate an interrupt and report to the CPU through the high-speed interface 210. After the CPU receives and analyzes the interrupt, it can read the calculation results through the high-speed interface.

[0084] Figure 3 shows the current consumption of a general-purpose computing multi-core processor 200 during operation. As shown in the upper diagram of Figure 3 , the computation of a task set (which may include one or more tasks executed on the same computing unit) may go through the process shown in the figure, T0-T5. The specific current change process is analyzed as follows:

[0085] After the chip is started, the workflow of steps 1) to 5) above is executed.

[0086] At T0, when the chip is initialized, the cache 207 is empty and misses the data, and the above step 6) of fetching data is started. Compared with T0, the current in this process is increased, but the overall current is not large (between T1 and T2).

[0087] Between T2 and T3, all data (for example, the last piece of data that the computing task depends on) arrives within a small adjacent time window. Thousands of computing units may start working simultaneously within a very short time window, causing the current to rise sharply (d1).

[0088] Between T3 and T4, when the computing unit is fully operational, the current reaches a very large value and remains at this value. There may be slight current fluctuations depending on the specific numbers calculated, but the overall value is high.

[0089] Between T4 and T5, when the computing task is completed, all computing units end the calculation almost simultaneously, and the current begins to drop sharply (d2).

[0090] After the calculation is completed, the calculation completion can be reported through an interrupt.

[0091] After completing the calculation of one task set, you can start the calculation of the next task set and repeat the above process.

[0092] As shown in Figure 3, excessive DIDT during the calculation process of a task set usually occurs at the beginning (for example, d1 or d3) and end (for example, d2 or d4) stages. During the entire calculation process, the efficiency of computing resource utilization is generally above 90%. Therefore, as long as the problem of rapid current change at the beginning and end of the task set execution is solved and the use of computing resources during intermediate calculations is controlled, the problem of excessive DIDT can be effectively solved.

[0093] To address these issues, the present disclosure proposes: after determining a computing task, the computing unit is "preheated" using a set of preset data during the data preparation period, ensuring that the current is slowly increased to the current level or computing density required for the calculation before the data arrives. After the calculation is completed, the computing unit is not allowed to terminate immediately. Instead, the computing unit's operating state is "stretched" using a set of preset data, ensuring that the chip's current (or computing density) slowly transitions to a low current level (or low computing density level).

[0094] In addition, if the next computing task (for example, the next task set) comes before the above-mentioned stretching process is completed, the "stretching" can be ended immediately and the "preheating" of the next task set can be turned to determine the new preheating temperature and equality based on the instructions included in the next task set.

[0095] By using the above-mentioned means, the change in current can obtain the result shown in Figure 3, thus solving the problem of excessive DIDT from a design perspective.

[0096] The embodiments of the present disclosure further improve the processor architecture. Specifically, FIG4 shows a schematic diagram of an improved general-purpose computing multi-core processor 400 according to an embodiment of the present disclosure. The main improvements are described below.

[0097] As shown in FIG4 , the general computing multi-core processor 400 adds a DIDT control unit 211 to each computing unit.

[0098] The input 224 of the DIDT control unit 211 comes from the instruction fetch unit 204 and includes the count of unfinished tasks (or pending tasks) in the single instruction multiple data (SIMD) structure (instruction fetch unit 204, decode & operand fetch unit 205, and computation unit 206) (as shown in Figure 5 below) and the power consumption classification count information of the pending instructions fetched by each task (as shown in Figures 6 and 7 below). The output of the DIDT control unit 211 directly controls the computation unit 206. The specific control method is shown in Figure 8.

[0099] FIG5 shows a schematic flowchart of counting unfinished tasks according to an embodiment of the present disclosure.

[0100] As shown in FIG5 , an initial reset is performed in block 501 , and the task count wave_cnt is assigned a value of 0.

[0101] In block 502, a check is performed at each clock cycle to determine whether a task has been assigned or completed. If a task has been assigned but not completed, the task counter wave_cnt is incremented by 1 in block 503. Otherwise, the process proceeds to block 504. In block 504, if a task has been completed but not assigned, the task counter wave_cnt is decremented by 1 in block 505. Otherwise, the process proceeds to block 506, where the task counter wave_cnt remains unchanged, and the process proceeds to the next clock cycle statistics.

[0102] Through statistics, the number of tasks that have not yet been completed can be known from the task count wave_cnt, and wave_cnt can be output to the DIDT control unit 211.

[0103] The statistical counts of the pending (or uncompleted) instructions retrieved by each task according to the predetermined categories are shown in FIG6 and FIG7 .

[0104] Specifically, FIG6 shows an example classification of instructions according to the power consumption of the instructions (or, according to the computational density level required to execute the instructions) according to an embodiment of the present disclosure. As shown in FIG6 , instructions can be classified into five categories based on the power consumption of the instructions or the computational density level required to execute the instructions, namely, low-density operation instructions, medium-low-density operation instructions, medium-density operation instructions, medium-high-density operation instructions, and high-density operation instructions. It should be understood that the categories shown in FIG6 are merely examples, and in other embodiments, the instructions may be divided into more or fewer categories, or any other categories.

[0105] The pending instructions of each task can be counted according to each category. The statistical method is similar to the task counting method shown in FIG5 . A schematic flowchart is shown in FIG7 .

[0106] Specifically, FIG7 shows a schematic flow chart of counting pending instructions of a specific category in an unfinished task according to an embodiment of the present disclosure.

[0107] As shown in Figure 7, taking the x instruction as an example (the x instruction can be, for example, any of the five instruction categories mentioned above), an initial reset is performed in block 701, and the x instruction count x_ist_cnt is assigned to 0. In block 702, it is checked in each clock cycle whether an x ​​instruction is retrieved and an x ​​instruction is executed. If only an x ​​instruction is retrieved but no x instruction is executed, then in block 703, the x instruction count x_ist_cnt is increased by 1, otherwise the process proceeds to block 704. In block 704, if only an x ​​instruction is executed but no x instruction is retrieved, then in block 705, the x instruction count x_ist_cnt is decreased by 1, otherwise the process proceeds to block 706, the x instruction count x_ist_cnt remains unchanged, and then the statistics of the next clock cycle are entered.

[0108] The statistical processes shown in Figures 5 and 7 can be executed in the instruction fetch unit 204. The instruction fetch unit 204 can send all statistical results (e.g., statistical counts of unfinished tasks, classified statistical counts of pending instructions for each task, etc.) to the DIDT control unit 211. Different tasks can be distinguished by their corresponding task identifiers, wave_ids. The information sent along with the statistical results can also include the wave_id of the currently executing task (e.g., wave_cur_id) and the wave_id of the next (or multiple subsequent) tasks to be switched to (or executed) (e.g., wave_nxt_id).

[0109] The DIDT control unit 211 processes the received statistical results or statistical information, and after processing, outputs preset data (eg, immediate data) to the calculation unit 206 for “preheating” and “stretching” the operation status of the calculation unit 206 .

[0110] FIG8 is a schematic diagram illustrating the source of data obtained by a computing unit according to an embodiment of the present disclosure. As shown in FIG8 , the data obtained by the computing unit 206 can come from two sources. The DIDT control unit 211 processes the statistical information 801 (e.g., task and instruction statistical information) sent thereto, and combines it with the instruction status currently being processed by the computing unit (e.g., instruction status 802 during normal SIMD instruction execution) to generate an immediate value 803 to be provided to the computing unit 206 for processing. This immediate value is selected from the data 804 sent from the normal SIMD instruction execution (e.g., via a multiplexer (MUX) 805). The selection control signal 806 can be an indication signal from the normal SIMD instruction execution. When a normal SIMD instruction is being executed, the selection control signal 806 instructs the multiplexer 805 to select the data 804 sent from the normal SIMD instruction execution to enter the computing unit 206 for normal computing task execution. Otherwise, the selection control signal 806 instructs the multiplexer 805 to select the immediate value 803. In this case, the data for the computing unit 206 comes from the DIDT control unit 211. This priority selection method can ensure that the performance of the calculation unit 206 during normal operation is not affected by the DIDT control unit 211.

[0111] FIG9 shows an example decision-making mechanism of the DIDT control unit 211 according to an embodiment of the present disclosure.

[0112] As shown in FIG9 , the control core 901 of the DIDT control unit 211 receives various information, including task statistics 902, classification statistics 903 of the instructions of each task, the identifier wave_cur_id of the currently executing task and the identifier wave_nxt_id (904) of the subsequent task to be executed, and the instruction 905 currently being executed by the computing unit. After processing and judging, the control core 901 outputs a control command 911 to the immediate number generation module 906. The immediate number generation module 906 searches the immediate number database 907 based on the command 911 to obtain data 912 (e.g., one or more immediate numbers). The immediate number generation module 906 then sends the data 912. The specific method of sending the data 912 will be described below. In some embodiments, one or more of the control core 901, the immediate number generation module 906, and the immediate number database 907 may be included in the DIDT control unit 211. In other embodiments, one or more of the immediate number generation module 906 and the immediate number database 907 may also be modules separate from the DIDT control unit 211 and communicatively coupled.

[0113] FIG10 shows an example control flow 1000 of the control core 901 according to an embodiment of the present disclosure.

[0114] As shown in FIG10 , in step 1001, the system is reset. After the system is reset, register configuration can be performed in step 1002. For example, the number of clock cycles required for the first stage preheating, the second stage preheating, and the stretching stage (i.e., the duration of each stage) can be configured. Then, in step 1003, the arrival of a task can be waited for. For example, as shown in FIG10 , it can be determined that a task is currently arriving by determining that the task count wave_cnt = 1 and the previous task count pre_wave_cnt = 0.

[0115] After a task arrives, a first-stage preheating command may be sent in step 1004 to increase the power consumption (equivalently, current, or computing density) of the computing unit to a configured computing density level (also referred to herein as a third computing density level). The computing density level to which the first-stage preheating is expected to be increased may be preconfigured. By setting the first-stage preheating before the second-stage preheating described below, the operating state of the computing unit may be gradually increased to the preset computing density level earlier, thereby improving processing efficiency while resolving the problem of excessive DIDT.

[0116] In some embodiments, the format of the first stage warm-up command can be {current computing density level encoding, target computing density level encoding, required number of clock cycles}. In some embodiments, corresponding to the five instruction categories shown in Figure 6, the computing density level can be classified and encoded into five categories, such as low computing density level (the corresponding encoding can be 1), medium-low computing density level (the corresponding encoding can be 2), medium computing density level (the corresponding encoding can be 3), medium-high computing density level (the corresponding encoding can be 4) and high computing density level (the corresponding encoding can be 5). In addition, a zero computing density level (the corresponding encoding can be 0) can also be set to indicate the computing density level when the computing unit stops running or has no tasks or no instructions to execute. In this embodiment, an example of the first stage warm-up command can be {0, 3, 1000}, which can indicate that the computing density of the computing unit needs to be gradually linearly or step-by-step increased (for example, by one or more gradually increasing immediate values ​​803) from the zero computing density level with no instruction execution to the medium computing density level within 1000 clock cycles.

[0117] Then, in step 1005, the instruction can be waited for to be retrieved. Next, in step 1006, the computing density level (e.g., target computing density level) corresponding to the instruction to be executed subsequently can be determined based on the current task identifier wave_cur_id and the next task identifier wave_nxt_id (or based on the statistical information of all retrieved pending instructions or pending instructions of each task). Then, in step 1007, the computing density preheating level for the second stage can be determined based on the current computing density level and the target computing density level.

[0118] The second-stage compute density warm-up level can be the difference between the current compute density level and the target compute density level. The current compute density level can be directly obtained from the current power consumption or current operating state of the computing unit. The target compute density level can be determined by averaging or weighted averaging the instructions to be processed. For example, assuming that based on the current task identifier wave_cur_id and the next task identifier wave_nxt_id, as well as the statistical information shown in Figures 5 and 7, the instructions to be processed include 10 low-density computing instructions, 20 medium-low-density computing instructions, 50 medium-density computing instructions, 30 medium-high-density computing instructions, and 20 high-density computing instructions, the target compute density level can be determined by taking the following weighted average and rounding it down: floor((10*1+20*2+50*3+30*4+20*5) / (10+20+50+30+20))=floor(3.23)=3, where floor() is a floor function. That is, the target computing density level determined by the above formula corresponds to a code of 3 (medium computing density level). In other embodiments, rounding up or other calculation methods may be used to determine the current computing density level and the target computing density level (and their corresponding codes), which are not limited herein.

[0119] Then, a second-stage preheating command may be sent in step 1008 to preheat the power consumption (equivalently, current, or computing density) of the computing unit to a desired level (e.g., a target computing density level, which may also be referred to herein as a first computing density level). The format of the second-stage preheating command may be the same as the format of the first-stage preheating command. For example, assuming that the current computing density level is a medium computing density level (correspondingly encoded as 3), and the target computing density level determined by the method described above is a high computing density level (correspondingly encoded as 5), the second-stage preheating command may be {3, 5, 500}, which may indicate that the computing density of the computing unit needs to be gradually linearly or step-wise increased (e.g., by one or more gradually increasing immediate values ​​803) from the medium computing density level to the high computing density level within 500 clock cycles.

[0120] In step 1009, it can be determined whether all tasks and / or all instructions have been executed. If the instructions have not been executed, steps 1006-1008 can be executed in a loop until the instructions of all tasks are executed. For example, in step 1009, it can be determined whether all tasks and / or all instructions have been executed according to a predetermined clock cycle. In some embodiments, at each judgment time point, if it is determined that there are still instructions to be executed, a new target computing density level can be determined in steps 1006-1008 based on the pending instructions included in the current task and / or the next task, and a corresponding preheating is performed before actually executing these pending instructions (for example, during the acquisition of the operands involved in these pending instructions). Through this processing method, the computing density level of the computing unit can also be adaptively pre-adjusted based on the instructions to be executed (for example, the unfinished instructions included in the current task and the pending instructions included in the next task) during the task execution process, thereby further avoiding the sudden increase of DIDT and the impact on chip performance.

[0121] In step 1010, after all tasks and instructions are executed, a stretch command may be sent to gradually reduce the power consumption of the computing unit linearly or in steps (e.g., by one or more gradually decreasing immediate values ​​803) to a set level (herein, referred to as a second computing density level), for example, the zero computing density level as described above. After the stretch command is sent, process 1000 may return to 1003 to wait for the arrival of the next round of task sets. The format of the stretch command may be the same as or different from the format of the warm-up command.

[0122] As described above, the command 911 (for example, the first stage preheating command, the second stage preheating command, the stretching command, etc.) sent by the control core 901 can reach the immediate number generation module 906. After the immediate number generation module 906 receives the command 911, it can search the immediate number database 907 according to the current computing density level and the target computing density level to obtain one or more immediate number data corresponding to the computing density gradient (for example, the required computing density gradually increases or decreases) and gradually take the data and release it based on the corresponding clock timing or clock cycle, and provide it to the computing unit 206 through a structure such as that shown in Figure 8 for preheating or stretching the operating state of the computing unit 206.

[0123] FIG12 shows a graph showing the change in current over time of the computing unit 206 during the preheating phase according to an embodiment of the present disclosure. As shown in FIG12 , the change in the overall current of the computing unit 206 over time may be in a step-like manner. The corresponding multiple immediate values ​​may be released or provided using the M / N frequency division method. M may be the number of clock cycles required for the preheating phase, which may be directly obtained from the command 911 sent. N is the number of steps required to go from the current computing density level to the target computing density level. How N is calculated will be described below.

[0124] FIG11 shows an example of an immediate database 907 according to an embodiment of the present disclosure. The immediate database 907 can be obtained by modeling and simulating the power consumption of the computing unit. Generally, under a given frequency and process, the power consumption of the computing unit is mainly affected by the bit width, the number of operands, and the precision. Therefore, when modeling the immediate database 907, the computing density level of the instruction can be distinguished by the bit width and precision. For example, according to the five instruction precisions of Int8, FP8, FP16, FP32, and FP64, five corresponding sub-databases are established, namely, a low-density operation immediate database 1101, a medium-low density operation immediate database 1102, a medium-density operation immediate database 1103, a medium-high density operation immediate database 1104, and a high-density operation immediate database 1105. These five sub-databases can correspond to the low computing density level (correspondingly coded as 1), the medium-low computing density level (correspondingly coded as 2), the medium computing density level (correspondingly coded as 3), the medium-high computing density level (correspondingly coded as 4), and the high computing density level (correspondingly coded as 5) as described above. In addition, different operands within each computational density correspond to different power consumption. Therefore, when modeling, 64 immediate values ​​can be further selected according to the corresponding power consumption ladder at each computational density to form the entire data model, that is, a total of 320 immediate values, correspondingly numbered 0, 1, 2, ..., 319, as shown in Figure 11. In some embodiments, each of the sequence numbers 0, 1, 2, ..., 319 can correspond to an immediate value or an immediate group. For simplicity of explanation, it is assumed that each sequence number corresponds to an immediate value.

[0125] For example, if the command 911 sent by the control core 901 is {0, 3, 1000}, it means that the computational density of the computing unit needs to be gradually increased linearly or in steps from the zero computational density level where no instructions are executed to the medium computational density level within 1000 clock cycles. In this case, this can be achieved by relatively evenly providing the computing unit with multiple immediate numbers with corresponding computational densities gradually increasing in steps within 1000 clock cycles. For example, assuming that the middle position number in the medium density operation immediate number library 1103 (i.e., the 32nd immediate number therein) is used to represent the medium computational density level, the multiple immediate numbers can be a total of 64+64+32=160 (i.e., the N value described above) immediate numbers from the first immediate number (serial number 0) in the low density operation immediate number library 1101 to the 32nd immediate number (serial number 159) in the medium density operation immediate number library 1103. At this time, the immediate number generation module 906 can evenly perform 160 data searches, replacements, and outputs within 1000 clock cycles, allowing the power consumption of the computing unit to gradually increase to a medium level of medium computing density, as shown in FIG12 , so that the chip's voltage regulator (VR) and other devices have time to respond.

[0126] In another example, if the command 911 (which may be a warm-up command or a stretching command) sent by the control core 901 is {2, 4, 500}, it means that the computational density of the computing unit needs to be gradually linearly or stepped from a medium-low computational density level to a medium-high computational density level within 500 clock cycles. In this case, it can be achieved by relatively evenly providing the computing unit with a plurality of immediate numbers with corresponding computational densities gradually increasing in steps within 500 clock cycles. For example, assuming that the middle position number of each computational density level (i.e., the 32nd immediate number therein) is used to represent the computational density level, the plurality of immediate numbers can be a total of 32+64+32=128 immediate numbers, from the 32nd immediate number (serial number 95) in the medium-low density computational immediate number library 1102 to the 32nd immediate number (serial number 223) in the medium-high density computational immediate number library 1104. At this moment, the immediate number generation module 906 can evenly perform 128 times of search, replacement and output of data in 500 clock cycles, so that the power consumption of the computing unit is gradually stepped to a medium level of medium and high computing density. In other embodiments, the computing density level can also be further refined. For example, in a one-to-one correspondence with the 320 immediate numbers included in the immediate number database 907 as shown in Figure 11, the computing density can be further refined into 320 computing density levels. At this moment, the first immediate number among the multiple immediate numbers provided to the computing unit can directly correspond to the current computing density level, and the last immediate number can directly correspond to the target computing density level.

[0127] Figure 13 shows the current of the computing unit during operation after processing by the DIDT control unit according to an embodiment of the present disclosure. As shown in Figure 13, after the DIDT control of the embodiment of the present disclosure, the current of the entire circuit rises slowly, avoiding a sudden increase in the DIDT. In the figure, A is the circuit current change during the first stage of preheating, B is the circuit current change during the second stage of preheating, C is the current of the computing unit during normal calculation (this part may vary within a certain range depending on the number of calculations, but the amplitude is not large), and D or D' is a figure for stretching after a task set is completed. If a new task set arrives during the stretching process (for example, arrives and will be processed by the computing unit), then E is the re-preheating phase for the new task set, F is the normal calculation phase for the new task set, and G is the stretching phase after the calculation of the new task set is completed (assuming no new task set arrives). As described above, in this document, the time period corresponding to A can also be referred to as the third time period, the time period corresponding to B can also be referred to as the first time period, the time period corresponding to D or D' or G can also be referred to as the second time period, and the time period corresponding to E can also be referred to as the fourth time period. It should be understood that in some embodiments, the first stage preheating as shown in FIG. A may be omitted, and the second stage preheating as shown in FIG. B may be started directly.

[0128] Next, FIG14 shows a method 1400 for task processing according to an embodiment of the present disclosure.

[0129] As shown in Figure 14, a method 1400 for task processing according to an embodiment of the present disclosure may include: in step S1401, determining a first computing density level for processing one or more tasks; in step S1402, gradually increasing the computing density of a computing unit for processing one or more tasks to the first computing density level within a first time period before processing the one or more tasks; in step S1403, processing the one or more tasks through the computing unit; and in step S1404, gradually reducing the computing density of the computing unit to a second computing density level within a second time period after the processing of the one or more tasks is completed.

[0130] In some embodiments, the one or more tasks may be one or more second tasks as described above, or a task set including one or more tasks. The first computing density level may be a target computing density level to which a computing unit for processing the one or more tasks is desired to be preheated. The first time period may be a time period for preheating the operating state of the computing unit (e.g., the second preheating time period shown in B of FIG13 ), and the second time period may be a time period for stretching the operating state of the computing unit. On the other hand, the first time period may be a time period for obtaining instructions to be processed included in the one or more tasks.

[0131] In some embodiments, determining a first computing density level for processing one or more tasks may include: determining a statistical number of instructions to be processed included in the one or more tasks according to a classification statistical number of predetermined categories; and determining the first computing density level for processing the one or more tasks based on the classification statistical number.

[0132] As described above, in some embodiments, the predetermined categories may include one or more of the following: low-density operation instructions, medium-low-density operation instructions, medium-density operation instructions, medium-high-density operation instructions, and high-density operation instructions.

[0133] In some embodiments, method 1400 may further include: gradually increasing the computing density of the computing unit to a preset third computing density level in a third time period before the first time period. The third time period may be the first warm-up time period shown in A of FIG13 .

[0134] In some embodiments, within a first time period before processing the one or more tasks, gradually increasing the computational density of a computing unit for processing the one or more tasks to a first computational density level may include: within the first time period, sequentially providing the computing unit with one or more first preset data having gradually increasing computational density for processing by the computing unit, wherein a first one of the one or more first preset data may correspond to a current computational density level of the computing unit, and a last one of the one or more first preset data may correspond to the first computational density level. For example, as described above in conjunction with Figures 10-12, the one or more first preset data may be one or more immediate values ​​from the immediate value database 907 (or pre-stored in the immediate value database 907).

[0135] In some embodiments, within a second time period after processing of one or more tasks is completed, gradually reducing the computational density of the computing unit to a second computational density level may include: within the second time period, sequentially providing one or more second preset data with gradually decreasing computational density to the computing unit for processing by the computing unit, wherein a first second preset data in the one or more second preset data may correspond to the current computational density level of the computing unit, and a last second preset data in the one or more second preset data may correspond to the second computational density level. For example, as described above in conjunction with Figures 10-12, the one or more second preset data may be one or more immediate values ​​from the immediate value database 907 (or pre-stored in the immediate value database 907).

[0136] In some embodiments, method 1400 may further include: if one or more new tasks arrive within the second time period, stopping the gradual reduction of the computing density; and gradually increasing the computing density of the computing unit to a fourth computing density level within a fourth time period starting from the time when the gradual reduction of the computing density is stopped. In some embodiments, the fourth computing density level may be a computing density level for processing the new one or more tasks.

[0137] FIG15 shows a schematic diagram of an apparatus 1500 for task processing according to an embodiment of the present disclosure.

[0138] As shown in Figure 15, an apparatus 1500 for task processing according to an embodiment of the present disclosure may include: a determination unit 1501, which can be configured to determine a first computing density level for processing one or more tasks; a first preheating unit 1502, which can be configured to gradually increase the computing density of the computing unit used to process one or more tasks to the first computing density level within a first time period before processing the one or more tasks; a computing unit 1503, which can be configured to process one or more tasks through the computing unit; and a stretching unit 1504, which can be configured to gradually reduce the computing density of the computing unit to a second computing density level within a second time period after the processing of the one or more tasks is completed.

[0139] In some embodiments, the one or more tasks may be one or more tasks or a second task as described above, or a task set including one or more tasks. The first computing density level may be a target computing density level to which a computing unit for processing the one or more tasks is desired to be preheated. The first time period may be a time period for preheating the operating state of the computing unit (e.g., the second preheating time period shown in B of FIG13 ), and the second time period may be a time period for stretching the operating state of the computing unit. On the other hand, the first time period may be a time period for obtaining instructions to be processed included in the one or more tasks.

[0140] In some embodiments, the determination unit 1501 is configured to determine a first computing density level for processing one or more tasks, which may include: determining a classified statistical number of instructions to be processed included in the one or more tasks according to a predetermined category; and determining the first computing density level for processing the one or more tasks based on the classified statistical number.

[0141] As described above, in some embodiments, the predetermined categories may include one or more of the following: low-density operation instructions, medium-low-density operation instructions, medium-density operation instructions, medium-high-density operation instructions, and high-density operation instructions.

[0142] In some embodiments, the apparatus 1500 may further include a second warm-up unit configured to gradually increase the computing density of the computing unit to a preset third computing density level within a third time period prior to the first time period. The third time period may be the first warm-up time period shown in A of FIG13 .

[0143] In some embodiments, the first preheating unit 1502 is configured to gradually increase the computing density of the computing unit used to process the one or more tasks to a first computing density level within a first time period before processing the one or more tasks, which may include: providing one or more first preset data with gradually increasing computing density to the computing unit in sequence within the first time period for processing by the computing unit, wherein the first first preset data in the one or more first preset data may correspond to the current computing density level of the computing unit, and the last first preset data in the one or more first preset data may correspond to the first computing density level. For example, as described above in conjunction with Figures 10-12, the one or more first preset data may be one or more immediate values ​​from the immediate value database 907 (or pre-stored in the immediate value database 907).

[0144] In some embodiments, the stretching unit 1504 is configured to gradually reduce the computational density of the computing unit to a second computational density level within a second time period after the processing of the one or more tasks is completed. This may include: sequentially providing one or more second preset data with gradually decreasing computational density to the computing unit within the second time period for processing by the computing unit, wherein a first second preset data in the one or more second preset data may correspond to the current computational density level of the computing unit, and a last second preset data in the one or more second preset data may correspond to the second computational density level. For example, as described above in conjunction with Figures 10-12, the one or more second preset data may be one or more immediate values ​​from the immediate value database 907 (or pre-stored in the immediate value database 907).

[0145] In some embodiments, the stretching unit 1504 may be further configured to stop gradually reducing the computing density if one or more new tasks arrive within the second time period. In some embodiments, the first preheating unit 1502 may be further configured to gradually increase the computing density of the computing unit to a fourth computing density level within a fourth time period starting from the time when the gradual reduction of the computing density is stopped. In some embodiments, the fourth computing density level may be a computing density level for processing the new one or more tasks.

[0146] FIG16 shows a schematic diagram of another apparatus 1600 for task processing according to an embodiment of the present disclosure.

[0147] As shown in FIG16 , an apparatus 1600 for task processing according to an embodiment of the present disclosure may include at least one processor 1601 and a storage device 1602. The storage device 1602 may store at least one program, which, when executed by the at least one processor 1601, may enable the at least one processor 1601 to implement any method according to an embodiment of the present disclosure, such as the method 1400 for task processing described above and any other method described in conjunction with FIG1 to FIG14 .

[0148] In addition, an embodiment of the present disclosure also provides an electronic device, which may include an apparatus 1500 or 1600 for task processing according to an embodiment of the present disclosure, and may execute or implement any method according to an embodiment of the present disclosure, such as the above-mentioned method 1400 for task processing and any other method described in combination with Figures 1-14.

[0149] It should be understood that the steps in the various methods or processes described in this disclosure in conjunction with the accompanying drawings are merely examples, and any one of them may be deleted from the corresponding method or process, added to other methods or processes, combined with each other in any manner, or performed in any suitable order.

[0150] In particular, according to an embodiment of the present disclosure, the method or process described above in conjunction with the embodiments of the present disclosure or the accompanying drawings can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by a processing device, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0151] Furthermore, an embodiment of the present disclosure provides a computer-readable storage medium having instructions stored thereon, which, when executed, cause a processor to perform any method for task processing according to an embodiment of the present disclosure.

[0152] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0153] The units involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a unit does not necessarily limit the unit itself.

[0154] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0155] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0156] The above description is merely an optional embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.

[0157] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0158] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A method for task processing, comprising: determining a first computational density level for processing one or more tasks; In a first time period before processing the one or more tasks, gradually increasing the computing density of computing units used to process the one or more tasks to the first computing density level; Processing the one or more tasks by the computing unit; as well as In a second time period after the processing of the one or more tasks is completed, the computing density of the computing unit is gradually reduced to a second computing density level.

2. The method according to claim 1, wherein: Determining a first computing density level for processing one or more tasks includes: Determining the statistical number of instructions to be processed included in the one or more tasks classified according to predetermined categories; and A first computational intensity level for processing the one or more tasks is determined based on the classification statistics.

3. The method according to claim 2, wherein: The predetermined categories include one or more of the following: Low-density computing instructions, medium-low-density computing instructions, medium-density computing instructions, medium-high-density computing instructions, and high-density computing instructions.

4. The method according to any one of claims 1 to 3, further comprising: In a third time period before the first time period, the computing density of the computing unit is gradually increased to a preset third computing density level.

5. The method according to any one of claims 1 to 4, wherein: The first time period is a time period for acquiring instructions to be processed included in the one or more tasks.

6. The method according to any one of claims 1 to 5, wherein: In a first time period before processing the one or more tasks, gradually increasing the computing density of a computing unit for processing the one or more tasks to the first computing density level includes: In the first time period, one or more first preset data with gradually increasing computing density are sequentially provided to the computing unit for processing by the computing unit, wherein the first first preset data of the one or more first preset data is consistent with the current computing density level of the computing unit. Correspondingly, and the last first preset data among the one or more first preset data corresponds to the first calculation density level.

7. The method according to claim 6, wherein: The one or more first preset data are pre-stored in a database.

8. The method according to any one of claims 1 to 7, wherein: In a second time period after the processing of the one or more tasks is completed, gradually reducing the computing density of the computing unit to a second computing density level includes: During the second time period, one or more second preset data with gradually decreasing computing density are provided to the computing unit in sequence for processing by the computing unit, wherein the first second preset data among the one or more second preset data corresponds to the current computing density level of the computing unit, and the last second preset data among the one or more second preset data corresponds to the second computing density level.

9. The method according to claim 8, wherein: The one or more second preset data are pre-stored in the database.

10. The method according to any one of claims 1 to 9, further comprising: When one or more new tasks arrive within the second time period, stopping the gradual reduction of the computing density; as well as In a fourth time period starting from the time when the stepwise reduction of the computing density is stopped, the computing density of the computing unit is gradually increased to a fourth computing density level, The fourth computing density level is a computing density level used to process the new one or more tasks.

11. A device for task processing, comprising: a determining unit configured to determine a first computing density level for processing one or more tasks; A first preheating unit is configured to gradually increase the computing density of a computing unit used for processing the one or more tasks to the first computing density level within a first time period before processing the one or more tasks; a computing unit configured to process the one or more tasks by the computing unit; as well as The stretching unit is configured to gradually reduce the computing density of the computing unit to a second computing density level within a second time period after the processing of the one or more tasks is completed.

12. The device according to claim 11, wherein The determining unit is configured to determine a first computing density level for processing one or more tasks, comprising: Determining the statistical number of instructions to be processed included in the one or more tasks classified according to predetermined categories; and A first computational intensity level for processing the one or more tasks is determined based on the classification statistics.

13. The device according to claim 12, wherein: The predetermined categories include one or more of the following: Low-density computing instructions, medium-low density computing instructions, medium-density computing instructions, medium-high density computing instructions, and high-density computing instructions.

14. The device according to any one of claims 11 to 13, further comprising: The second preheating unit is configured to gradually increase the computing density of the computing unit to a preset third computing density level in a third time period before the first time period.

15. The device according to any one of claims 11 to 14, wherein: The first time period is a time period for acquiring instructions to be processed included in the one or more tasks.

16. The device according to any one of claims 11 to 15, wherein: The first preheating unit is configured to gradually increase the computing density of the computing unit used to process the one or more tasks to the first computing density level within a first time period before processing the one or more tasks, including: During the first time period, one or more first preset data with gradually increasing computing density are provided to the computing unit in sequence for processing by the computing unit, wherein the first first preset data among the one or more first preset data corresponds to the current computing density level of the computing unit, and the last first preset data among the one or more first preset data corresponds to the first computing density level.

17. The device according to any one of claims 11 to 16, wherein: The stretching unit is configured to gradually reduce the computing density of the computing unit to a second computing density level within a second time period after the processing of the one or more tasks is completed, comprising: In the second time period, one or more second preset data with gradually decreasing computing density are sequentially provided to the computing unit for processing by the computing unit, wherein the first second preset data of the one or more second preset data is consistent with the current computing density level of the computing unit. Correspondingly, and the last second preset data among the one or more second preset data corresponds to the second calculation density level.

18. The device according to any one of claims 11 to 17, wherein: The stretching unit is further configured as: When one or more new tasks arrive within the second time period, stopping the gradual reduction of the computing density; and The first preheating unit is further configured to: gradually increase the computing density of the computing unit to a fourth computing density level within a fourth time period starting from the time when the gradual decrease of the computing density is stopped, The fourth computing density level is a computing density level used to process the new one or more tasks.

19. An apparatus for task processing, comprising: at least one processor; and A storage device having at least one program stored thereon, which, when the at least one program is executed by the at least one processor, enables the at least one processor to implement the method according to any one of claims 1 to 10.

20. An electronic device comprising the device for task processing according to any one of claims 11 to 18 or 19.

21. A non-transitory storage medium containing computer executable instructions, wherein: The computer executable instructions are used to perform the method according to any one of claims 1 to 10 when executed by a computer processor.

Citation Information

Patent Citations

  • Task processing method, equipment, device and storage medium

    CN117762614A

  • Method, system and device for improving chip computing performance and medium

    CN111176731A

  • Methods and apparatuses for reducing step loads of processors

    US20090070607A1

  • Control scheme to temporarily raise supply voltage in response to sudden change in current demand

    US20140380066A1